Useful Google IO 2012 Talks

Google I/O 2012 - The Web Can Do That!? - YouTube

I wasn't expecting HTML5 to contribute to such an inspiring Google I/O 2012 video!!

Useful link:
Media Capture and Streams

That was one of the best talks. Here's the slides: The Web Can Do That!? - Google IO 2012

Also, extra awesome nerdsauce: the slides are built on AngularJS. Oh yeah.

Oh, and, while unrelated to Google IO, this is an intermediary in my own research after watching the V8 performance video:

http://blog.mrale.ph

V8 is incredibly powerful, and I'm just now realizing just how much I've abused it, when I should really be using my own debugging/assembly knowledge to understand its internals and code optimization to try to generate highly-efficient code for V8 (I only code for Chrome and Node.js these days, which leaves me just working with V8).
 
Last edited:
That was one of the best talks. Here's the slides: The Web Can Do That!? - Google IO 2012

Also, extra awesome nerdsauce: the slides are built on AngularJS. Oh yeah.

Oh, and, while unrelated to Google IO, this is an intermediary in my own research after watching the V8 performance video:

http://blog.mrale.ph

V8 is incredibly powerful, and I'm just now realizing just how much I've abused it, when I should really be using my own debugging/assembly knowledge to understand its internals and code optimization to try to generate highly-efficient code for V8 (I only code for Chrome and Node.js these days, which leaves me just working with V8).


I know what you mean about abusing V8- me too. That video you posted "breaking the JavaScript Speed Limit" is great, however, I'm going to have to really re-learn JavaScript for V8 opts for a project Timebomb and I are working on which, of course, involves V8 and friends. Thanks for that link.

I say, if you're still thinking about doing basic css/html5 tutorials, well, it would be great for me if you would put together some cutting-edge V8 tutorials :D
 
It's quite involved. I need to actually read the v8 source to see what causes contexts to get thrown into dictionary mode and exactly what triggers dictionary mode in objects. I don't believe v8 has a flag that will indicate when an object is turned into dictionary mode, but that's one of the huge performance points. Perhaps you can do it with a heap analyzer but that seems a bit involved. Maybe I can patch v8 to have a way to better explain when objects are dictionary-moded, in a perf-like output that shows in what function it occured, etc. Here's 2 examples I came up with last night that blew my mind:

Code:
function Foo() {
    this.abc = 7;
}
Foo.prototype.test = function (n) {
    this.abc = n;
}

function Bar() {
    this.abc = 7;  // I am unsure at this point if Foo and Bar share a hidden class, that's also a curious point
}

Bar.prototype = {
    test: function (n) {
        this.abc = n;
    }
};

function runTest(cur, name) {
    var timer = name + ' speed';
    console.time(timer);
    // If Foo and Bar objects share a hidden class, this function is monomorphic
    for (var i = 0; i < 10000000; ++i) {
        cur.test(i); // Should remain SMI since 10M < 1073741824
    }
    console.timeEnd(timer);
}

var cur = new Foo();
runTest(cur, 'foo');

cur = new Bar();
runTest(cur, 'bar');

If you look at the above code, you understand Foo and Bar are functionally and logically IDENTICAL. The only thing that differs between them is the construction of their prototype. The output of this test here on my machine is:

foo speed: 16ms
bar speed: 140ms

What that says to me, very obviously, is that constructing an object with functions as members instantly drops it into dictionary mode. I confirmed this on the v8-users discussion board with a quick search. It's an edge case, and I'm really hoping they optimize this sometime because this is a REALLY common pattern to create prototypes.

Also, @ hidden class question in code, running with --trace-deopt shows:

Code:
[deoptimize context: fef14679]foo speed: 19ms
**** DEOPT: runTest at bailout #7, address 0x0, frame size 96
[deoptimizing: begin 0xfff48c89 runTest @7]
  translating runTest => node=41, height=16
    0x0025f128: [top + 64] <- 0xd9306451 ; rdi 00000000D9306451 <JS Global Objec
t>
    0x0025f120: [top + 56] <- 0xfff6b581 ; rcx 00000000FFF6B581 <a Bar>
    0x0025f118: [top + 48] <- 00000000D9304121 <undefined> ; literal
    0x0025f110: [top + 40] <- 0xb73902d8 ; caller's pc
    0x0025f108: [top + 32] <- 0x0025f168 ; caller's fp
    0x0025f100: [top + 24] <- 0xfff107d1; context
    0x0025f0f8: [top + 16] <- 0xfff48c89; function
    0x0025f0f0: [top + 8] <- 0xfff6b601 ; rbx 00000000FFF6B601 <String[9]: bar s
peed>
    0x0025f0e8: [top + 0] <- 0 ; rax (smi)
[deoptimizing: end 0xfff48c89 runTest => node=41, pc=0xb73906d3, state=NO_REGIST
ERS, alignment=no padding, took 10.000 ms]
[removing optimized code for: runTest]
bar speed: 140ms

Which would indicate that no, they do not share a hidden class at all, which is a pretty important point. The objects have identical shapes with different prototype pointers so they have different hidden classes. I'm not sure if two objects with different prototype chains can share a hidden class, though I'll try to see if I can hack a way to do it with Object.create sometime.

This is a really, really, really, really, really, really important point, as it's the primary reason code gets really bad. You're dropping into a polymorphic state instead of a monomorphic one. This should be avoided as much as possible. For example, if I repeat both of those loops after the initial Foo/Bar test, we'll be running the tests again with a polymorphic non-optimized function, which gives us this result:

foo speed: 18ms
bar speed: 119ms
foo speed: 128ms
bar speed: 165ms

Ouch. Don't do that. Don't ever do that. Oddballs can cause this too.

As for the second point, it has to do with closures and their effect on the parent scope. At every invocation of the parent function, it optimistically binds the closed variable into a closure scope. The following shows two functionally equivalent pieces of code, one with a closure, one without, notice the very existence of a closure results in a huge performance slowdown, even if the closure is never used (and a good static analysis would reveal that this closure is never called). It's important to note that a variable bound to a closure context is accessed in that context, not as a local variable anymore (it's boxed into a closure context which I believe is essentially an object, which adds overhead as it isn't an inline register access or stack access anymore):

Code:
function testPresence(N) {
    var z = 7;
    function closureFunction(n) {
        z = n;
    }
    console.time('presence speed');
    for (var i = 0; i < N; ++i) {
        z = i;
    }
    console.timeEnd('presence speed');
}


function testAbsence(N) {
    var z = 7;
    function nonclosureFunction(n, scope) {
        scope.z = n;
    }
    console.time('absence speed');
    for (var i = 0; i < N; ++i) {
        z = i;
    }
    console.timeEnd('absence speed');
}


(function () {
    var i = 0,
        origLog = console.log,
        lastSpeed = -1,
        speeds = [[],[]],
        N = 10000000,
        M = 50;
    console.log = function (x, y, b) {
        lastSpeed = b;
    };


    for (var i = 0; i < M; i++) {
        testPresence(N);
        speeds[0].push(lastSpeed);
        testAbsence(N);    
        speeds[1].push(lastSpeed);
    }
    function calcMean(arr) {
        return arr.reduce(function (prev, current, index) {
            if (-1 === prev) {
                return current;
            }
            return prev + ((current - prev)/index)
        }, -1);
    }
    origLog('Presence: ', Math.round(100*calcMean(speeds[0]))/100, 'ms');
    origLog('Absence: ', Math.round(100*calcMean(speeds[1]))/100, 'ms');
    console.log = origLog;
}());

Result:

Presence: 20.86 ms
Absence: 12.67 ms

Again, I really need to read a LOT more into v8's code to understand what exactly it does so I can better know what to do and not to do and how to optimize code like this. These are both extremely significant examples in hot code (in general, not a big deal, but in a performance-critical hot loop or something that you need to run very fast, it's huge).

I believe binding closures against the global scope also caused dramatic performance issues, but as of the latest V8 build, that issue seems to be resolved (running Node 0.8.1 vs 0.8.2 vs d8). They're constantly improving performance (and sometimes breaking it), so it's kindof a moving target. Once I know everything about how V8 optimizes JS, I'd need to keep informed on patches to know just how they're modifying things.

I'm also considering modifying crankshaft on function optimization to have an opt-in feature to check for tail recursion (not generic tail call elimination, as that's WAY more complicated). If the function is tail-recursive, add a test to verify that the recurse is self otherwise deoptimize, else do not push a new frame on the stack, and simply jmp back to the beginning of the function. That would be opt-in from a runtime switch and would only work in strict mode because access to arguments.caller or callee would result in an exception.

Oh, and to optimize that runTest, you'd create 2 functions:

Code:
// Monomorphic on cur.klass = Foo, name.klass = String
function runTestFoo(cur, name) {
    var timer = name + ' speed';
    console.time(timer);
    // If Foo and Bar objects share a hidden class, this function is monomorphic
    for (var i = 0; i < 10000000; ++i) {
        cur.test(i); // Should remain SMI since 10M < 1073741824
    }
    console.timeEnd(timer);
}

// Monomorphic on cur.klass = Bar, name.klass = String
function runTestBar(cur, name) {
    var timer = name + ' speed';
    console.time(timer);
    // If Foo and Bar objects share a hidden class, this function is monomorphic
    for (var i = 0; i < 10000000; ++i) {
        cur.test(i); // Should remain SMI since 10M < 1073741824
    }
    console.timeEnd(timer);
}

// runTest will be polymorphic, but it's not hot, so we don't care!
function runTest(cur, name) {
    if (cur instanceof Foo) {
        return runTestFoo(cur, name);
    }
    return runTestBar(cur, name);
}



Or just explicitly call runTestFoo/Bar, or bind the runTest to the prototype so each class has its own runType so you're forced to use the right hidden class. And you can hide those 2 functions behind function scope so nobody can cheat, but if someone is dick-move enough to modify the klass of their Foo or Bar by modifying the object directly, they can go to hell:

Code:
cur = new Bar();
cur.lol = 1;
runTest(cur, 'bar');
delete cur.lol;
runTest(cur, 'bar');

bar speed: 183ms
bar speed: 270ms
 
Last edited:
For your first implementation I get this:
Javascript Mozilla Pastebin - collaborative debugging tool
Code:
foo speed: 16ms
bar speed: 61ms
baz speed: 76ms
qux speed: 75ms
foo speed: 82ms
bar speed: 81ms
baz speed: 75ms
qux speed: 76ms
I noticed that in my version your conclusions were almost the same. I noticed that once bar is polymorphic, it becomes the same as foo, as we expected bar to be in the first place. But something new, if you add the method to this inside of the constructor, or if you add the property to the prototype instead of this (basically, if either the constructor or the prototype contain more than one property- counting methods), then it acts polymorphic the first and every time you use it. but just maybe a wee bit faster than if you put the properties in different places.

Testing the other ones..

------ EDIT -------

I feel like a Moron... This is weird, If you run each of the Foo, Bar, Baz, and Qux examples individually, they are exactly the same.

Code:
foo speed: 16ms
bar speed: 17ms
baz speed: 16ms
qux speed: 16ms

Javascript Mozilla Pastebin - collaborative debugging tool
 
Last edited:
For your first implementation I get this:
Javascript Mozilla Pastebin - collaborative debugging tool
Code:
foo speed: 16ms
bar speed: 61ms
baz speed: 76ms
qux speed: 75ms
foo speed: 82ms
bar speed: 81ms
baz speed: 75ms
qux speed: 76ms
I noticed that in my version your conclusions were almost the same. I noticed that once bar is polymorphic, it becomes the same as foo, as we expected bar to be in the first place. But something new, if you add the method to this inside of the constructor, or if you add the property to the prototype instead of this (basically, if either the constructor or the prototype contain more than one property- counting methods), then it acts polymorphic the first and every time you use it. but just maybe a wee bit faster than if you put the properties in different places.

Testing the other ones..

It's not the class or instance that becomes polymorphic, it's the function. In this case, runTest's first parameter is how it decides what to do. The IC we're really hoping gets optimized is cur.test, we want cur.test to be a direct lookup, because we do it 10 million times, so it's extremely hot. When you pass a 'cur' with a different hidden class into runTest, that IC that has been optimized is no longer valid, so it bails out and jumps back into the non-optimized version and then re-optimizes to be polymorphic (which is significantly slower).

You can polymorph on the prototype's string to create custom runTest copies for each prototype you pass in, but again, hidden classes can be modified outside of constructors. We don't have access to hidden classes in JS, and unfortunately sealing and preventing extensions of an object puts it in dictionary mode (which would be one way of forcing hidden class). That's another change I hope V8 makes, since we can use those tools to lock an object into a given hidden class and make very fast code (essentially, we already can implement types in an untyped language with those 2 methods, with the assumption that prototype name = unique).

Here's my modification to your runTest and the results:

Code:
// Polymorphic, slow dispatcher
var protoNameRe = /^function (\w+)\(/,
    fastTests = {};
function runTest(cur, name) {
    var protoName = protoNameRe.exec(cur.constructor)[1],
        timer = name + ' speed',
        fastTest = fastTests[protoName];
    if (!fastTest) {
        fastTest = fastTests[protoName] = function (cur) {
            for (var i = 0; i < 10000000; ++i) {
                cur.test(i);
            }
        }
    }
    console.time(timer);
    fastTest(cur);
    console.timeEnd(timer);
}

foo speed: 16ms
bar speed: 135ms
baz speed: 41ms
qux speed: 45ms
foo speed: 14ms
bar speed: 133ms
baz speed: 38ms
qux speed: 39ms
 
It's not the class or instance that becomes polymorphic, it's the function. In this case, runTest's first parameter is how it decides what to do. The IC we're really hoping gets optimized is cur.test, we want cur.test to be a direct lookup, because we do it 10 million times, so it's extremely hot. When you pass a 'cur' with a different hidden class into runTest, that IC that has been optimized is no longer valid, so it bails out and jumps back into the non-optimized version and then re-optimizes to be polymorphic (which is significantly slower).

You can polymorph on the prototype's string to create custom runTest copies for each prototype you pass in, but again, hidden classes can be modified outside of constructors. We don't have access to hidden classes in JS, and unfortunately sealing and preventing extensions of an object puts it in dictionary mode (which would be one way of forcing hidden class). That's another change I hope V8 makes, since we can use those tools to lock an object into a given hidden class and make very fast code (essentially, we already can implement types in an untyped language with those 2 methods, with the assumption that prototype name = unique).

Here's my modification to your runTest and the results:

Code:
// Polymorphic, slow dispatcher
var protoNameRe = /^function (\w+)\(/,
    fastTests = {};
function runTest(cur, name) {
    var protoName = protoNameRe.exec(cur.constructor)[1],
        timer = name + ' speed',
        fastTest = fastTests[protoName];
    if (!fastTest) {
        fastTest = fastTests[protoName] = function (cur) {
            for (var i = 0; i < 10000000; ++i) {
                cur.test(i);
            }
        }
    }
    console.time(timer);
    fastTest(cur);
    console.timeEnd(timer);
}

foo speed: 16ms
bar speed: 135ms
baz speed: 41ms
qux speed: 45ms
foo speed: 14ms
bar speed: 133ms
baz speed: 38ms
qux speed: 39ms

Here's where our configuration similarities differ:
Code:
runTest(new Foo(), 'foo'); // foo speed: 55ms
runTest(new Bar(), 'bar'); // bar speed: 66ms
runTest(new Baz(), 'baz'); // baz speed 61ms
runTest(new Qux(), 'qux'); // qux speed 82ms


runTest(new Foo(), 'foo'); // foo speed: 53ms
runTest(new Bar(), 'bar'); // bar speed: 81ms
runTest(new Baz(), 'baz'); // baz speed 89ms
runTest(new Qux(), 'qux'); // qux speed 85ms

And if you comment all of those out except one of the Qux (which is the largest) I get closer to the same as foo:
Code:
qux speed: 56ms

Same goes for foo, bar, and baz.

I'm not getting any closer as to how this works........
 
What are you running the tests in?

Also, Baz and Foo should have very similar performance metrics (Baz may be slightly faster). Though, for some reason, Baz suffers from the same self-deoptimization as Qux (described below). This is probably a bug in my version of v8, though. test doesn't modify the object so it shouldn't change the hidden class.

Bar is really slow no matter what.

Qux is tricky. You start by putting the default value on the prototype, and then set 'this.x'. It will look at 'this' and see there's no own property called 'x' so it'll add it to the object, changing its hidden class. But because the function isn't optimized until the function gets hot, the IC will tell the optimizer later that it's stable, so it'll optimize it for the case where you've added a this.x. Now if you run the test on a new Qux object, the first iteration will force the hot loop to be polymorphed:

Code:
runTest(new Qux(), 1, 'qux'); // qux speed: 19ms
runTest(new Qux(), 1, 'qux'); // qux speed: 91ms

Two points: avoid slow access mode in prototypes and keep hot loops monomorphic.
 
I am using node.js v0.6.7
"* Upgrade V8 to 3.6.6.15", https://raw.github.com/joyent/node/master/ChangeLog

Let me upgrade...
code:
http://pastebin.mozilla.org/1707943
Mk now I'm up-to-date:

Observing many:
Code:
spence@Shock:~/projects/v8/test3$ node -v
v0.8.2
spence@Shock:~/projects/v8/test3$ node run
foo speed: 55ms
bar speed: 80ms
baz speed: 54ms
qux speed: 89ms
foo speed: 55ms
bar speed: 119ms
baz speed: 85ms
qux speed: 91ms

Observing only one at a time:
Code:
spence@Shock:~/projects/v8/test3$ node run
foo speed: 55ms
spence@Shock:~/projects/v8/test3$ node run
bar speed: 82ms
spence@Shock:~/projects/v8/test3$ node run
baz speed: 48ms
spence@Shock:~/projects/v8/test3$ node run
qux speed: 55ms

And if we compact into one script without using require() at all,
http://pastebin.mozilla.org/1707944

Code:
foo speed: 13ms
bar speed: 79ms
baz speed: 30ms
qux speed: 90ms
foo speed: 13ms
bar speed: 119ms
baz speed: 87ms
qux speed: 92ms

Code:
spence@Shock:~/projects/v8/test3$ node run
foo speed: 13ms
spence@Shock:~/projects/v8/test3$ node run
bar speed: 81ms
spence@Shock:~/projects/v8/test3$ node run
baz speed: 14ms
spence@Shock:~/projects/v8/test3$ node run
qux speed: 13ms
I feel like I can make more sense out of these numbers.. Well, not really.. Why come our results are different? I'm on Linux 64-bit.
 
Last edited:
Back