Repository navigation
Cut the allocations of a tracking frame and of a TrackedValue - #21650
NullVoxPopuli-ai-agent wants to merge 5 commits into
Conversation
aae7903 to
c59a238
Compare
|
moved to draft while I explore if the specifics of this PR are a good idea |
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
…it later Blocks and component cache groups pass their previous tag to endTrackFrame(), so a block that re-renders with the same dependencies keeps its tag and its memoized revision. The tag comparison against the tracker's live entries comes from emberjs#21650, which is below this commit. A guard that fails is dropped on the spot, with no tracking frame: a block that changed is likely to change again, and a frame on every miss is what made all-rows-change workloads slower. A dropped guard comes back after eight unguarded updates, so a block that changed once and then stayed still is skipped again, while a block that changes on every update pays for one frame in every eight. A re-render re-arms the guard right away. An unguarded block costs what it did before guards existed: one frame push. A re-render of a dropped block opens its own tracking frame, because the append VM closes one when the block exits. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012wPae8QnFbKHKn5xbgaQVq
emberjs#21650 now has the commit that makes the value field of a TrackedValue general (bd88a65). This branch had the same change as f04dc41. The merge changes no file content. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> # Conflicts: # packages/@glimmer/validator/lib/tracked-value.ts
|
The store of "Before" is Memory: no change. Bytes that V8 allocates for one operation on one tracked value, for 11 kinds of value:
Time: no change. Geometric mean of after / before for 1,000 tracked values of each kind, 8 rounds. "Control" is the code of "before" from a second build, so it shows the noise.
The event that the store removes. A read loop is hot, then one other cell gets the first number. The time is what the next 5,000 read loops lose.
One group where "after" is slower, and why the store is not the causeAfter the event, the read loop is 3% to 4% slower in "after". That is 0.06 ns to 0.09 ns for each read. Two more builds separate the store from the deopt:
The kinds of value, and the method
The mitata codeThe files are in the research folder, with the full tables for each case. node research/tracked-value-field/micro-run.mjs --cpu=1 --rounds=8 \
--sources=before=<ember-source>,after=<ember-source>
/**
* The kinds of value that the micro cases store in a tracked value.
*/
export const COUNT = 1000;
/**
* Each read loop does the work that fits the kind of value,
* so that V8 has a reason to use what it knows about the field.
*/
function sumNumber(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) n += cells[i].value;
return n;
}
function sumLength(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) n += cells[i].value.length;
return n;
}
function sumField(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) n += cells[i].value.i;
return n;
}
function countTrue(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) if (cells[i].value === true) n++;
return n;
}
function countUndefined(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) if (cells[i].value === undefined) n++;
return n;
}
function countFunction(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) if (typeof cells[i].value === 'function') n++;
return n;
}
const PROBE = Symbol('probe');
function countProbe(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) if (cells[i].value === PROBE) n++;
return n;
}
function countBig(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) if (cells[i].value > 500n) n++;
return n;
}
function countNumber(cells) {
let n = 0;
for (let i = 0; i < cells.length; i++) if (typeof cells[i].value === 'number') n++;
return n;
}
/**
* `a` is the first value of cell `i`. A write pass stores `b`, then `a` again.
* `number` says that the kind stores a number in the field without help.
*/
export const kinds = [
{ name: 'small integer', number: true, a: (i) => i, b: (i) => i + 1, read: sumNumber },
{ name: 'double', number: true, a: (i) => i + 0.5, b: (i) => i + 1.5, read: sumNumber },
{ name: 'string', a: (i) => `a${i}`, b: (i) => `b${i}`, read: sumLength },
{ name: 'boolean', a: (i) => i % 2 === 0, b: (i) => i % 2 === 1, read: countTrue },
{
name: 'undefined and null',
a: (i) => (i % 2 === 0 ? undefined : null),
b: (i) => (i % 2 === 0 ? null : undefined),
read: countUndefined,
},
{ name: 'object', a: (i) => ({ i }), b: (i) => ({ i: i + 1 }), read: sumField },
{ name: 'array', a: (i) => [i], b: (i) => [i, i], read: sumLength },
{ name: 'function', a: (i) => () => i, b: (i) => () => i + 1, read: countFunction },
{
name: 'symbol',
a: (i) => (i % 2 === 0 ? PROBE : Symbol()),
b: (i) => (i % 2 === 1 ? PROBE : Symbol()),
read: countProbe,
},
{ name: 'bigint', a: (i) => BigInt(i), b: (i) => BigInt(i + 1), read: countBig },
];
const mixedKinds = kinds.slice();
kinds.push({
name: 'all kinds mixed',
number: true,
a: (i) => mixedKinds[i % mixedKinds.length].a(i),
b: (i) => mixedKinds[(i + 1) % mixedKinds.length].b(i),
read: countNumber,
});
/**
* Measures one micro case of `tracked(value)` in this process.
*
* `micro-run.mjs` starts this file one time per build, per case, per round,
* so the `#value` field of the class sees only the values of one case.
*
* EMBER_SOURCE=<ember-source folder> node --expose-gc micro.mjs --index=0
* node micro.mjs --list
*/
import { writeFileSync } from 'node:fs';
import { parseArgs } from 'node:util';
import { measure } from 'mitata';
import { COUNT, kinds } from './kinds.mjs';
const { values: args } = parseArgs({
options: {
index: { type: 'string' },
list: { type: 'boolean', default: false },
out: { type: 'string' },
'min-cpu-ms': { type: 'string', default: '500' },
},
});
const cases = [];
for (let kind of kinds) {
cases.push({ kind, op: 'create' }, { kind, op: 'read' }, { kind, op: 'write' });
}
for (let kind of kinds) {
if (kind.number) continue;
cases.push({ kind, op: 'read, after the first number' }, { kind, op: 'first number' });
}
for (let c of cases) c.name = `${c.kind.name}: ${c.op}`;
if (args.list) {
console.log(JSON.stringify(cases.map((c) => c.name)));
process.exit(0);
}
const { load } = await import('../../adapters/ember-source.mjs');
const { tracked } = await load('@glimmer/tracking');
const { default: setGlobalContext } = await load('@glimmer/global-context');
setGlobalContext({ scheduleRevalidate() {} });
const { kind, op, name } = cases[Number(args.index)];
const A = [];
const B = [];
for (let i = 0; i < COUNT; i++) {
A.push(kind.a(i));
B.push(kind.b(i));
}
function create(target) {
for (let i = 0; i < COUNT; i++) target[i] = tracked(A[i]);
}
const cells = new Array(COUNT);
create(cells);
/**
* The spare cell is not in `cells`,
* so the read loop never gets the number.
*/
const spare = tracked(A[0]);
let sink = 0;
let flip = false;
function write() {
flip = !flip;
let from = flip ? B : A;
for (let i = 0; i < COUNT; i++) cells[i].value = from[i];
}
function read() {
sink += kind.read(cells);
}
let result;
if (op === 'first number') {
/**
* The cost of the event itself:
* the read loop is hot, then one cell gets the first number.
*
* `extra` is the time of the next calls above the time before the write.
*/
const CALLS = 5000;
const now = process.hrtime.bigint;
for (let r = 0; r < 20000; r++) read();
let before = now();
for (let r = 0; r < CALLS; r++) read();
let steady = Number(now() - before);
spare.value = 1;
let start = now();
read();
let first = Number(now() - start);
for (let r = 1; r < CALLS; r++) read();
let after = Number(now() - start);
result = { case: name, p50: after - steady, first, steady: steady / CALLS };
} else {
let run = read;
if (op === 'create') {
let target = new Array(COUNT);
run = () => create(target);
} else if (op === 'write') {
run = write;
} else if (op === 'read, after the first number') {
for (let r = 0; r < 20000; r++) read();
spare.value = 1;
}
let stats = await measure(run, { min_cpu_time: Number(args['min-cpu-ms']) * 1e6 });
result = { case: name, p50: stats.p50, min: stats.min };
}
if (sink === -1) console.log(sink);
if (args.out) {
writeFileSync(args.out, JSON.stringify(result));
} else {
console.log(result);
}
/**
* Compares builds of ember-source on the micro cases of `micro.mjs`.
*
* node research/tracked-value-field/micro-run.mjs \
* --sources=before=<ember-source folder>,after=<ember-source folder> \
* [--rounds=6] [--cpu=1] [--min-cpu-ms=500] [--case=<regex>] [--out=<file without extension>]
*
* - the first source is the baseline
* - each case of each build runs in its own process
* - the order of the builds rotates, and is mirrored between rounds
* - a cell is the median of the rounds
*/
import { spawnSync } from 'node:child_process';
import { readFileSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { resolve } from 'node:path';
import { fileURLToPath } from 'node:url';
import { parseArgs } from 'node:util';
const { values } = parseArgs({
options: {
sources: { type: 'string' },
rounds: { type: 'string', default: '6' },
cpu: { type: 'string' },
'min-cpu-ms': { type: 'string', default: '500' },
case: { type: 'string' },
out: { type: 'string' },
},
});
const worker = fileURLToPath(new URL('./micro.mjs', import.meta.url));
const rounds = Number(values.rounds);
const filter = values.case ? new RegExp(values.case) : null;
const sides = [];
const sources = {};
for (let pair of values.sources.split(',')) {
let [side, path] = pair.split('=');
sides.push(side);
sources[side] = resolve(path);
}
function node(extra, env) {
let args = ['--expose-gc', worker].concat(extra);
let command = process.execPath;
if (values.cpu) {
args.unshift('-c', values.cpu, command);
command = 'taskset';
}
let { status, stdout } = spawnSync(command, args, {
env: { ...process.env, ...env },
stdio: ['ignore', 'pipe', 'inherit'],
});
if (status !== 0) throw new Error(`micro.mjs ${extra.join(' ')} failed with exit code ${status}`);
return String(stdout);
}
function runOne(side, index) {
let out = `${tmpdir()}/tracked-value-micro-${process.pid}.json`;
node([`--index=${index}`, `--out=${out}`, `--min-cpu-ms=${values['min-cpu-ms']}`], {
EMBER_SOURCE: sources[side],
});
let report = JSON.parse(readFileSync(out, 'utf8'));
rmSync(out);
return report;
}
function median(list) {
let sorted = list.slice().sort((a, b) => a - b);
let middle = sorted.length >> 1;
return sorted.length % 2 ? sorted[middle] : (sorted[middle - 1] + sorted[middle]) / 2;
}
function time(ns) {
if (Math.abs(ns) >= 1e6) return `${(ns / 1e6).toFixed(2)} ms`;
if (Math.abs(ns) >= 1e4) return `${(ns / 1e3).toFixed(2)} µs`;
return `${ns.toFixed(0)} ns`;
}
/**
* The order of round `r`: the list starts at side `r / 2`,
* and each odd round runs the order of the round before it in reverse.
*/
function orderOf(round) {
let shift = (round >> 1) % sides.length;
let order = sides.slice(shift).concat(sides.slice(0, shift));
return round % 2 === 0 ? order : order.reverse();
}
const names = JSON.parse(node(['--list'], {}));
const samples = names.map(() => Object.fromEntries(sides.map((side) => [side, []])));
for (let round = 0; round < rounds; round++) {
let order = orderOf(round);
let started = Date.now();
for (let index = 0; index < names.length; index++) {
if (filter && !filter.test(names[index])) continue;
for (let side of order) {
samples[index][side].push(runOne(side, index).p50);
}
}
console.error(`round ${round + 1}/${rounds} (${Math.round((Date.now() - started) / 1000)} s)`);
}
const base = sides[0];
const lines = [];
const eventLines = [];
for (let index = 0; index < names.length; index++) {
if (filter && !filter.test(names[index])) continue;
let row = samples[index];
let baseline = median(row[base]);
let isEvent = names[index].endsWith(': first number');
let cells = [names[index].replace(': first number', '')];
for (let side of sides) {
let value = median(row[side]);
if (isEvent || side === base) {
cells.push(time(value));
continue;
}
let ratios = row[side].map((sample, round) => sample / row[base][round]).sort((x, y) => x - y);
let low = ratios[0].toFixed(2);
let high = ratios[ratios.length - 1].toFixed(2);
cells.push(`${time(value)} (${(value / baseline).toFixed(2)}x, ${low} to ${high})`);
}
(isEvent ? eventLines : lines).push(`| ${cells.join(' | ')} |`);
}
const head = `| ${sides.join(' | ')} |`;
const rule = `| --- |${sides.map(() => ' ---: |').join('')}`;
const report = [`| case, 1,000 tracked values ${head}`, rule]
.concat(lines, ['', `| extra time of 5,000 read loops after the first number ${head}`, rule])
.concat(eventLines, [
'',
`- Median of ${rounds} rounds. A time is the p50 from mitata for one pass over 1,000 tracked values.`,
`- The ratio compares with "${base}". The range is the lowest and the highest ratio of one round.`,
'- The extra time is the time of 5,000 read loops after one cell gets a number, minus the time of 5,000 read loops before it.',
`- node ${process.version}`,
])
.join('\n');
console.log(report);
if (values.out) {
writeFileSync(`${values.out}.md`, `${report}\n`);
writeFileSync(`${values.out}.json`, JSON.stringify({ names, samples, node: process.version }));
}
/**
* Measures the bytes that V8 allocates for create, read and write
* of a `tracked(value)`, for one kind of value, in this process.
*
* EMBER_SOURCE=<ember-source folder> node --expose-gc \
* --max-semi-space-size=256 --min-semi-space-size=256 alloc.mjs --index=0
* node alloc.mjs --list
*
* The young generation is large, so no GC runs inside a measurement.
* The bytes are then the growth of the used heap.
*/
import { parseArgs } from 'node:util';
import v8 from 'node:v8';
import { COUNT, kinds } from './kinds.mjs';
const { values: args } = parseArgs({
options: {
index: { type: 'string' },
list: { type: 'boolean', default: false },
},
});
if (args.list) {
console.log(JSON.stringify(kinds.map((kind) => kind.name)));
process.exit(0);
}
const { load } = await import('../../adapters/ember-source.mjs');
const { tracked } = await load('@glimmer/tracking');
const { default: setGlobalContext } = await load('@glimmer/global-context');
setGlobalContext({ scheduleRevalidate() {} });
const kind = kinds[Number(args.index)];
const A = [];
const B = [];
for (let i = 0; i < COUNT; i++) {
A.push(kind.a(i));
B.push(kind.b(i));
}
const cells = new Array(COUNT);
const target = new Array(COUNT);
function create(into) {
for (let i = 0; i < COUNT; i++) into[i] = tracked(A[i]);
}
let sink = 0;
let flip = false;
function write() {
flip = !flip;
let from = flip ? B : A;
for (let i = 0; i < COUNT; i++) cells[i].value = from[i];
}
function read() {
sink += kind.read(cells);
}
create(cells);
/**
* Newer versions of V8 count every allocated byte.
* Older versions report only the used heap, which is the same number while no GC runs.
*/
function allocated() {
let total = v8.getHeapStatistics().total_allocated_bytes;
return total === undefined ? process.memoryUsage().heapUsed : total;
}
function bytesPerOperation(run, passes) {
for (let r = 0; r < 20000; r++) run();
globalThis.gc();
let before = allocated();
for (let r = 0; r < passes; r++) run();
let after = allocated();
return (after - before) / (passes * COUNT);
}
const result = {
kind: kind.name,
create: bytesPerOperation(() => create(target), 1000),
read: bytesPerOperation(read, 5000),
write: bytesPerOperation(write, 5000),
};
if (sink === -1) console.log(sink);
console.log(JSON.stringify(result)); |
Each tracking frame made one Tracker, one Set, and one array from that Set. Frames are strictly nested, so beginTrackFrame now takes the tracker for its depth from a pool. A tracker keeps its tags in an array. A tag keeps the index at which a tracker took it last, in a new field `slot`. A tracker has the tag if its entry at that index is the tag, so a tag that the frame consumes again costs one comparison. resetTracking clears the pooled trackers. An untrack frame takes a depth but no tracker, so the pool can have holes. Split out of #21650. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The `for...of` loop made one iterator for each computation of a combined tag. The loop keeps `Math.max`. A `>` comparison ignores the NaN revision of the volatile tag. Split out of #21650. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each TrackedValue made four arrow functions and one options object. `get`, `set`, `update` and `freeze` are now accessors that make the bound function on its first read, and keep it. Each one has a setter, so code can still assign to them. The default options are one frozen object. Split out of #21650. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
All instances have one hidden class, and V8 records which kind of value the `#value` field held so far. If no instance held a number yet, the first number makes V8 throw away the optimized code of each function that reads a TrackedValue. The constructor now stores a number before the value, so the field is general from the first instance. Split out of #21650. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
endTrackFrame takes the tag that the same frame produced the last time it ran. If the frame consumed the same tags again, the result is that tag: it keeps its memoized revision, and the frame allocates no tag and no array. getValue passes the tag of its cache. Split out of emberjs#21650. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
5a5c129 to
8610111
Compare
With the check in `combine`, V8 did not inline the method into the end of the frame, and each small frame paid for a call. Against main, a frame with one tag was 4% to 7% slower, and `kairo: avoidable propagation` was 13% slower. The check is now in a function of its own, `combineOrReuse`, which only frames with two or more tags call. No case is more than 2% slower than main, and `kairo: mux` and `batch` keep their gain. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only `getValue` passes the last tag, so only a cache can reuse it. The reuse is now in code that only `getValue` calls: `endCacheFrame` and `Tracker#combineForCache`. `endTrackFrame` and `Tracker#combine` are the code of main again, with no argument. The frames of the render VM and of the curly component manager end there, so this change cannot make them slower. The two tests that passed a tag to `endTrackFrame` are gone. The tests of `createCache` cover the reuse. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tests of this change passed with the code of main too, because they only checked that a cache still follows its tags. The new tests read the tag of a cache through an outer frame. One of them fails on main: a cache that consumes the same tags again must keep its tag. Two more need a new tag: for other tags, and for the same tags in another order. Those two fail if the check says "same tags" for every combined tag. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each cache now records a step with a text that says what it reads, for example "cache reads tag1 and tag3". A test checks the steps after each read, in place of a counter that went from 1 to 5. A read that must not run the cache checks for no step. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Four of the five parts of this PR are merged: #21663, #21665, #21666 and #21667. The part that is left is #21664.
This branch is now rebased on
main. It has the same five commits as #21664, so review and merge #21664, then close this PR.TrackedValueon their first use0first in the value field of aTrackedValueTrackedValueaccessor has no setter onmain.backup/tracking-frame-allocations-before-rebasein the fork.The combined change
A tracking frame allocates less, and
tracked(value)allocates less.createCachegraph updates in about half the time: the weighted mean is 0.6x ofmain.pnpm benchis 2.1% faster in script time.packages/@glimmer/validatorchanges.mainAll 20 cases, with alien-signals for scale
The benchmark is https://github.com/NullVoxPopuli-ai-agent/ember-reactivity-bench. One measurement is the writes of one frame, then one flush that brings every output up to date. Each case runs in its own process. The numbers are the median of 4 rounds on Node 24.20.
Its
research/validator-speedfolder has the profile and the experiment that this PR comes from.What changes
beginTrackFrametakes aTrackerfrom a pool, one for each depth. Before, each frame made oneTracker, oneSet, and one array.endTrackFrame(previous)returnspreviousif the frame consumed the same tags again.getValuepasses the tag of its cache.get,set,updateandfreezeof aTrackedValueare accessors that make the bound function on its first read. Before, each instance made four functions and one options object.TrackedValuestores0before the value, so the first number in a later instance does not make V8 throw away optimized code.How the tracker finds a tag that the frame consumed before
slot.main, the tracker, itsSetandbeginTrackFrametook 27% to 45% of the samples.Why the constructor stores
0first, with a micro benchmarkAll instances of
TrackedValuehave one hidden class, and V8 records which kind of value#valueheld so far. Take a first instance that held a string, an object orundefined. The first number in any later instance then made V8 throw away the optimized code of each function that reads aTrackedValue. The store of0makes the field general in the first instance, before V8 optimizes any code.The store costs no time and no memory, and it removes a stall of 1 ms to 31 ms for one hot loop.
Memory. The bytes that V8 allocates are the same with and without the store, for 11 kinds of value, on Node 24.20 and Node 26.10.
Time. A mitata benchmark makes 1,000
tracked(value)cells for each kind: small integer, double, string, boolean,undefinedandnull, object, array, function, symbol, bigint, and all kinds mixed. Each case of each build runs in its own process, in 8 rounds. The numbers are the geometric mean of the ratios to the head before the store. "Control" is that same code from a second build, so it shows the noise.The event. A read loop is hot, then one other cell gets the first number. The time is what the next 5,000 read loops lose.
One group needs a note. After the event, the read loop is 3% to 4% faster in the builds that had the deopt, which is 0.06 ns to 0.09 ns for each read. The deopt makes V8 compile the loop a second time, and the second code is the faster one. Two more builds show that the store is not the cause. A build that stores
undefinedand then the value has the deopt, and equals the build with one store. A build with one store that makes two instances at load has no deopt, and equals the build with the store of0.The scripts, the
--trace-deoptcounts and the full tables: research page. The mitata code is also in this comment.What changes for callers
endTrackFrame()with no argument works as before.slot.get,set,updateandfreezeof aTrackedValueare still bound, each read gives the same function, and code can still assign to them. They are now accessors on the prototype, soObject.keys()and spread do not list them.Two edge cases
@glimmer/runtimeand the curly component manager do not pass a tag toendTrackFramein this PR. They get the pool and the array, but not the reuse of the tag. A test of that change showed no more gain inpnpm bench: comment.Rendering:
pnpm benchagainstmainScript time for the full run is 2.1% lower, with 50 samples for each side.
clearItems2is 13.1% slower, and the cause is the position of a GC, not the code.Phases, and the cause of the slower clear
clearItems2is 13.1% slower. A GC moves from the end of the phase before it into that phase, because the PR allocates less: comment.c5544a2699. The later commits were not run throughpnpm bench.Tests
The full suite passes locally: 9507 pass, 18 skipped, 0 failed. The type checks, ESLint and Prettier pass.
New tests and other checks
resetTracking()with an open frame and after a frame inside untrack frames.TrackedValue, with a replacedset.@glimmer/validator: trackingpasses: 39 tests.bd88a65ab1. The head54ee916923changes comments only.type-check:internals,type-check:typesandlint:docspass./** */blocks.pnpm dlx nollm@0.7.0 --diff <main>reports no problem for the added lines.Changes after the first review
main.TrackedValueaccessor has a setter, and the shared default options are frozen.resetTracking()skips the depths that have no tracker. An untrack frame takes a depth but no tracker, so a track frame that begins inside one can leave a hole in the pool.resetTracking()then calledclear()onundefined. The QUnit setup callsresetTracking()after each test, so a filtered run of the tracking tests stopped at the first test after the hole.TrackedValuestores0before the value (bd88a65ab1).🤖 Generated with Claude Code