Skip to content

Cut the allocations of a tracking frame and of a TrackedValue - #21650

Draft
NullVoxPopuli-ai-agent wants to merge 5 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/tracking-frame-allocations
Draft

NullVoxPopuli-ai-agent wants to merge 5 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/tracking-frame-allocations

Conversation

@NullVoxPopuli-ai-agent

@NullVoxPopuli-ai-agent NullVoxPopuli-ai-agent commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Four of the five parts of this PR are merged: #21663, #21665, #21666 and #21667. The part that is left is #21664.

This branch is now rebased on main. It has the same five commits as #21664, so review and merge #21664, then close this PR.

PR change state
#21663 Pool the trackers of tracking frames merged
#21664 Reuse the combined tag of a tracking frame open
#21665 Index loop over the subtags of a combined tag merged
#21666 Make the functions of a TrackedValue on their first use merged
#21667 Store 0 first in the value field of a TrackedValue merged
  • The text below describes the branch before the split. It stays as the reference for the combined numbers.
  • The merged PRs changed in review: a TrackedValue accessor has no setter on main.
  • The head before the rebase is the branch backup/tracking-frame-allocations-before-rebase in the fork.
  • The numbers for each part: table and method.

The combined change

A tracking frame allocates less, and tracked(value) allocates less.

  • A createCache graph updates in about half the time: the weighted mean is 0.6x of main.
  • Rendering in pnpm bench is 2.1% faster in script time.
  • Only packages/@glimmer/validator changes.
case main this PR
propagate: 100 chains x 100 deep 619 µs 277 µs (0.4x)
kairo: diamond 371 ns 178 ns (0.5x)
kairo: mux 18.43 µs 9.01 µs (0.5x)
rows: 1000 rows, write 1 6.88 µs 6.22 µs (0.9x)
rows: 1000 rows, write all 103.8 µs 51.2 µs (0.5x)
batch: 10 writes, 1 output 580 ns 248 ns (0.4x)
create: 1000 signals 16.94 µs 9.96 µs (0.6x)
create: 1000 computeds, read each 38.2 µs 21.6 µs (0.6x)
All 20 cases, with alien-signals for scale
case ember: main ember: this PR alien-signals
propagate: 1 chains x 1 deep 100 ns 56 ns (0.6x) 54 ns (0.5x)
propagate: 10 chains x 10 deep 5.44 µs 2.43 µs (0.4x) 2.78 µs (0.5x)
propagate: 100 chains x 100 deep 619.49 µs 276.72 µs (0.4x) 505.74 µs (0.8x)
propagate: 1 chains x 1000 deep 51.78 µs 22.75 µs (0.4x) 24.93 µs (0.5x)
propagate: 1000 chains x 1 deep 87.07 µs 43.77 µs (0.5x) 50.89 µs (0.6x)
kairo: avoidable propagation 326 ns 146 ns (0.4x) 74 ns (0.2x)
kairo: broad propagation 6.63 µs 3.22 µs (0.5x) 3.67 µs (0.6x)
kairo: deep propagation 2.52 µs 1.14 µs (0.5x) 1.13 µs (0.4x)
kairo: diamond 371 ns 178 ns (0.5x) 164 ns (0.4x)
kairo: mux 18.43 µs 9.01 µs (0.5x) 5.78 µs (0.3x)
kairo: repeated observers 345 ns 223 ns (0.6x) 99 ns (0.3x)
kairo: triangle 649 ns 316 ns (0.5x) 299 ns (0.5x)
kairo: unstable 506 ns 310 ns (0.6x) 295 ns (0.6x)
rows: 1000 rows, write 1 6.88 µs 6.22 µs (0.9x) 5.94 µs (0.9x)
rows: 1000 rows, write all 103.75 µs 51.16 µs (0.5x) 60.02 µs (0.6x)
batch: 10 writes, 1 output 580 ns 248 ns (0.4x) 228 ns (0.4x)
avoidable: write the same value 5 ns 6 ns (1.1x) 9 ns (1.7x)
create: 1000 signals 16.94 µs 9.96 µs (0.6x) 3.37 µs (0.2x)
create: 1000 computeds, read each 38.20 µs 21.59 µs (0.6x) 24.80 µs (0.6x)
create: 1000 outputs 44.18 µs 28.59 µs (0.6x) 30.71 µs (0.7x)
weighted geometric mean 1.0x 0.6x 0.6x

The benchmark is https://github.com/NullVoxPopuli-ai-agent/ember-reactivity-bench. One measurement is the writes of one frame, then one flush that brings every output up to date. Each case runs in its own process. The numbers are the median of 4 rounds on Node 24.20.

Its research/validator-speed folder has the profile and the experiment that this PR comes from.

What changes

  1. beginTrackFrame takes a Tracker from a pool, one for each depth. Before, each frame made one Tracker, one Set, and one array.
  2. endTrackFrame(previous) returns previous if the frame consumed the same tags again. getValue passes the tag of its cache.
  3. get, set, update and freeze of a TrackedValue are accessors that make the bound function on its first read. Before, each instance made four functions and one options object.
  4. The constructor of a TrackedValue stores 0 before the value, so the first number in a later instance does not make V8 throw away optimized code.
How the tracker finds a tag that the frame consumed before
  • A tracker keeps its tags in an array. A tag keeps the index at which a tracker took it last, in a new field slot.
  • A tracker has the tag if its entry at that index is the tag. A tag that the frame consumes again costs one comparison.
  • No counter is necessary, and a stale index cannot hide a tag.
  • A frame counter was the first design. It leaves the small-integer range of V8 after about one billion frames, and the benchmark was then 1.1 to 1.3 times slower.
  • The loop over the subtags of a combined tag uses an index in place of an iterator.
  • In CPU profiles of three graphs on main, the tracker, its Set and beginTrackFrame took 27% to 45% of the samples.
Why the constructor stores 0 first, with a micro benchmark

All instances of TrackedValue have one hidden class, and V8 records which kind of value #value held so far. Take a first instance that held a string, an object or undefined. The first number in any later instance then made V8 throw away the optimized code of each function that reads a TrackedValue. The store of 0 makes the field general in the first instance, before V8 optimizes any code.

The store costs no time and no memory, and it removes a stall of 1 ms to 31 ms for one hot loop.

Memory. The bytes that V8 allocates are the same with and without the store, for 11 kinds of value, on Node 24.20 and Node 26.10.

operation Node 24.20 Node 26.10
create, most kinds 184 bytes 240 bytes
create, double 200 bytes 256 bytes
read 0 bytes 0 bytes
write, double 16 bytes 16 bytes
write, other kinds 0 bytes 0 bytes

Time. A mitata benchmark makes 1,000 tracked(value) cells for each kind: small integer, double, string, boolean, undefined and null, object, array, function, symbol, bigint, and all kinds mixed. Each case of each build runs in its own process, in 8 rounds. The numbers are the geometric mean of the ratios to the head before the store. "Control" is that same code from a second build, so it shows the noise.

group Node 24.20, control Node 24.20, with the store Node 26.10, control Node 26.10, with the store
create, 11 kinds 0.993 0.988 0.997 0.991
read, 11 kinds 0.988 0.993 0.984 1.005
write, 11 kinds 0.999 0.994 0.989 0.996

The event. A read loop is hot, then one other cell gets the first number. The time is what the next 5,000 read loops lose.

Node 24.20 Node 26.10
without the store, 8 kinds 14 ms to 31 ms 1.1 ms to 16 ms
with the store, 8 kinds -1.3 ms to 1.0 ms -1.1 ms to 1.2 ms

One group needs a note. After the event, the read loop is 3% to 4% faster in the builds that had the deopt, which is 0.06 ns to 0.09 ns for each read. The deopt makes V8 compile the loop a second time, and the second code is the faster one. Two more builds show that the store is not the cause. A build that stores undefined and then the value has the deopt, and equals the build with one store. A build with one store that makes two instances at load has no deopt, and equals the build with the store of 0.

The scripts, the --trace-deopt counts and the full tables: research page. The mitata code is also in this comment.

What changes for callers

  • endTrackFrame() with no argument works as before.
  • Each tag has one more field, slot.
  • get, set, update and freeze of a TrackedValue are still bound, each read gives the same function, and code can still assign to them. They are now accessors on the prototype, so Object.keys() and spread do not list them.
Two edge cases
  • A tag that a nested frame takes at another index, between two consumptions of the outer frame, is in the outer frame two times. The combined tag has the same revision. If the nested frame takes the tag at the same index, the outer frame has the tag one time.
  • @glimmer/runtime and the curly component manager do not pass a tag to endTrackFrame in this PR. They get the pool and the array, but not the reuse of the tag. A test of that change showed no more gain in pnpm bench: comment.

Rendering: pnpm bench against main

Script time for the full run is 2.1% lower, with 50 samples for each side. clearItems2 is 13.1% slower, and the cause is the position of a GC, not the code.

Phases, and the cause of the slower clear
  • Select is 9 to 10% faster, and the second update of each 10th row is 13.7% faster.
  • Three render phases are 3.5 to 3.8% faster. The second append is 4.4% faster.
  • clearItems2 is 13.1% slower. A GC moves from the end of the phase before it into that phase, because the PR allocates less: comment.
  • The table, the tracerbench PDF and the method are in this comment.
  • The first version of this PR was 3.3% faster, and the two intervals overlap.
  • This run is from head c5544a2699. The later commits were not run through pnpm bench.

Tests

The full suite passes locally: 9507 pass, 18 skipped, 0 failed. The type checks, ESLint and Prettier pass.

New tests and other checks
  • New tests for the reuse of the tag, and for a tag that a nested frame consumes.
  • A new test for an outer cache that follows the reused tag of an inner cache.
  • New tests for resetTracking() with an open frame and after a frame inside untrack frames.
  • New tests for the functions of a TrackedValue, with a replaced set.
  • A filtered run of @glimmer/validator: tracking passes: 39 tests.
  • The full suite ran on bd88a65ab1. The head 54ee916923 changes comments only.
  • type-check:internals, type-check:types and lint:docs pass.
  • The comments that this PR adds are /** */ blocks. pnpm dlx nollm@0.7.0 --diff <main> reports no problem for the added lines.
Changes after the first review
  • The tracker finds a repeat tag by its index, not by a frame number. The frame number also made a computed that reads five computeds of one tag keep a combined tag with five entries. It has one tag again, as on main.
  • Each TrackedValue accessor has a setter, and the shared default options are frozen.
  • resetTracking() skips the depths that have no tracker. An untrack frame takes a depth but no tracker, so a track frame that begins inside one can leave a hole in the pool. resetTracking() then called clear() on undefined. The QUnit setup calls resetTracking() after each test, so a filtered run of the tracking tests stopped at the first test after the hole.
  • The constructor of a TrackedValue stores 0 before the value (bd88a65ab1).

🤖 Generated with Claude Code

@NullVoxPopuli-ai-agent NullVoxPopuli-ai-agent changed the title [PERF] Cut the allocations of a tracking frame and of a TrackedValue Cut the allocations of a tracking frame and of a TrackedValue Oct 2, 2026
@NullVoxPopuli-ai-agent
NullVoxPopuli-ai-agent force-pushed the nvp/tracking-frame-allocations branch from aae7903 to c59a238 Compare October 2, 2026 15:11
@NullVoxPopuli
NullVoxPopuli marked this pull request as draft October 2, 2026 15:32
@NullVoxPopuli

Copy link
Copy Markdown
Contributor

moved to draft while I explore if the specifics of this PR are a good idea

@NullVoxPopuli-ai-agent

This comment was marked as outdated.

@NullVoxPopuli-ai-agent

This comment was marked as outdated.

@NullVoxPopuli-ai-agent

This comment was marked as outdated.

@NullVoxPopuli-ai-agent

This comment was marked as outdated.

NullVoxPopuli-ai-agent pushed a commit to NullVoxPopuli-ai-agent/ember.js that referenced this pull request Oct 5, 2026
…it later

Blocks and component cache groups pass their previous tag to
endTrackFrame(), so a block that re-renders with the same dependencies
keeps its tag and its memoized revision. The tag comparison against the
tracker's live entries comes from emberjs#21650, which is below this commit.

A guard that fails is dropped on the spot, with no tracking frame: a block
that changed is likely to change again, and a frame on every miss is what
made all-rows-change workloads slower. A dropped guard comes back after
eight unguarded updates, so a block that changed once and then stayed
still is skipped again, while a block that changes on every update pays
for one frame in every eight. A re-render re-arms the guard right away.

An unguarded block costs what it did before guards existed: one frame
push. A re-render of a dropped block opens its own tracking frame, because
the append VM closes one when the block exits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wPae8QnFbKHKn5xbgaQVq
NullVoxPopuli-ai-agent pushed a commit to NullVoxPopuli-ai-agent/ember.js that referenced this pull request Oct 6, 2026
emberjs#21650 now has the commit that makes the value field of a TrackedValue
general (bd88a65). This branch had the same change as f04dc41.
The merge changes no file content.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

# Conflicts:
#	packages/@glimmer/validator/lib/tracked-value.ts
@NullVoxPopuli-ai-agent

Copy link
Copy Markdown
Contributor Author

The store of 0 on line 56 of tracked-value.ts costs no time and no memory. It removes a stall of 1 ms to 31 ms for one hot loop, one time for each page load.

"Before" is 50eef0f8c1. "After" is the head 54ee916923, with the store. Node 24.20 and Node 26.10 give the same answer.

Memory: no change. Bytes that V8 allocates for one operation on one tracked value, for 11 kinds of value:

operation Node 24.20, before and after Node 26.10, before and after
create, most kinds 184 bytes 240 bytes
create, double 200 bytes 256 bytes
read 0 bytes 0 bytes
write, double 16 bytes 16 bytes
write, other kinds 0 bytes 0 bytes

Time: no change. Geometric mean of after / before for 1,000 tracked values of each kind, 8 rounds. "Control" is the code of "before" from a second build, so it shows the noise.

group Node 24.20, control Node 24.20, after Node 26.10, control Node 26.10, after
create, 11 kinds 0.993 0.988 0.997 0.991
read, 11 kinds 0.988 0.993 0.984 1.005
write, 11 kinds 0.999 0.994 0.989 0.996

The event that the store removes. A read loop is hot, then one other cell gets the first number. The time is what the next 5,000 read loops lose.

Node 24.20 Node 26.10
before, 8 kinds 14 ms to 31 ms 1.1 ms to 16 ms
after, 8 kinds -1.3 ms to 1.0 ms -1.1 ms to 1.2 ms
One group where "after" is slower, and why the store is not the cause

After the event, the read loop is 3% to 4% slower in "after". That is 0.06 ns to 0.09 ns for each read.

Two more builds separate the store from the deopt:

build constructor deopt read after the event, Node 24.20 Node 26.10
control one store yes 1.002 0.990
undefined-store this.#value = undefined, then the value yes 1.000 0.993
after this.#value = 0, then the value no 1.038 1.040
prime one store. The module makes two instances at load, with 0 and with undefined no 1.044 1.031
  • A second store is free: "undefined-store" equals "before".
  • "Prime" has no second store and is as slow as "after".
  • The two slower builds are the two with no deopt. --trace-opt shows that the deopt makes V8 compile the read loop a second time, and the second code is the faster one.
  • The field has the same representation in all builds at that point. So the difference comes from which compile of the loop the benchmark measures.
The kinds of value, and the method
  • The kinds: small integer, double, string, boolean, undefined and null, object, array, function, symbol, bigint, and all kinds mixed.
  • Each case of each build runs in its own process, pinned to one core. The #value field then sees only the values of one case.
  • The order of the builds rotates and is mirrored between rounds. A cell is the median of the rounds.
  • The read loop does work that fits the kind, for example a sum for numbers and a sum of length for strings.
  • Memory on Node 26.10 is total_allocated_bytes of V8. On Node 24.20 it is the growth of the used heap, with a young generation large enough that no GC runs.
  • A write of a double allocates 16 bytes in all builds. The class field starts as undefined, so V8 puts each double in a new heap number.
The mitata code

The files are in the research folder, with the full tables for each case.

node research/tracked-value-field/micro-run.mjs --cpu=1 --rounds=8 \
  --sources=before=<ember-source>,after=<ember-source>

kinds.mjs

/**
 * The kinds of value that the micro cases store in a tracked value.
 */
export const COUNT = 1000;

/**
 * Each read loop does the work that fits the kind of value,
 * so that V8 has a reason to use what it knows about the field.
 */
function sumNumber(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) n += cells[i].value;
  return n;
}

function sumLength(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) n += cells[i].value.length;
  return n;
}

function sumField(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) n += cells[i].value.i;
  return n;
}

function countTrue(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) if (cells[i].value === true) n++;
  return n;
}

function countUndefined(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) if (cells[i].value === undefined) n++;
  return n;
}

function countFunction(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) if (typeof cells[i].value === 'function') n++;
  return n;
}

const PROBE = Symbol('probe');
function countProbe(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) if (cells[i].value === PROBE) n++;
  return n;
}

function countBig(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) if (cells[i].value > 500n) n++;
  return n;
}

function countNumber(cells) {
  let n = 0;
  for (let i = 0; i < cells.length; i++) if (typeof cells[i].value === 'number') n++;
  return n;
}

/**
 * `a` is the first value of cell `i`. A write pass stores `b`, then `a` again.
 * `number` says that the kind stores a number in the field without help.
 */
export const kinds = [
  { name: 'small integer', number: true, a: (i) => i, b: (i) => i + 1, read: sumNumber },
  { name: 'double', number: true, a: (i) => i + 0.5, b: (i) => i + 1.5, read: sumNumber },
  { name: 'string', a: (i) => `a${i}`, b: (i) => `b${i}`, read: sumLength },
  { name: 'boolean', a: (i) => i % 2 === 0, b: (i) => i % 2 === 1, read: countTrue },
  {
    name: 'undefined and null',
    a: (i) => (i % 2 === 0 ? undefined : null),
    b: (i) => (i % 2 === 0 ? null : undefined),
    read: countUndefined,
  },
  { name: 'object', a: (i) => ({ i }), b: (i) => ({ i: i + 1 }), read: sumField },
  { name: 'array', a: (i) => [i], b: (i) => [i, i], read: sumLength },
  { name: 'function', a: (i) => () => i, b: (i) => () => i + 1, read: countFunction },
  {
    name: 'symbol',
    a: (i) => (i % 2 === 0 ? PROBE : Symbol()),
    b: (i) => (i % 2 === 1 ? PROBE : Symbol()),
    read: countProbe,
  },
  { name: 'bigint', a: (i) => BigInt(i), b: (i) => BigInt(i + 1), read: countBig },
];

const mixedKinds = kinds.slice();
kinds.push({
  name: 'all kinds mixed',
  number: true,
  a: (i) => mixedKinds[i % mixedKinds.length].a(i),
  b: (i) => mixedKinds[(i + 1) % mixedKinds.length].b(i),
  read: countNumber,
});

micro.mjs

/**
 * Measures one micro case of `tracked(value)` in this process.
 *
 * `micro-run.mjs` starts this file one time per build, per case, per round,
 * so the `#value` field of the class sees only the values of one case.
 *
 *   EMBER_SOURCE=<ember-source folder> node --expose-gc micro.mjs --index=0
 *   node micro.mjs --list
 */
import { writeFileSync } from 'node:fs';
import { parseArgs } from 'node:util';

import { measure } from 'mitata';

import { COUNT, kinds } from './kinds.mjs';

const { values: args } = parseArgs({
  options: {
    index: { type: 'string' },
    list: { type: 'boolean', default: false },
    out: { type: 'string' },
    'min-cpu-ms': { type: 'string', default: '500' },
  },
});

const cases = [];
for (let kind of kinds) {
  cases.push({ kind, op: 'create' }, { kind, op: 'read' }, { kind, op: 'write' });
}
for (let kind of kinds) {
  if (kind.number) continue;
  cases.push({ kind, op: 'read, after the first number' }, { kind, op: 'first number' });
}
for (let c of cases) c.name = `${c.kind.name}: ${c.op}`;

if (args.list) {
  console.log(JSON.stringify(cases.map((c) => c.name)));
  process.exit(0);
}

const { load } = await import('../../adapters/ember-source.mjs');
const { tracked } = await load('@glimmer/tracking');
const { default: setGlobalContext } = await load('@glimmer/global-context');

setGlobalContext({ scheduleRevalidate() {} });

const { kind, op, name } = cases[Number(args.index)];

const A = [];
const B = [];
for (let i = 0; i < COUNT; i++) {
  A.push(kind.a(i));
  B.push(kind.b(i));
}

function create(target) {
  for (let i = 0; i < COUNT; i++) target[i] = tracked(A[i]);
}

const cells = new Array(COUNT);
create(cells);

/**
 * The spare cell is not in `cells`,
 * so the read loop never gets the number.
 */
const spare = tracked(A[0]);

let sink = 0;
let flip = false;

function write() {
  flip = !flip;
  let from = flip ? B : A;
  for (let i = 0; i < COUNT; i++) cells[i].value = from[i];
}

function read() {
  sink += kind.read(cells);
}

let result;

if (op === 'first number') {
  /**
   * The cost of the event itself:
   * the read loop is hot, then one cell gets the first number.
   *
   * `extra` is the time of the next calls above the time before the write.
   */
  const CALLS = 5000;
  const now = process.hrtime.bigint;

  for (let r = 0; r < 20000; r++) read();

  let before = now();
  for (let r = 0; r < CALLS; r++) read();
  let steady = Number(now() - before);

  spare.value = 1;

  let start = now();
  read();
  let first = Number(now() - start);
  for (let r = 1; r < CALLS; r++) read();
  let after = Number(now() - start);

  result = { case: name, p50: after - steady, first, steady: steady / CALLS };
} else {
  let run = read;

  if (op === 'create') {
    let target = new Array(COUNT);
    run = () => create(target);
  } else if (op === 'write') {
    run = write;
  } else if (op === 'read, after the first number') {
    for (let r = 0; r < 20000; r++) read();
    spare.value = 1;
  }

  let stats = await measure(run, { min_cpu_time: Number(args['min-cpu-ms']) * 1e6 });

  result = { case: name, p50: stats.p50, min: stats.min };
}

if (sink === -1) console.log(sink);

if (args.out) {
  writeFileSync(args.out, JSON.stringify(result));
} else {
  console.log(result);
}

micro-run.mjs

/**
 * Compares builds of ember-source on the micro cases of `micro.mjs`.
 *
 *   node research/tracked-value-field/micro-run.mjs \
 *     --sources=before=<ember-source folder>,after=<ember-source folder> \
 *     [--rounds=6] [--cpu=1] [--min-cpu-ms=500] [--case=<regex>] [--out=<file without extension>]
 *
 * - the first source is the baseline
 * - each case of each build runs in its own process
 * - the order of the builds rotates, and is mirrored between rounds
 * - a cell is the median of the rounds
 */
import { spawnSync } from 'node:child_process';
import { readFileSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { resolve } from 'node:path';
import { fileURLToPath } from 'node:url';
import { parseArgs } from 'node:util';

const { values } = parseArgs({
  options: {
    sources: { type: 'string' },
    rounds: { type: 'string', default: '6' },
    cpu: { type: 'string' },
    'min-cpu-ms': { type: 'string', default: '500' },
    case: { type: 'string' },
    out: { type: 'string' },
  },
});

const worker = fileURLToPath(new URL('./micro.mjs', import.meta.url));
const rounds = Number(values.rounds);
const filter = values.case ? new RegExp(values.case) : null;

const sides = [];
const sources = {};
for (let pair of values.sources.split(',')) {
  let [side, path] = pair.split('=');
  sides.push(side);
  sources[side] = resolve(path);
}

function node(extra, env) {
  let args = ['--expose-gc', worker].concat(extra);
  let command = process.execPath;

  if (values.cpu) {
    args.unshift('-c', values.cpu, command);
    command = 'taskset';
  }

  let { status, stdout } = spawnSync(command, args, {
    env: { ...process.env, ...env },
    stdio: ['ignore', 'pipe', 'inherit'],
  });

  if (status !== 0) throw new Error(`micro.mjs ${extra.join(' ')} failed with exit code ${status}`);

  return String(stdout);
}

function runOne(side, index) {
  let out = `${tmpdir()}/tracked-value-micro-${process.pid}.json`;

  node([`--index=${index}`, `--out=${out}`, `--min-cpu-ms=${values['min-cpu-ms']}`], {
    EMBER_SOURCE: sources[side],
  });

  let report = JSON.parse(readFileSync(out, 'utf8'));

  rmSync(out);

  return report;
}

function median(list) {
  let sorted = list.slice().sort((a, b) => a - b);
  let middle = sorted.length >> 1;

  return sorted.length % 2 ? sorted[middle] : (sorted[middle - 1] + sorted[middle]) / 2;
}

function time(ns) {
  if (Math.abs(ns) >= 1e6) return `${(ns / 1e6).toFixed(2)} ms`;
  if (Math.abs(ns) >= 1e4) return `${(ns / 1e3).toFixed(2)} µs`;

  return `${ns.toFixed(0)} ns`;
}

/**
 * The order of round `r`: the list starts at side `r / 2`,
 * and each odd round runs the order of the round before it in reverse.
 */
function orderOf(round) {
  let shift = (round >> 1) % sides.length;
  let order = sides.slice(shift).concat(sides.slice(0, shift));

  return round % 2 === 0 ? order : order.reverse();
}

const names = JSON.parse(node(['--list'], {}));
const samples = names.map(() => Object.fromEntries(sides.map((side) => [side, []])));

for (let round = 0; round < rounds; round++) {
  let order = orderOf(round);
  let started = Date.now();

  for (let index = 0; index < names.length; index++) {
    if (filter && !filter.test(names[index])) continue;

    for (let side of order) {
      samples[index][side].push(runOne(side, index).p50);
    }
  }

  console.error(`round ${round + 1}/${rounds} (${Math.round((Date.now() - started) / 1000)} s)`);
}

const base = sides[0];
const lines = [];
const eventLines = [];

for (let index = 0; index < names.length; index++) {
  if (filter && !filter.test(names[index])) continue;

  let row = samples[index];
  let baseline = median(row[base]);
  let isEvent = names[index].endsWith(': first number');
  let cells = [names[index].replace(': first number', '')];

  for (let side of sides) {
    let value = median(row[side]);

    if (isEvent || side === base) {
      cells.push(time(value));
      continue;
    }

    let ratios = row[side].map((sample, round) => sample / row[base][round]).sort((x, y) => x - y);
    let low = ratios[0].toFixed(2);
    let high = ratios[ratios.length - 1].toFixed(2);

    cells.push(`${time(value)} (${(value / baseline).toFixed(2)}x, ${low} to ${high})`);
  }

  (isEvent ? eventLines : lines).push(`| ${cells.join(' | ')} |`);
}

const head = `| ${sides.join(' | ')} |`;
const rule = `| --- |${sides.map(() => ' ---: |').join('')}`;

const report = [`| case, 1,000 tracked values ${head}`, rule]
  .concat(lines, ['', `| extra time of 5,000 read loops after the first number ${head}`, rule])
  .concat(eventLines, [
    '',
    `- Median of ${rounds} rounds. A time is the p50 from mitata for one pass over 1,000 tracked values.`,
    `- The ratio compares with "${base}". The range is the lowest and the highest ratio of one round.`,
    '- The extra time is the time of 5,000 read loops after one cell gets a number, minus the time of 5,000 read loops before it.',
    `- node ${process.version}`,
  ])
  .join('\n');

console.log(report);

if (values.out) {
  writeFileSync(`${values.out}.md`, `${report}\n`);
  writeFileSync(`${values.out}.json`, JSON.stringify({ names, samples, node: process.version }));
}

alloc.mjs

/**
 * Measures the bytes that V8 allocates for create, read and write
 * of a `tracked(value)`, for one kind of value, in this process.
 *
 *   EMBER_SOURCE=<ember-source folder> node --expose-gc \
 *     --max-semi-space-size=256 --min-semi-space-size=256 alloc.mjs --index=0
 *   node alloc.mjs --list
 *
 * The young generation is large, so no GC runs inside a measurement.
 * The bytes are then the growth of the used heap.
 */
import { parseArgs } from 'node:util';
import v8 from 'node:v8';

import { COUNT, kinds } from './kinds.mjs';

const { values: args } = parseArgs({
  options: {
    index: { type: 'string' },
    list: { type: 'boolean', default: false },
  },
});

if (args.list) {
  console.log(JSON.stringify(kinds.map((kind) => kind.name)));
  process.exit(0);
}

const { load } = await import('../../adapters/ember-source.mjs');
const { tracked } = await load('@glimmer/tracking');
const { default: setGlobalContext } = await load('@glimmer/global-context');

setGlobalContext({ scheduleRevalidate() {} });

const kind = kinds[Number(args.index)];

const A = [];
const B = [];
for (let i = 0; i < COUNT; i++) {
  A.push(kind.a(i));
  B.push(kind.b(i));
}

const cells = new Array(COUNT);
const target = new Array(COUNT);

function create(into) {
  for (let i = 0; i < COUNT; i++) into[i] = tracked(A[i]);
}

let sink = 0;
let flip = false;

function write() {
  flip = !flip;
  let from = flip ? B : A;
  for (let i = 0; i < COUNT; i++) cells[i].value = from[i];
}

function read() {
  sink += kind.read(cells);
}

create(cells);

/**
 * Newer versions of V8 count every allocated byte.
 * Older versions report only the used heap, which is the same number while no GC runs.
 */
function allocated() {
  let total = v8.getHeapStatistics().total_allocated_bytes;

  return total === undefined ? process.memoryUsage().heapUsed : total;
}

function bytesPerOperation(run, passes) {
  for (let r = 0; r < 20000; r++) run();

  globalThis.gc();

  let before = allocated();
  for (let r = 0; r < passes; r++) run();
  let after = allocated();

  return (after - before) / (passes * COUNT);
}

const result = {
  kind: kind.name,
  create: bytesPerOperation(() => create(target), 1000),
  read: bytesPerOperation(read, 5000),
  write: bytesPerOperation(write, 5000),
};

if (sink === -1) console.log(sink);

console.log(JSON.stringify(result));

NullVoxPopuli added a commit that referenced this pull request Oct 6, 2026
Each tracking frame made one Tracker, one Set, and one array from that
Set. Frames are strictly nested, so beginTrackFrame now takes the
tracker for its depth from a pool.

A tracker keeps its tags in an array. A tag keeps the index at which a
tracker took it last, in a new field `slot`. A tracker has the tag if
its entry at that index is the tag, so a tag that the frame consumes
again costs one comparison.

resetTracking clears the pooled trackers. An untrack frame takes a depth
but no tracker, so the pool can have holes.

Split out of #21650.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
NullVoxPopuli added a commit that referenced this pull request Oct 6, 2026
The `for...of` loop made one iterator for each computation of a
combined tag.

The loop keeps `Math.max`. A `>` comparison ignores the NaN revision of
the volatile tag.

Split out of #21650.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
NullVoxPopuli added a commit that referenced this pull request Oct 7, 2026
Each TrackedValue made four arrow functions and one options object.
`get`, `set`, `update` and `freeze` are now accessors that make the
bound function on its first read, and keep it. Each one has a setter,
so code can still assign to them.

The default options are one frozen object.

Split out of #21650.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
NullVoxPopuli added a commit that referenced this pull request Oct 7, 2026
All instances have one hidden class, and V8 records which kind of value
the `#value` field held so far. If no instance held a number yet, the
first number makes V8 throw away the optimized code of each function
that reads a TrackedValue.

The constructor now stores a number before the value, so the field is
general from the first instance.

Split out of #21650.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
endTrackFrame takes the tag that the same frame produced the last time
it ran. If the frame consumed the same tags again, the result is that
tag: it keeps its memoized revision, and the frame allocates no tag and
no array.

getValue passes the tag of its cache.

Split out of emberjs#21650.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@NullVoxPopuli-ai-agent
NullVoxPopuli-ai-agent force-pushed the nvp/tracking-frame-allocations branch from 5a5c129 to 8610111 Compare October 7, 2026 17:18
NullVoxPopuli and others added 3 commits October 7, 2026 14:40
With the check in `combine`, V8 did not inline the method into the end
of the frame, and each small frame paid for a call. Against main, a
frame with one tag was 4% to 7% slower, and `kairo: avoidable
propagation` was 13% slower.

The check is now in a function of its own, `combineOrReuse`, which only
frames with two or more tags call. No case is more than 2% slower than
main, and `kairo: mux` and `batch` keep their gain.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only `getValue` passes the last tag, so only a cache can reuse it.
The reuse is now in code that only `getValue` calls:
`endCacheFrame` and `Tracker#combineForCache`.

`endTrackFrame` and `Tracker#combine` are the code of main again, with
no argument. The frames of the render VM and of the curly component
manager end there, so this change cannot make them slower.

The two tests that passed a tag to `endTrackFrame` are gone. The tests
of `createCache` cover the reuse.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tests of this change passed with the code of main too, because they
only checked that a cache still follows its tags.

The new tests read the tag of a cache through an outer frame. One of
them fails on main: a cache that consumes the same tags again must keep
its tag. Two more need a new tag: for other tags, and for the same tags
in another order. Those two fail if the check says "same tags" for
every combined tag.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each cache now records a step with a text that says what it reads, for
example "cache reads tag1 and tag3". A test checks the steps after each
read, in place of a counter that went from 1 to 5.

A read that must not run the cache checks for no step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants