Skip to content

Keep the value and the tag of a tracked field in one cell - #21661

Draft
NullVoxPopuli-ai-agent wants to merge 3 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/tracked-cell
Draft

NullVoxPopuli-ai-agent wants to merge 3 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/tracked-cell

Conversation

@NullVoxPopuli-ai-agent

@NullVoxPopuli-ai-agent NullVoxPopuli-ai-agent commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

A read or a write of a @tracked field now does one WeakMap lookup. On main it does three map lookups.

  • rere-benchmark ember apps, against the current main: 6.2% faster (geometric mean of 13 benches). 1 item, 100k updates is 14.9% faster.
  • pnpm bench, against the current main: no significant change for the full run (+0.5%).
  • Costs in pnpm bench: swapRows1 is 6.9% slower, removeSecondRow1 is 4.6% slower, and swapRows2 is 3.5% slower. Each of these phases takes about 50 ms.

Extracted from the spike #21656. This PR does not depend on another PR.

How it works

On main, the getter of a tracked field calls tagFor(self, key) and then reads the value:

  1. TRACKED_TAGS.get(self), a WeakMap lookup
  2. tags.get(key), a Map lookup
  3. values.get(self), a second WeakMap lookup

The setter does the same through dirtyTagFor.

With this PR, each tracked field has one WeakMap from the instance to a cell:

interface TrackedCell<V> {
  value: V;
  tag: UpdatableTag;
  initialized: boolean;
}
  • A read is one lookup, then consumeTag(cell.tag).
  • A write is one lookup, then DIRTY_TAG(cell.tag, true).
  • The cell takes its tag from the tag registry when it is made. So tagFor and dirtyTagFor work on the tag of the field, as on main. Observers, computed chains and notifyPropertyChange get the tag from the registry, and they can ask for it before the first read or write of the field.

Commits

  1. 4427f536f3 Keep a tracked field's value and tag in one cell
  2. fcfd26b0ea Make the value field of a tracked cell general from the first cell
  3. e2e0b1b558 merges main. The one conflict was an import line in tracking-test.ts.

Against the current main

Control is main at 9bec1cb2a8, which has the merged parts of #21650. Experiment is this branch with main merged in, e2e0b1b558.

rere-benchmark, with main, this PR, and #21662 in 4 mirrored cycles: 8 runs for each build, 0 void runs, geometric mean of 13 benches.

pair change cycle 1 cycle 2 cycle 3 cycle 4
cell (#21661) against main -6.2% -0.0% -8.0% -6.4% -11.0%
TrackedValue (#21662) against main -3.7% +2.7% -7.9% -0.1% -5.2%
TrackedValue against the cell +2.7% +2.7% +0.1% +6.7% +6.5%
  • The main build did not drift in this round: -0.3% between the first two cycles and the last two.
  • 1 item, 100k updates is 14.9% faster with the cell, and 11.3% faster with the TrackedValue.
  • The cell is ahead of the TrackedValue in each cycle, by 0.1% to 6.7%.

pnpm bench, 50 paired rounds: tracerbench-report.pdf

  • Script time for the full run without GC: +0.5%, which is not significant.
  • 4 of 23 phases have a significant change in script time: clearManyItems2 is 1.0% faster, swapRows1 is 6.9% slower, removeSecondRow1 is 4.6% slower, and swapRows2 is 3.5% slower.
  • The slower clear phases of the run before the merge do not show here. I did not look for the cause of the slower swap and remove phases. Store each @tracked field in a TrackedValue #21662 also shows removeSecondRow1 4.2% slower against the same control.

The two sections below are from before the merge, with main at f693f240ee as the control.

rere-benchmark

main, the first commit and the second commit ran in 4 mirrored cycles: 8 runs for each build, 5 samples per bench in each run. The table is for the second commit. Each number is the median of 40 samples, in ms.

Bench main this PR change
1 item, 1k updates 5.7 4.7 -17.5%
1 item, 100k updates 23.4 19.4 -17.1%
1k items, 1 update each (sequentially) 27.2 25.0 -7.7%
1k items 1 update on 5% (random, async) 57.6 53.8 -6.5%
1k items 1 update on 5% (random) 14.7 13.8 -6.1%
1 value, 1k consumers, 10k updates (bursts of 1000) 108 103 -5.4%
1 value, 1k consumers, 10k updates (single burst) 31.0 29.7 -4.3%
Incrementing Render Effect 1549 1497 -3.4%
1k items, 1 update each (sequentially, async) 610 593 -2.7%
1 item, 1k updates (async) 35.1 34.3 -2.1%
1 item, 100k updates (async) 669 658 -1.6%
1k items 1 update on 25% (random) 17.8 18.0 +1.1%
1k items 1 update on 25% (random, async) 181 184 +1.4%
DB Monitor w/ chat simulation (fps) 113 112 -1.0% fps
  • Over 13 benches, the second commit is 5.7% faster than main (geometric mean). The four cycles read -3.5%, -7.0%, -3.1%, and -6.7%.
  • The first commit reads -6.0% in the same round. The two commits do not differ: +0.2%, with cycles on both sides of 0.
  • An earlier round, on a machine with more other load, gave -2.0% for the first commit, with one cycle of four above 0.
  • 1 item, 100k updates is faster in every round: -12.2%, -12.7%, and -17.1%.
  • The list benches with a synchronous render take 14 to 27 ms, and their results change sign between rounds. I do not count one row of them as a change.
  • The geometric mean leaves out DB Monitor and fan-out bursts of 100. Both depend on frame timing.

pnpm bench

Tracerbench compare of the second commit against main, 50 rounds: tracerbench-report.pdf.

phase main script ms script time full phase
all phases 5296 no change no change
clearItems2 41 +4.8% slower +1.0% slower
clearManyItems1 403 +1.4% slower +1.2% slower

The table shows the full run and each phase with a significant change. The full run reads +0.4%, which is not significant.

The benchmark app has one tracked field, selected on the state service. Each row reads it one time in a render. So this bench shows little of the gain.

selectFirstRow1 and the second commit

The first commit alone made selectFirstRow1 7.3% slower (report). With the second commit it reads -0.5%, which is not significant.

  • Cause: all cells have one hidden class, and V8 records which kind of value each field held so far. selected is undefined until the first select of a row. That select stores the first number in a cell, and V8 throws away the optimized code that reads the field.
  • Proof: a run of the benchmark app with --trace-deopt shows 3 deopts with the reason dependent field representation changed. All 3 are in selectFirstRow1: isSelected, the tracked getter, and one function with a minified name. main has none.
  • Fix, in the second commit: each cell starts with a number in value and gets undefined right after. The field is then general from the first cell on.
  • With the second commit, the trace shows no deopt with that reason.

I did not look for the cause of the two slower clear phases.

Memory

  • Each tracked field of each instance now has a cell object with three fields.
  • A field that is written and never read now gets its tag at the first write. On main, it gets the tag at the first read.
  • I did not measure memory. The finalGc phase of pnpm bench shows no change.
  • Store each @tracked field in a TrackedValue #21662 is an alternative that keeps each field in a TrackedValue. In rere it is 2.7% slower than this cell on the current main, and it was 3.9% slower in the round before.
How this was measured
  • ember.js pnpm bench settings: tracerbench compare on smoke-tests/benchmark-app, headless, fidelity 50. Chrome and tracerbench ran pinned to one CPU core.
  • The host cores have a limit of about 3 GHz, which makes the CPU about 2.2 times slower, so the CPU throttle is 4x and not the usual 8x. The times are close to those of 8x with no limit.
  • Benchmark app change for this run, on control and experiment alike: gc() at the top of runBenchmark(), before the first mark, and a measured finalGc phase that calls gc() after clearItems4.
  • Control is main f693f240ee. Experiment is the second commit, fcfd26b0ea, which is on top of f693f240ee.
  • Table values are the median of 50 per-round ratios. Tracerbench runs control and experiment back to back in each round, so both sides of a pair share the machine state. A value counts as a change when its bootstrap 95% interval leaves out 0 and the Wilcoxon p is below 0.05.
  • "Script time" runs from the Start mark of the phase to the end of the task that holds it: the click plus the render of Ember, with main-thread GC removed. "Full phase" is the phase of tracerbench with main-thread GC removed.
  • rere-benchmark: ember apps in headless Chrome 154 with a GPU and no frame rate limit, 4x throttle on the same host limit, Chrome on 3 pinned cores. A calibration bench runs before and after each run, and a run in a slow period runs again. No run needed that.
  • Deopt trace: headless Chrome with --js-flags=--trace-deopt and no CPU throttle, one run of the benchmark app for each build. A wrapper around performance.mark prints each mark to the same output, so each deopt line belongs to a phase.

Tests

  • 3 new tests in tracking-test.ts. They check that the field and the tag registry use one tag, in each order of tagFor, read, write and dirtyTagFor. An earlier version of this change made a tag of its own in the cell, and these tests cover that bug.
  • The full suite passes locally with main merged in: 9526 pass, 18 skipped, 0 failed.
  • tsc --noEmit, ESLint, Prettier and pnpm lint:docs pass.

🤖 Generated with Claude Code

NullVoxPopuli and others added 2 commits October 5, 2026 13:50
A read or a write of a tracked field went through three maps: the tag
registry (a WeakMap, then a Map per object) and a WeakMap of values.
Each field now has one WeakMap from the instance to a cell with the
value and the tag, so a read is one lookup and consumeTag, and a write
is one lookup and DIRTY_TAG.

The cell takes its tag from the tag registry when it is made. Other code
can ask the registry for the tag before the first read or write of the
field (an observer, a computed chain, notifyPropertyChange), and that
tag must stay the tag of the field.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
All cells have one hidden class, and V8 records which kind of value a
field held so far. If the tracked fields of an app hold no number at
first, the first number makes V8 throw away the optimized code of each
function that reads a cell.

The benchmark app has one tracked field, `selected`, which is
`undefined` until the first select of a row. `--trace-deopt` showed 3
deopts with the reason "dependent field representation changed" in that
phase, and the phase was 7.3% slower than main.

Each cell now starts with a number and gets `undefined` right after, so
the field is general before that code is optimized.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Both sides added an import to tracking-test.ts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants