Skip to content

Skip unchanged blocks in the updating VM - #21660

Closed
NullVoxPopuli-ai-agent wants to merge 8 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/block-guards-on-21650
Closed

NullVoxPopuli-ai-agent wants to merge 8 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/block-guards-on-21650

Conversation

@NullVoxPopuli-ai-agent

Copy link
Copy Markdown
Contributor

The updating VM skips a block whose inputs did not change. A list with 1,000 rows where a few rows change now pays for those rows, and one tag check for each of the others.

  • rere-benchmark ember apps: 15% faster than main (geometric mean of 13 benches). The list benches are 9% to 46% faster.
  • pnpm bench: script time for the full run is 3.4% lower. Update of each 10th row is 40% faster. Select, append, remove and swap are 12% to 25% faster.
  • Costs: the clear phases are 3% to 7% slower, and the second render of 10,000 rows is 13% slower as a full phase.

This is the change of #21612, on top of #21650. Extracted from the spike #21656.

Stacked on #21650

The first five commits are #21650 (the tracker pool, and endTrackFrame(previous)). This PR needs both. The commits of this PR are the last three:

  1. 28d8a42a17 Skip unchanged blocks in the updating VM with a per-block tag
  2. 88cb6b3c47 Reuse block and cache group tags; drop a missed guard, re-arm it later
  3. 4ed19218ae Tests

How it works

  • Each block of the updating VM (the body of an {{#if}}, an {{#each}} item, and each other block that can render again) keeps the combined tag of everything it consumed in its last render or update.
  • On an update, the block validates that tag first. If it still holds, the block consumes the tag and returns. Its opcodes do not run.
  • A guard that fails is dropped at once, with no tracking frame. A block that changed is likely to change again, and a frame on every miss made workloads slower where every row changes.
  • A dropped guard comes back after 8 unguarded updates. So a block that changed once is skipped again later, and a block that changes on every update pays for one frame in 8.
  • A re-render of a block puts its guard back at once.
  • alwaysRevalidate goes past the guards.

rere-benchmark

main, #21650, and this PR ran in 4 mirrored cycles together with other builds: 8 runs for each build, 5 samples per bench in each run. Each number is the median of 40 samples, in ms.

Bench main #21650 this PR against main against #21650
1k items, 1 update each (sequentially, async) 768 747 412 -46.4% -44.9%
1k items 1 update on 25% (random, async) 243 235 132 -45.5% -43.6%
1k items 1 update on 5% (random, async) 76.9 68.4 43.3 -43.7% -36.6%
1k items 1 update on 5% (random) 18.0 17.5 14.3 -20.6% -18.3%
1k items 1 update on 25% (random) 23.9 24.0 20.5 -14.4% -14.8%
1k items, 1 update each (sequentially) 32.9 33.0 29.8 -9.4% -10.0%
1 item, 100k updates (async) 883 833 824 -6.6% -1.1%
Incrementing Render Effect 1958 1879 1902 -2.8% +1.2%
1 item, 1k updates (async) 41.6 42.8 41.7 +0.2% -2.6%
1 value, 1k consumers, 10k updates (single burst) 40.9 40.1 41.6 +1.6% +3.5%
1 item, 1k updates 6.4 6.0 6.6 +3.1% +10.0%
1 item, 100k updates 29.4 30.2 31.8 +7.8% +5.1%
1 value, 1k consumers, 10k updates (bursts of 1000) 124 125 149 +20.5% +19.0%
DB Monitor w/ chat simulation (fps) 97.0 98.7 100 +3.4% fps +1.5% fps
  • Over 13 benches, this PR is 14.8% faster than main (geometric mean). The four cycles read -16.5%, -14.6%, -13.5%, and -12.1%. Against Cut the allocations of a tracking frame and of a TrackedValue #21650 it is 12.7% faster.
  • The geometric mean leaves out DB Monitor and fan-out bursts of 100. Both depend on frame timing, and the run medians of fan-out bursts of 100 go from 792 to 1503 ms for main alone.
  • Fan-out bursts of 1000 changes all 1,000 consumers on every render, so every guard fails. This PR reads 19% slower there. That bench is also noisy: its run medians go from 97 to 161 ms for main and from 113 to 205 ms for this PR.

pnpm bench

Tracerbench compare against main, 50 rounds: tracerbench-report.pdf.

phase main script ms script time full phase
all phases 7067 -3.4% no change
clearItems1 79 no change +17.5% slower
render1000Items2 227 -9.1% -5.7%
render10000Items1 1802 -1.3% no change
clearManyItems1 495 +6.9% slower +5.7% slower
render10000Items2 1372 no change +13.4% slower
clearManyItems2 469 +2.7% slower +2.5% slower
render1000Items3 129 +30.3% slower (see below) no change
append1000Items1 232 -18.6% -6.8%
append1000Items2 183 -15.5% -11.7%
updateEvery10thItem1 45 -40.2% -11.3%
updateEvery10thItem2 49 -41.3% -16.1%
selectFirstRow1 79 -18.1% -10.4%
selectSecondRow1 74 -11.5% -7.2%
removeFirstRow1 72 -12.5% no change
removeSecondRow1 73 -24.6% -3.4%
swapRows1 56 -19.4% -10.4%
swapRows2 65 -13.9% -3.7%
clearItems4 150 +3.1% slower +2.7% slower
finalGc 841 -21.8% -21.8%

The table shows only the phases with a significant change.

  • render1000Items3 is not slower. Its script time with GC included reads -2.0%, which is not significant. On main, a collection runs inside that phase and the table removes it. With this PR, the collection runs one phase earlier.
  • The clear phases are 3% to 7% slower in both columns. I did not look for the cause yet.
  • render10000Items2 and clearItems1 are slower only as a full phase. The script time is the same, so the extra time is after the render task. I did not find the cause yet.
  • The final gc() is 22% faster.
How this was measured
  • ember.js pnpm bench settings: tracerbench compare on smoke-tests/benchmark-app, headless, fidelity 50. Chrome and tracerbench ran pinned to one CPU core.
  • The host cores have a limit of about 3 GHz, which makes the CPU about 2.2 times slower, so the CPU throttle is 4x and not the usual 8x. The times are close to those of 8x with no limit.
  • Benchmark app change for this run, on control and experiment alike: gc() at the top of runBenchmark(), before the first mark, and a measured finalGc phase that calls gc() after clearItems4.
  • Control is main f693f240ee. Experiment is this PR at 88cb6b3c47, which is on top of f693f240ee. The last commit only adds tests.
  • Table values are the median of 50 per-round ratios. Tracerbench runs control and experiment back to back in each round, so both sides of a pair share the machine state. A value counts as a change when its bootstrap 95% interval leaves out 0 and the Wilcoxon p is below 0.05.
  • "Script time" runs from the Start mark of the phase to the end of the task that holds it: the click plus the render of Ember, with main-thread GC removed. "Full phase" is the phase of tracerbench with main-thread GC removed.
  • rere-benchmark: ember apps in headless Chrome 154 with a GPU and no frame rate limit, 4x throttle on the same host limit, Chrome on 3 pinned cores. A calibration bench runs before and after each run, and a run in a slow period runs again. No run needed that.

Tests

  • 6 new tests in block-guards-test.ts. Each one runs a block, a nested block, or a list through 20 updates of one kind and then changes something else, so the block is skipped, dropped, and guarded again while the test checks the output.
  • The full suite passes locally: 9513 pass, 18 skipped, 0 failed.
  • tsc --noEmit, ESLint and Prettier pass.

🤖 Generated with Claude Code

NullVoxPopuli and others added 8 commits October 2, 2026 01:05
A tracking frame allocated one tracker, one Set, and one array from
that Set. In CPU profiles of three createCache graphs, the tracker, its
Set and beginTrackFrame took 27% to 45% of the samples.

Frames are strictly nested, so one tracker for each depth is enough.
beginTrackFrame takes the tracker of its depth from a pool.

A tracker keeps its tags in an array. Each frame has a number, and the
tracker writes that number on each tag that it takes, so a tag that the
frame consumes again costs one comparison and no Set.

endTrackFrame now takes the tag that the same frame produced the last
time. If the frame consumed the same tags again, that tag is the
result, with its memoized revision and no allocation. getValue passes
the tag of its cache.

The loop over the subtags of a combined tag uses an index in place of
an iterator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A TrackedValue made four arrow functions and one options object for
each instance. A value that the code reads and writes through `value`
did not use any of them.

`get`, `set`, `update` and `freeze` are now accessors that make the
bound function on its first read, and keep it. They are bound as
before, and each read gives the same function.

`trackedValue(value)` with no options shares one options object.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tracker wrote the number of its frame on each tag. That counter
leaves the small-integer range of V8 after about one billion frames,
and the benchmark is then 1.1 to 1.3 times slower. A nested frame also
wrote its own number on a shared tag, so the outer frame took that tag
again each time: a computed that reads five computeds of one tag had a
combined tag with five entries.

A tag now keeps the index at which a tracker took it. A tracker has the
tag if its entry at that index is the tag. This needs no counter, and a
stale index cannot hide a tag, because the entry is then another tag or
no tag. A nested frame that takes the tag at the same index keeps the
check true, so the five computeds give one tag again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`get`, `set`, `update` and `freeze` were own properties before, so
code could assign to them. Each accessor now has a setter. A write
through `value` uses an assigned `set`, as it did before.

The shared default options are frozen.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An untrack frame takes a depth but no tracker. A track frame that begins
inside one, at a depth that never had a tracker, leaves a hole below it.
resetTracking then called clear() on undefined. The QUnit setup calls
resetTracking after each test, so a filtered run of the tracking module
stopped at the first test after the hole. The full suite fills every
depth before that, so it did not show the error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every block opcode (try, list, list item) now records the combined tag of
what its render or last update consumed, the way a component cache group
does. On update, a block whose tag still validates is skipped as a whole,
and its tag is consumed into the parent frame so parents stay correct.

Before, an unchanged row in a {{#each}} cost one validation per dynamic
reference: on the js-framework-benchmark row that is five validateTag
calls and five megamorphic opcode evaluations per row per update. Now it
costs one validation, and the five opcodes never run.

The append VM opens the frame before a block's opcode is constructed,
because the list block reads its iterable in the constructor, and closes it
when the block exits. On the updating side the frame is closed when the
block's updating frame finishes. When a block re-renders after a thrown
assertion, the append VM closes the frame on exit and the updating VM pops
the frame without closing it again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wPae8QnFbKHKn5xbgaQVq
…it later

Blocks and component cache groups pass their previous tag to
endTrackFrame(), so a block that re-renders with the same dependencies
keeps its tag and its memoized revision. The tag comparison against the
tracker's live entries comes from emberjs#21650, which is below this commit.

A guard that fails is dropped on the spot, with no tracking frame: a block
that changed is likely to change again, and a frame on every miss is what
made all-rows-change workloads slower. A dropped guard comes back after
eight unguarded updates, so a block that changed once and then stayed
still is skipped again, while a block that changes on every update pays
for one frame in every eight. A re-render re-arms the guard right away.

An unguarded block costs what it did before guards existed: one frame
push. A re-render of a dropped block opens its own tracking frame, because
the append VM closes one when the block exits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wPae8QnFbKHKn5xbgaQVq
A dropped guard comes back after 8 updates. Each test runs a block, a
nested block, or a list through 20 updates of one kind and then changes
something else, so the block is skipped, dropped, and guarded again
while the test checks the output.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@NullVoxPopuli-ai-agent

Copy link
Copy Markdown
Contributor Author

Replaced by #21669. It has the same change on main alone, so it needs no other PR. The gain holds without #21650: 11.3% in rere-benchmark and 2.7% of script time in pnpm bench.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants