Skip to content

Skip unchanged blocks in the updating VM - #21669

Draft
NullVoxPopuli-ai-agent wants to merge 3 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/block-guards-alone
Draft

NullVoxPopuli-ai-agent wants to merge 3 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/block-guards-alone

Conversation

@NullVoxPopuli-ai-agent

Copy link
Copy Markdown
Contributor

The updating VM skips a block whose inputs did not change. A list with 1,000 rows where a few rows change now pays for those rows, and one tag check for each of the others.

  • rere-benchmark ember apps: 11.3% faster than main (geometric mean of 13 benches). The list benches with one render for each update are 33% to 45% faster.
  • pnpm bench: script time for the full run is 2.7% lower. Update of each 10th row is 39% to 41% faster. Append, remove and swap are 9% to 26% faster.
  • Costs: clearItems2 is 3.0% slower, and two fan-out benches read 8% to 10% slower.

This PR replaces #21660, which had the same change on top of #21650. This version needs no other PR. Extracted from the spike #21656, and first proposed as #21612.

Commits

  1. dfe3529fe0 Skip unchanged blocks in the updating VM with a per-block tag
  2. fdb505e907 Drop a missed guard, and bring it back later
  3. d7a316d71e Test blocks through more updates than a guard needs to come back

How it works

  • Each block of the updating VM (the body of an {{#if}}, an {{#each}} item, and each other block that can render again) keeps the combined tag of everything it consumed in its last render or update.
  • On an update, the block validates that tag first. If it still holds, the block consumes the tag and returns. Its opcodes do not run.
  • A guard that fails is dropped at once, with no tracking frame. A block that changed is likely to change again, and a frame on every miss made workloads slower where every row changes.
  • A dropped guard comes back after 8 unguarded updates. So a block that changed once is skipped again later, and a block that changes on every update pays for one frame in 8.
  • A re-render of a block puts its guard back at once.
  • alwaysRevalidate goes past the guards.

rere-benchmark

main, this PR, and #21660 ran in 4 mirrored cycles: 8 runs for each build, 5 samples per bench in each run, 0 void runs. Each number is the median of 40 samples, in ms.

Bench main this PR change
1k items, 1 update each (sequentially, async) 607 334 -44.9%
1k items 1 update on 25% (random, async) 183 120 -34.5%
1k items 1 update on 5% (random, async) 54.8 36.9 -32.7%
1k items 1 update on 5% (random) 15.3 12.5 -18.0%
1k items 1 update on 25% (random) 18.1 15.5 -13.9%
1 item, 1k updates (async) 35.2 34.2 -2.8%
1 item, 100k updates (async) 664 658 -0.9%
1 item, 100k updates 24.0 24.1 +0.4%
1k items, 1 update each (sequentially) 26.8 27.0 +0.6%
Incrementing Render Effect 1522 1541 +1.3%
1 item, 1k updates 5.5 5.8 +4.5%
1 value, 1k consumers, 10k updates (bursts of 1000) 111 120 +8.1%
1 value, 1k consumers, 10k updates (single burst) 29.4 32.4 +10.2%
DB Monitor w/ chat simulation (fps) 113 114 +0.7% fps

pnpm bench

Tracerbench compare against main, 50 rounds: tracerbench-report.pdf.

phase main script ms script time full phase
all phases 7741 -2.7% -3.1%
clearItems1 84 no change +8.2% slower
render1000Items2 248 -5.7% -3.6%
clearItems2 59 +3.0% slower +1.9% slower
render10000Items2 1525 no change -2.7%
render1000Items3 141 -5.1% -9.4%
append1000Items1 251 -18.0% -10.6%
append1000Items2 200 -13.7% -9.0%
updateEvery10thItem1 49 -40.7% -9.0%
updateEvery10thItem2 49 -38.7% -4.7%
selectFirstRow1 83 -11.3% -7.7%
selectSecondRow1 76 -2.7% no change
removeFirstRow1 76 -9.0% no change
removeSecondRow1 78 -26.3% -5.9%
swapRows1 64 -19.8% -2.7%
swapRows2 69 -15.6% no change
finalGc 862 -27.1% -27.1%

The table shows only the phases with a significant change.

  • clearItems1 is slower only as a full phase. The script time is the same, so the extra time is after the render task. I did not find the cause.
  • The final gc() is 27% faster.
How this was measured
  • ember.js pnpm bench settings: tracerbench compare on smoke-tests/benchmark-app, headless, fidelity 50. Chrome and tracerbench ran pinned to one CPU core.
  • The host cores have a limit of about 3 GHz, which makes the CPU about 2.2 times slower, so the CPU throttle is 4x and not the usual 8x. The times are close to those of 8x with no limit.
  • Benchmark app change for this run, on control and experiment alike: gc() at the top of runBenchmark(), before the first mark, and a measured finalGc phase that calls gc() after clearItems4.
  • Control is main f693f240ee. Experiment is this PR at d7a316d71e, which is on top of f693f240ee.
  • Table values are the median of 50 per-round ratios. Tracerbench runs control and experiment back to back in each round, so both sides of a pair share the machine state. A value counts as a change when its bootstrap 95% interval leaves out 0 and the Wilcoxon p is below 0.05.
  • "Script time" runs from the Start mark of the phase to the end of the task that holds it: the click plus the render of Ember, with main-thread GC removed. "Full phase" is the phase of tracerbench with main-thread GC removed.
  • rere-benchmark: ember apps in headless Chrome 154 with a GPU and no frame rate limit, 4x throttle on the same host limit, Chrome on 3 pinned cores. A calibration bench runs before and after each run, and a run in a slow period runs again. No run needed that.

Tests

  • 6 new tests in block-guards-test.ts. Each one runs a block, a nested block, or a list through 20 updates of one kind and then changes something else, so the block is skipped, dropped, and guarded again while the test checks the output.
  • The full suite passes locally: 9501 pass, 18 skipped, 0 failed.
  • tsc --noEmit, ESLint and Prettier pass.

🤖 Generated with Claude Code

NullVoxPopuli and others added 3 commits October 5, 2026 17:13
Every block opcode (try, list, list item) now records the combined tag of
what its render or last update consumed, the way a component cache group
does. On update, a block whose tag still validates is skipped as a whole,
and its tag is consumed into the parent frame so parents stay correct.

Before, an unchanged row in a {{#each}} cost one validation per dynamic
reference: on the js-framework-benchmark row that is five validateTag
calls and five megamorphic opcode evaluations per row per update. Now it
costs one validation, and the five opcodes never run.

The append VM opens the frame before a block's opcode is constructed,
because the list block reads its iterable in the constructor, and closes it
when the block exits. On the updating side the frame is closed when the
block's updating frame finishes. When a block re-renders after a thrown
assertion, the append VM closes the frame on exit and the updating VM pops
the frame without closing it again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wPae8QnFbKHKn5xbgaQVq
A guard that fails is dropped on the spot, with no tracking frame: a block
that changed is likely to change again, and a frame on every miss made
all-rows-change workloads slower. A dropped guard comes back after eight
unguarded updates, so a block that changed once and then stayed still is
skipped again, while a block that changes on every update pays for one
frame in every eight. A re-render brings the guard back at once.

An unguarded block costs what it did before guards existed: one frame
push. A re-render of a dropped block opens its own tracking frame, because
the append VM closes one when the block exits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A dropped guard comes back after 8 updates. Each test runs a block, a
nested block, or a list through 20 updates of one kind and then changes
something else, so the block is skipped, dropped, and guarded again
while the test checks the output.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants