Repository navigation
Skip unchanged blocks in the updating VM - #21660
Closed
NullVoxPopuli-ai-agent wants to merge 8 commits into
Closed
NullVoxPopuli-ai-agent wants to merge 8 commits into
NullVoxPopuli-ai-agent wants to merge 8 commits into
Conversation
A tracking frame allocated one tracker, one Set, and one array from that Set. In CPU profiles of three createCache graphs, the tracker, its Set and beginTrackFrame took 27% to 45% of the samples. Frames are strictly nested, so one tracker for each depth is enough. beginTrackFrame takes the tracker of its depth from a pool. A tracker keeps its tags in an array. Each frame has a number, and the tracker writes that number on each tag that it takes, so a tag that the frame consumes again costs one comparison and no Set. endTrackFrame now takes the tag that the same frame produced the last time. If the frame consumed the same tags again, that tag is the result, with its memoized revision and no allocation. getValue passes the tag of its cache. The loop over the subtags of a combined tag uses an index in place of an iterator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A TrackedValue made four arrow functions and one options object for each instance. A value that the code reads and writes through `value` did not use any of them. `get`, `set`, `update` and `freeze` are now accessors that make the bound function on its first read, and keep it. They are bound as before, and each read gives the same function. `trackedValue(value)` with no options shares one options object. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tracker wrote the number of its frame on each tag. That counter leaves the small-integer range of V8 after about one billion frames, and the benchmark is then 1.1 to 1.3 times slower. A nested frame also wrote its own number on a shared tag, so the outer frame took that tag again each time: a computed that reads five computeds of one tag had a combined tag with five entries. A tag now keeps the index at which a tracker took it. A tracker has the tag if its entry at that index is the tag. This needs no counter, and a stale index cannot hide a tag, because the entry is then another tag or no tag. A nested frame that takes the tag at the same index keeps the check true, so the five computeds give one tag again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`get`, `set`, `update` and `freeze` were own properties before, so code could assign to them. Each accessor now has a setter. A write through `value` uses an assigned `set`, as it did before. The shared default options are frozen. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An untrack frame takes a depth but no tracker. A track frame that begins inside one, at a depth that never had a tracker, leaves a hole below it. resetTracking then called clear() on undefined. The QUnit setup calls resetTracking after each test, so a filtered run of the tracking module stopped at the first test after the hole. The full suite fills every depth before that, so it did not show the error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every block opcode (try, list, list item) now records the combined tag of
what its render or last update consumed, the way a component cache group
does. On update, a block whose tag still validates is skipped as a whole,
and its tag is consumed into the parent frame so parents stay correct.
Before, an unchanged row in a {{#each}} cost one validation per dynamic
reference: on the js-framework-benchmark row that is five validateTag
calls and five megamorphic opcode evaluations per row per update. Now it
costs one validation, and the five opcodes never run.
The append VM opens the frame before a block's opcode is constructed,
because the list block reads its iterable in the constructor, and closes it
when the block exits. On the updating side the frame is closed when the
block's updating frame finishes. When a block re-renders after a thrown
assertion, the append VM closes the frame on exit and the updating VM pops
the frame without closing it again.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wPae8QnFbKHKn5xbgaQVq
…it later Blocks and component cache groups pass their previous tag to endTrackFrame(), so a block that re-renders with the same dependencies keeps its tag and its memoized revision. The tag comparison against the tracker's live entries comes from emberjs#21650, which is below this commit. A guard that fails is dropped on the spot, with no tracking frame: a block that changed is likely to change again, and a frame on every miss is what made all-rows-change workloads slower. A dropped guard comes back after eight unguarded updates, so a block that changed once and then stayed still is skipped again, while a block that changes on every update pays for one frame in every eight. A re-render re-arms the guard right away. An unguarded block costs what it did before guards existed: one frame push. A re-render of a dropped block opens its own tracking frame, because the append VM closes one when the block exits. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012wPae8QnFbKHKn5xbgaQVq
A dropped guard comes back after 8 updates. Each test runs a block, a nested block, or a list through 20 updates of one kind and then changes something else, so the block is skipped, dropped, and guarded again while the test checks the output. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The updating VM skips a block whose inputs did not change. A list with 1,000 rows where a few rows change now pays for those rows, and one tag check for each of the others.
main(geometric mean of 13 benches). The list benches are 9% to 46% faster.pnpm bench: script time for the full run is 3.4% lower. Update of each 10th row is 40% faster. Select, append, remove and swap are 12% to 25% faster.This is the change of #21612, on top of #21650. Extracted from the spike #21656.
Stacked on #21650
The first five commits are #21650 (the tracker pool, and
endTrackFrame(previous)). This PR needs both. The commits of this PR are the last three:28d8a42a17Skip unchanged blocks in the updating VM with a per-block tag88cb6b3c47Reuse block and cache group tags; drop a missed guard, re-arm it later4ed19218aeTestsHow it works
{{#if}}, an{{#each}}item, and each other block that can render again) keeps the combined tag of everything it consumed in its last render or update.alwaysRevalidategoes past the guards.rere-benchmark
main, #21650, and this PR ran in 4 mirrored cycles together with other builds: 8 runs for each build, 5 samples per bench in each run. Each number is the median of 40 samples, in ms.mainmainmain(geometric mean). The four cycles read -16.5%, -14.6%, -13.5%, and -12.1%. Against Cut the allocations of a tracking frame and of a TrackedValue #21650 it is 12.7% faster.mainalone.mainand from 113 to 205 ms for this PR.pnpm benchTracerbench compare against
main, 50 rounds: tracerbench-report.pdf.The table shows only the phases with a significant change.
render1000Items3is not slower. Its script time with GC included reads -2.0%, which is not significant. Onmain, a collection runs inside that phase and the table removes it. With this PR, the collection runs one phase earlier.render10000Items2andclearItems1are slower only as a full phase. The script time is the same, so the extra time is after the render task. I did not find the cause yet.gc()is 22% faster.How this was measured
pnpm benchsettings: tracerbench compare onsmoke-tests/benchmark-app, headless, fidelity 50. Chrome and tracerbench ran pinned to one CPU core.gc()at the top ofrunBenchmark(), before the first mark, and a measuredfinalGcphase that callsgc()afterclearItems4.f693f240ee. Experiment is this PR at88cb6b3c47, which is on top off693f240ee. The last commit only adds tests.Tests
block-guards-test.ts. Each one runs a block, a nested block, or a list through 20 updates of one kind and then changes something else, so the block is skipped, dropped, and guarded again while the test checks the output.tsc --noEmit, ESLint and Prettier pass.🤖 Generated with Claude Code