Skip to content

Skip unchanged {{#each}} item subtrees during updates - #21544

Closed
NullVoxPopuli-ai-agent wants to merge 2 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:extract-each-skip
Closed

NullVoxPopuli-ai-agent wants to merge 2 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:extract-each-skip

Conversation

@NullVoxPopuli-ai-agent

Copy link
Copy Markdown
Contributor

Re-cut of #21512 (closed during the #21520 spike consolidation) onto current main, now with the full-suite verification that was outstanding: 9443 tests, 0 failures locally.

During revalidation, each {{#each}} item collects the tags its subtree consumed (via a tracking-frame finalizer on the item's try-frame); on later passes, an item whose collected tag validates clean is skipped entirely instead of walking every opcode of its subtree. A triviality gate (≤2 opcodes and no nested block) keeps the bookkeeping off items too small to profit. Measured standalone at ~1.9x on an 8x-throttled dbmon (walk-heavy, many clean items per frame); no observable behavior change — the same subtrees re-render for the same reasons, cheaper.

This was the campaign's first lever and its most reviewed artifact; it needs no RFC, no flag, and composes with (but does not require) the scheduler work in #21493/#21520.

🤖 Generated with Claude Code

@NullVoxPopuli-ai-agent

This comment was marked as outdated.

@NullVoxPopuli

This comment was marked as outdated.

Comment thread packages/@glimmer/runtime/lib/vm/update.ts
@NullVoxPopuli

This comment was marked as outdated.

@NullVoxPopuli-ai-agent

This comment was marked as outdated.

@NullVoxPopuli-ai-agent

This comment was marked as outdated.

The UpdatingVM walks every updating opcode of every list item on every
render: cache groups (JumpIfNotModifiedOpcode) exist only at component
boundaries, so a list of plain template rows revalidates every binding
even when nothing in a row changed.

Collect each item's consumed tags in a tracking frame (via a new
frame-finalizer hook on UpdatingVMFrame) and skip the item's entire
subtree while that combined tag validates.

Trivial items opt out: for a text node or two, validating a combined
tag costs as much as updating, so collection would be pure overhead.
An item is trivial when it has <= 2 opcodes and no nested block -- a
nested block child means an arbitrarily large subtree hides behind a
small top-level count.

dbmon-style workloads (fat rows, sparse changes): ~1.6x fps at 8x CPU
throttle, ~6x (rAF-capped) at 4x. Dense-change / tiny-item workloads
and the krausest bench: neutral.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@NullVoxPopuli-ai-agent

Copy link
Copy Markdown
Contributor Author

Tracerbench PDF for a run with gc() before the first measurement and a measured gc() at the end: tracerbench-report.pdf

Ember's script time is 40 to 41% lower on update-every-10th, 21 to 22% on remove, and 29 to 31% on swap. Selecting the second row is 15.7% slower, as in both earlier runs. The final gc() takes as long as on main, so this PR leaves no extra garbage behind.

In the same setup, #21557 measured update -40%, remove -20 to -25%, swap -23%, selectSecondRow1 +9.9% slower, and a 6.6% slower first append.

phase main script ms (8x) script time full phase
all phases 10199 -1.3% no change
updateEvery10thItem1 87 -40.9% -4.8%
updateEvery10thItem2 86 -40.3% -3.8%
selectSecondRow1 127 +15.7% slower +11.1% slower
removeFirstRow1 124 -21.9% -4.9%
removeSecondRow1 122 -21.3% -4.2%
swapRows1 107 -31.3% -8.9%
swapRows2 107 -29.1% -10.0%
How this was measured
  • ember.js pnpm bench settings: tracerbench compare on smoke-tests/benchmark-app, 8x CPU throttle, headless, fidelity 50.
  • Benchmark app change for this run, on control and experiment alike: gc() at the top of runBenchmark(), before the first mark, and a measured finalGc phase that calls gc() after clearItems4. Chrome runs with --js-flags=--expose-gc. Patch: benchmark-app-gc.patch
  • Control is main 153364bb9a. Experiment is this PR on top of 153364bb9a (merged locally when the branch is behind).
  • Chrome and tracerbench ran pinned to one CPU core. I also added --disable-background-networking,--disable-component-update to the Chrome flags and raised --sampleTimeout to 180 s.
  • The PDF is tracerbench's own report. The only change is that its results folder reads results/<label> instead of a local path.
  • Table values are the median of 50 per-round ratios. Tracerbench runs control and experiment back to back in each round, so both sides of a pair share the machine state. A value counts as a change when its bootstrap 95% CI excludes 0 and a Wilcoxon signed-rank test gives p < 0.05.
  • "Script time" runs from the phase's Start mark to the end of the task that holds it: the click plus Ember's render. It leaves out layout, paint, the benchmark's row checks and idle time. Main-thread GC inside that window is removed.
  • "Full phase" is tracerbench's own phase with main-thread GC removed.
  • An earlier main-vs-main run (without the gc() calls) had false hits of up to +5% on script time.
  • Raw results for every run: bench-reports branch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants