Skip to content

Reuse the combined tag of a cache that reads the same tags again - #21664

Open
NullVoxPopuli-ai-agent wants to merge 5 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/reuse-frame-tag
Open

NullVoxPopuli-ai-agent wants to merge 5 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/reuse-frame-tag

Conversation

@NullVoxPopuli-ai-agent

@NullVoxPopuli-ai-agent NullVoxPopuli-ai-agent commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

A cache that consumes the same tags as in its last run returns the same combined tag. It then allocates no tag and no array, and the tag keeps its memoized revision.

This helps createCache, @cached and invokeHelper. A cache that reads two or more tracked values runs again 25% to 28% faster. The worst case found is 5% slower: a cache with 100 tags where only the last tag changes.

case main this PR ratio ratio of each round
wide: 2 signals, same signals 127 ns 92 ns 0.72x 0.68 to 0.73
wide: 2 signals, other signals 148 ns 148 ns 1.00x 0.98 to 1.08
wide: 10 signals, same signals 297 ns 221 ns 0.75x 0.74 to 0.84
wide: 10 signals, other signals 300 ns 299 ns 1.00x 0.96 to 1.04
wide: 100 signals, same signals 1.88 µs 1.41 µs 0.75x 0.74 to 0.75
wide: 100 signals, other signals 2.14 µs 2.15 µs 1.00x 0.99 to 1.17
branch: a || b || c 105 ns 103 ns 0.98x 0.95 to 1.33
branch: 10 or 20 signals 376 ns 378 ns 1.00x 0.83 to 1.03
branch: 100 signals, another last signal 1.87 µs 1.96 µs 1.05x 1.03 to 1.06
branch: 10 random signals of 20 398 ns 405 ns 1.02x 1.00 to 1.11
branch: 10 random choices of 2 signals 441 ns 444 ns 1.01x 0.99 to 1.03
branch: a random number of signals, 1 to 20 318 ns 311 ns 0.98x 0.96 to 1.10
branch: 100 signals, a new set each frame 2.09 µs 2.12 µs 1.01x 1.00 to 1.02
branch: 100 signals, a new set each frame, nothing shared 2.15 µs 2.16 µs 1.00x 0.99 to 1.01
kairo: mux 12.06 µs 9.70 µs 0.80x 0.77 to 0.82
batch: 10 writes, 1 output 322 ns 291 ns 0.90x 0.80 to 0.92
  • "Same signals" is the usual case: a @cached getter runs again and reads the same tracked values.
  • The gain does not depend on the number of values.
  • "Other signals" and the branch cases have a cache that reads other values in each run, so the check fails.
  • The branch rows are from a later run of 10 rounds.
  • A change of the number of values costs nothing, because the check compares the two lengths first. a || b || c is such a case.
  • A whole new set of 100 values in every run costs 0% to 1%, for the same reason as a random set.
  • A random set of values in every run costs 0% to 2%. The old list and the new list differ at one of the first entries, so the check stops there.
  • The worst case is a list of tags that is equal up to the last entry. The check then compares every tag and fails. That is about 1 ns for each tag, and 5% with 100 tags.
  • The code of rendering does not change. The frames of the render VM and of the curly component manager end in endTrackFrame, which is the code of main.

The branch is five commits on main (9bec1cb2a8), and it merges alone.

What changes

  • getValue ends its frame with a new function, endCacheFrame(previous). It passes the tag that the cache has from its last run.
  • If the cache consumed the same tags again, in the same order, the result is that tag.
  • endTrackFrame and Tracker#combine do not change.

What changes for callers

Nothing. No exported function changes its signature, and endCacheFrame is not exported.

Why the reuse is in code of its own

The first two versions of this PR had the reuse in Tracker#combine, behind an argument of endTrackFrame.

  1. With the check as one more branch of combine, three cases with small frames were 4% to 13% slower than main. The method was too large for V8 to put it inline into the end of the frame, so each frame paid for a call.
  2. With the check in a function of its own, combineOrReuse, that loss went away, apart from 1% to 4% in two cases.
  3. Now only getValue reaches the reuse, through endCacheFrame and Tracker#combineForCache. The small frames are level with main again, and all other frames run the code of main.

Passing the last tag from the render VM and from the curly component manager was also tested. It showed no gain in pnpm bench: comment on #21650.

All 26 cases
case main this PR ratio ratio of each round
propagate: 1 chains x 1 deep 58 ns 56 ns 0.97x 0.69 to 1.07
propagate: 10 chains x 10 deep 2.60 µs 2.57 µs 0.99x 0.88 to 1.05
propagate: 100 chains x 100 deep 308.41 µs 298.03 µs 0.97x 0.95 to 1.00
propagate: 1 chains x 1000 deep 23.09 µs 22.94 µs 0.99x 0.85 to 1.00
propagate: 1000 chains x 1 deep 46.55 µs 46.17 µs 0.99x 0.70 to 1.01
kairo: avoidable propagation 139 ns 141 ns 1.01x 0.98 to 1.09
kairo: broad propagation 3.43 µs 3.40 µs 0.99x 0.97 to 1.01
kairo: deep propagation 1.24 µs 1.19 µs 0.96x 0.91 to 1.02
kairo: diamond 186 ns 186 ns 1.00x 0.97 to 1.02
kairo: mux 12.06 µs 9.70 µs 0.80x 0.77 to 0.82
kairo: repeated observers 246 ns 246 ns 1.00x 0.65 to 1.08
kairo: triangle 327 ns 326 ns 1.00x 0.92 to 1.06
kairo: unstable 334 ns 331 ns 0.99x 0.98 to 1.00
rows: 1000 rows, write 1 6.70 µs 6.69 µs 1.00x 0.96 to 1.24
rows: 1000 rows, write all 54.52 µs 54.43 µs 1.00x 0.92 to 1.01
batch: 10 writes, 1 output 322 ns 291 ns 0.90x 0.80 to 0.92
avoidable: write the same value 6 ns 6 ns 1.00x 1.00 to 1.00
create: 1000 signals 9.36 µs 9.35 µs 1.00x 1.00 to 1.00
create: 1000 computeds, read each 23.43 µs 23.83 µs 1.02x 0.90 to 1.10
create: 1000 outputs 34.16 µs 34.15 µs 1.00x 0.99 to 1.00
wide: 2 signals, same signals 127 ns 92 ns 0.72x 0.68 to 0.73
wide: 2 signals, other signals 148 ns 148 ns 1.00x 0.98 to 1.08
wide: 10 signals, same signals 297 ns 221 ns 0.75x 0.74 to 0.84
wide: 10 signals, other signals 300 ns 299 ns 1.00x 0.96 to 1.04
wide: 100 signals, same signals 1.88 µs 1.41 µs 0.75x 0.74 to 0.75
wide: 100 signals, other signals 2.14 µs 2.15 µs 1.00x 0.99 to 1.17

The benchmark is https://github.com/NullVoxPopuli-ai-agent/ember-reactivity-bench. One measurement is the writes of one frame, then one flush that brings every output up to date. Each case of each build runs in its own process, pinned to one core. The numbers are the median of 8 mirrored rounds on Node 24.20.

  • main is 9bec1cb2a8. This PR is 6ab63a155e.
  • Every computed of this benchmark is a createCache. No case measures rendering.
  • The six wide cases and the eight branch cases are new. The table above has the branch cases, and this page has their run.
  • kairo: avoidable propagation is 1.01x, and slower in 7 of 8 rounds. That is 1.5 ns for 6 caches.
  • kairo: repeated observers has two stable speeds in this benchmark, also between two runs of the same code.
  • Rendering benchmarks did not run. The benchmark app of Ember and the Ember apps of rere-benchmark have no @cached, no createCache and no invokeHelper, so they cannot reach the new code.
Tests
  • The tests of @glimmer/validator: tracking pass: 40 of 40. The full suite passed before the last two commits, which change tests only: 9525 pass, 18 skipped, 0 failed.
  • type-check:internals, ESLint and Prettier pass.
  • Each cache in the new tests records a step that says what it reads, and the test checks the steps after each read.
  • One new test fails on main: a cache that consumes the same tags in its next run keeps its tag.
  • Two new tests need a new tag: for other tags, and for the same tags in another order. They fail if the check answers "same tags" for every combined tag.
  • Two more tests check that a cache still follows its tags. One cache changes its tags between two runs, and one outer cache reads an inner cache with a reused tag. These pass on main too.

Related

This is the last open part of #21650. The other parts are merged: #21663, #21665, #21666 and #21667. #21650 has the same commits as this PR now, and it can close when this PR merges.

🤖 Generated with Claude Code

endTrackFrame takes the tag that the same frame produced the last time
it ran. If the frame consumed the same tags again, the result is that
tag: it keeps its memoized revision, and the frame allocates no tag and
no array.

getValue passes the tag of its cache.

Split out of emberjs#21650.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
NullVoxPopuli and others added 2 commits October 7, 2026 14:40
With the check in `combine`, V8 did not inline the method into the end
of the frame, and each small frame paid for a call. Against main, a
frame with one tag was 4% to 7% slower, and `kairo: avoidable
propagation` was 13% slower.

The check is now in a function of its own, `combineOrReuse`, which only
frames with two or more tags call. No case is more than 2% slower than
main, and `kairo: mux` and `batch` keep their gain.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Only `getValue` passes the last tag, so only a cache can reuse it.
The reuse is now in code that only `getValue` calls:
`endCacheFrame` and `Tracker#combineForCache`.

`endTrackFrame` and `Tracker#combine` are the code of main again, with
no argument. The frames of the render VM and of the curly component
manager end there, so this change cannot make them slower.

The two tests that passed a tag to `endTrackFrame` are gone. The tests
of `createCache` cover the reuse.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@NullVoxPopuli-ai-agent NullVoxPopuli-ai-agent changed the title Reuse the combined tag of a tracking frame Reuse the combined tag of a cache that reads the same tags again Oct 7, 2026
The tests of this change passed with the code of main too, because they
only checked that a cache still follows its tags.

The new tests read the tag of a cache through an outer frame. One of
them fails on main: a cache that consumes the same tags again must keep
its tag. Two more need a new tag: for other tags, and for the same tags
in another order. Those two fail if the check says "same tags" for
every combined tag.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@NullVoxPopuli

NullVoxPopuli commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Tests pass here: NullVoxPopuli/limber#2289 / https://test-ember-source-nvp-reuse.limber-glimdown.pages.dev/
(as does some spot-checking in the deploy)

however, the whole codebase only has 4 explicit @cached
https://github.com/search?q=repo%3ANullVoxPopuli%2Flimber+%40cached&type=code

(the rest of the createCache would be internals to ember)

*/
function combineOrReuse(previous: Tag | undefined, tags: (Tag | null)[], size: number): Tag {
if (previous !== undefined && isCombinationOf(previous, tags, size)) {
return previous;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is the reuse here -- no new tracked things were encountered, so re-calling combine is a little silly

let second = tagOf(cache);

assert.strictEqual(count, 2, 'the cache ran again');
assert.strictEqual(second, first, 'the cache has the tag of its first run');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

noice

Each cache now records a step with a text that says what it reads, for
example "cache reads tag1 and tag3". A test checks the steps after each
read, in place of a counter that went from 1 to 5.

A read that must not run the cache checks for no step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@NullVoxPopuli
NullVoxPopuli marked this pull request as ready for review October 7, 2026 22:58

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants