All rendering optimizations, measured again on top of the scheduler - #21656
Draft
NullVoxPopuli-ai-agent wants to merge 12 commits into
Draft
NullVoxPopuli-ai-agent wants to merge 12 commits into
NullVoxPopuli-ai-agent wants to merge 12 commits into
Conversation
Adds the `@ember/scheduler` package proposed by RFC 0957:
- `render`, `layout`, `composite`, `next` and `idle` phase functions,
each returning a promise that resolves according to the registered
scheduling strategy
- `registerStrategy`, for providing the scheduling strategy when defining
the Application
- `@ember/scheduler/strategy`, the default strategy implementation, which
flushes the render/layout/composite phases in order via ordered
requestAnimationFrame callbacks within a single frame, prior to paint
The deprecations of @ember/runloop and RSVP described by the RFC are left
to follow-up work; this is the additive API surface.
Fix Safari flake and idle() starvation in the default strategy
The "composite while composite is flushing" test asserted ordering across
two independent channels: a setTimeout task scheduled during frame 1
versus frame 2's requestAnimationFrame callbacks. The HTML spec does not
order pending timer tasks against the next rendering opportunity, and
Safari 15.6 runs the next frame's rAF callbacks first. The test now
anchors entirely to the rAF channel, using a raw requestAnimationFrame
registered ahead of the rescheduled phase windows as the frame-2
boundary.
idle() also armed requestIdleCallback without a timeout; fully-idle or
backgrounded pages can starve rIC indefinitely, leaving the promise
unresolvable. Cap the wait with { timeout: 500 }.
This is the RFC 957 end state on top of the scheduler interface (the previous commit, from emberjs#21552). Tag invalidation tells the renderer's scheduler directly. Nothing in the framework schedules through @ember/runloop, and backburner and RSVP are gone. - The default @ember/scheduler strategy is the renderer's clock: one render per tick, with bounded settle rounds, and a microtask, animation-frame, or task leg chosen by a microtask-window tick classifier. - @ember/runloop is a compatibility layer without dependencies: run, join, and bind are plain calls, queues become microtasks, and timers are native timers. Its public types are the same as on main. - Native promises replace RSVP in the router, the route managers, PromiseProxyMixin, and the tests. - Render settledness is reported as edges (isRenderPending on @ember/renderer), for test waiters. - Destroys drain synchronously when an engine instance is destroyed. This commit holds only the scheduler parts of the closed spike emberjs#21520. Its VM and allocation changes, and the "notify once per tick" latch, are left out, so that each one can be measured on its own. Known gap: the browser suite stops early. Tests that set state in runTask and then assert the DOM in the same task expect a synchronous render. The first ones are in the debug render tree tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A set walked the notify chain every time: the renderer callback and the settledness sample. A 100k-set loop now notifies once and then checks one boolean per set, until the end of the tick re-arms it. The flag only latches when a renderer heard the notification, so dirt during app boot cannot swallow later invalidations. Registering a renderer re-arms it too. From the spike emberjs#21520 (c93641e, d6be883). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A tracking frame allocated one tracker, one Set, and one array from that Set. In CPU profiles of three createCache graphs, the tracker, its Set and beginTrackFrame took 27% to 45% of the samples. Frames are strictly nested, so one tracker for each depth is enough. beginTrackFrame takes the tracker of its depth from a pool. A tracker keeps its tags in an array. Each frame has a number, and the tracker writes that number on each tag that it takes, so a tag that the frame consumes again costs one comparison and no Set. endTrackFrame now takes the tag that the same frame produced the last time. If the frame consumed the same tags again, that tag is the result, with its memoized revision and no allocation. getValue passes the tag of its cache. The loop over the subtags of a combined tag uses an index in place of an iterator. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A TrackedValue made four arrow functions and one options object for each instance. A value that the code reads and writes through `value` did not use any of them. `get`, `set`, `update` and `freeze` are now accessors that make the bound function on its first read, and keep it. They are bound as before, and each read gives the same function. `trackedValue(value)` with no options shares one options object. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tracker wrote the number of its frame on each tag. That counter leaves the small-integer range of V8 after about one billion frames, and the benchmark is then 1.1 to 1.3 times slower. A nested frame also wrote its own number on a shared tag, so the outer frame took that tag again each time: a computed that reads five computeds of one tag had a combined tag with five entries. A tag now keeps the index at which a tracker took it. A tracker has the tag if its entry at that index is the tag. This needs no counter, and a stale index cannot hide a tag, because the entry is then another tag or no tag. A nested frame that takes the tag at the same index keeps the check true, so the five computeds give one tag again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`get`, `set`, `update` and `freeze` were own properties before, so code could assign to them. Each accessor now has a setter. A write through `value` uses an assigned `set`, as it did before. The shared default options are frozen. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A tracked read went through three map hops: the central tag registry (a WeakMap to a Map per object) and a separate values WeakMap. Now each field has one WeakMap from the instance to a cell with its value and tag, so a read is one hop and consumeTag, and a write is one hop and DIRTY_TAG. The cell registers its tag in the central registry when it is made, so notifyPropertyChange and computed chains dirty and read the same tag as the field. From the spike emberjs#21520 (c93641e, d6be883). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a fresh iterator yields the same keys in the same order, as a derived array in a getter does on every render, the list block now updates each item's refs in place. It skips the diff bookkeeping, the marker DOM, and rebuilding the children. On the first mismatch, a PrefixedIterator replays the consumed items through the full sync. From the spike emberjs#21520 (the list part of d9f1250). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
combine() now flattens nested combinators and drops constant tags, up to 64 sub-tags, so validating a combined tag is one flat loop instead of a tree walk. A combinator keeps the tags it was made from in inputs. The tracker from emberjs#21650 compares a frame's tags with those to reuse the previous combinator, which the flattened list would no longer match. From the spike emberjs#21520 (the combine part of d9f1250). The inputs field is new. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every template property read consumed a tag for (object, key), so a set() on a plain object re-renders. Array values also consumed the '[]' tag, and each read checked for unknownProperty. This removes all three: a plain-data read no longer joins the reactivity graph, so only @Tracked fields, tracked collections, and replacing the object trigger a re-render. This breaks apps that set() plain objects shown in templates, rely on EmberArray '[]' invalidation, or render ObjectProxy content. It needs an RFC and a deprecation before it can land. It is in this branch to measure the ceiling. From the spike emberjs#21520 (2607466, 908f68b). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
NullVoxPopuli
added a commit
to NullVoxPopuli/rere-benchmark
that referenced
this pull request
Oct 4, 2026
An untrack frame takes a depth but no tracker. A track frame that begins inside one, at a depth that never had a tracker, leaves a hole below it. resetTracking then called clear() on undefined. The QUnit setup calls resetTracking after each test, so a filtered run of the tracking module stopped at the first test after the hole. The full suite fills every depth before that, so it did not show the error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
All the rendering optimizations that won in a fresh measurement, on top of the scheduler (#21655). Up to
61b679903a, this is 57% faster thanmainon the rere-benchmark ember apps (geometric mean of 15 benches), and every bench is faster. The last commit is a breaking change that adds 5%.Supersedes #21520. That spike was tuned with unpinned, unprofiled benches. Here, every candidate was measured again, one at a time, and only the winners stayed.
Commits
f3ce59665d,eb9b76b539main0c057ed0687732ae26d5..6e08894792TrackedValuefunctions8315e8f1c3d6f42cb25c61b679903a62e7898dbfset(),[],unknownProperty)Measured and dropped:
{{#each}}item subtrees (#21544)2b703acdc3)5530b80746)Final result
Mirrored round:
main, F1, F2, F2, F1,main. F1 is61b679903a, and F2 is62e7898dbf. The twomainruns differ by 0.9%.mainTimes are medians in ms over two runs each, and negative is faster. F1 is -57.3% and F2 is -59.3% (geometric mean).
The breaking commit helps Incrementing Render Effect (-13%), DB Monitor (+9% fps), and
5% (random)(-17%), and makes fan-out bursts of 100 15% slower.Method
Where the time goes now
From the profiles of F1:
@trackedwrites:dirtyTagFor(17%) and the trackedset(13%) on1 item, 100k updates. The one-cell change coverstrackedData, but the decorator still goes through the central tag registry.evaluate,execute,nextStatement): 15% to 25% of the list and fan-out benches.queueMicrotask: 7% of Incrementing Render Effect, from the scheduler.COMPUTE): 4% to 8%.Machine
Known gaps
runTaskexpect a synchronous render.🤖 Generated with Claude Code