Skip to content

All rendering optimizations, measured again on top of the scheduler - #21656

Draft
NullVoxPopuli-ai-agent wants to merge 12 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/all-optimizations-v2-final
Draft

NullVoxPopuli-ai-agent wants to merge 12 commits into
emberjs:mainfrom
NullVoxPopuli-ai-agent:nvp/all-optimizations-v2-final

Conversation

@NullVoxPopuli-ai-agent

Copy link
Copy Markdown
Contributor

All the rendering optimizations that won in a fresh measurement, on top of the scheduler (#21655). Up to 61b679903a, this is 57% faster than main on the rere-benchmark ember apps (geometric mean of 15 benches), and every bench is faster. The last commit is a breaking change that adds 5%.

Supersedes #21520. That spike was tuned with unpinned, unprofiled benches. Here, every candidate was measured again, one at a time, and only the winners stayed.

Commits

Commit Change Measured
f3ce59665d, eb9b76b539 #21655: the RFC 957 scheduler, without the runloop, RSVP, or backburner -47% vs main
0c057ed068 Notify the renderer once per tick -12% vs the scheduler
7732ae26d5 .. 6e08894792 #21650: pooled trackers, reused combined tags, lazy TrackedValue functions -11%
8315e8f1c3 One cell per tracked field (value and tag together) -10%
d6f42cb25c Update a list in place when its keys keep their order -11% (screening)
61b679903a Flatten nested combined tags -4% (screening)
62e7898dbf Breaking: template reads skip the legacy tags (POJO set(), [], unknownProperty) -5%

Measured and dropped:

Change Result
Skip unchanged {{#each}} item subtrees (#21544) +10%: the per-item tracking frame costs more than skipping small item templates saves
Lock path references onto their field's tag (spike 2b703acdc3) +4%
Updating VM without a frame object per block (spike 5530b80746) +15% on top of the list path

Final result

Mirrored round: main, F1, F2, F2, F1, main. F1 is 61b679903a, and F2 is 62e7898dbf. The two main runs differ by 0.9%.

Bench main F1 change
1 item, 100k updates (async) 2193 100 -95.4%
1k items, 1 update each (sequentially, async) 1658 129 -92.2%
1k items 1 update on 25% (random, async) 613 97.8 -84.1%
1 item, 1k updates (async) 164 30.2 -81.6%
1k items 1 update on 5% (random, async) 260 77.7 -70.1%
1 value, 1k consumers, 10k updates (bursts of 100) 2175 706 -67.5%
1 item, 100k updates 77.7 59.7 -23.1%
1 value, 1k consumers, 10k updates (bursts of 1000) 383 314 -18.2%
1 value, 1k consumers, 10k updates (single burst) 129 115 -10.7%
1k items, 1 update each (sequentially) 133 119 -10.1%
1k items 1 update on 5% (random) 88.1 79.4 -9.9%
1 item, 1k updates 24.6 23.4 -5.1%
1k items 1 update on 25% (random) 101 95.8 -4.9%
Incrementing Render Effect 4510 4460 -1.1%
DB Monitor w/ chat simulation (fps) 23.2 33.5 +44% fps

Times are medians in ms over two runs each, and negative is faster. F1 is -57.3% and F2 is -59.3% (geometric mean).

The breaking commit helps Incrementing Render Effect (-13%), DB Monitor (+9% fps), and 5% (random) (-17%), and makes fan-out bursts of 100 15% slower.

Method

  • rere-benchmark ember apps, headless, 8x CPU throttle, 25 samples per bench, Chrome pinned to one CPU, every sample profiled (rere-benchmark#132). Profiling slows samples by about 1.6x, so these times are only comparable with each other.
  • The VM's host changes speed: at times the whole VM rendered about 60% slower, with no steal time. So every run starts only after a short calibration bench (fan-out, 3 samples, the same npm build every time) is within 15% of the round's reference, and the load is below 4. Runs from the slow period are thrown away.
  • Levers were measured in a chain (each step against the two runs before it) and then screened again one at a time on the base (no throttle, 3 samples, 3 interleaved rounds). Only the final two builds got the long mirrored round above.

Where the time goes now

From the profiles of F1:

  • Work outside JavaScript (DOM, style, layout): 57% of DB Monitor and 59% of fan-out bursts of 100.
  • @tracked writes: dirtyTagFor (17%) and the tracked set (13%) on 1 item, 100k updates. The one-cell change covers trackedData, but the decorator still goes through the central tag registry.
  • VM dispatch (evaluate, execute, nextStatement): 15% to 25% of the list and fan-out benches.
  • queueMicrotask: 7% of Incrementing Render Effect, from the scheduler.
  • Tracking frames (COMPUTE): 4% to 8%.

Machine

Item Value
VM QEMU Q35 with KVM
CPU 8 vCPUs of an AMD Ryzen 9 7900X, 1 thread per core
Memory 22.9 GiB RAM, 19 GiB swap
GPU virtio GPU, no WebGL
OS Pop!_OS 24.04, kernel 7.0.11
Software Chrome 152.0.7977.64, Node 26.10.0

Known gaps

🤖 Generated with Claude Code

NullVoxPopuli and others added 11 commits October 3, 2026 02:57
Adds the `@ember/scheduler` package proposed by RFC 0957:

- `render`, `layout`, `composite`, `next` and `idle` phase functions,
  each returning a promise that resolves according to the registered
  scheduling strategy
- `registerStrategy`, for providing the scheduling strategy when defining
  the Application
- `@ember/scheduler/strategy`, the default strategy implementation, which
  flushes the render/layout/composite phases in order via ordered
  requestAnimationFrame callbacks within a single frame, prior to paint

The deprecations of @ember/runloop and RSVP described by the RFC are left
to follow-up work; this is the additive API surface.

Fix Safari flake and idle() starvation in the default strategy

The "composite while composite is flushing" test asserted ordering across
two independent channels: a setTimeout task scheduled during frame 1
versus frame 2's requestAnimationFrame callbacks. The HTML spec does not
order pending timer tasks against the next rendering opportunity, and
Safari 15.6 runs the next frame's rAF callbacks first. The test now
anchors entirely to the rAF channel, using a raw requestAnimationFrame
registered ahead of the rescheduled phase windows as the frame-2
boundary.

idle() also armed requestIdleCallback without a timeout; fully-idle or
backgrounded pages can starve rIC indefinitely, leaving the promise
unresolvable. Cap the wait with { timeout: 500 }.
This is the RFC 957 end state on top of the scheduler interface
(the previous commit, from emberjs#21552). Tag invalidation tells the
renderer's scheduler directly. Nothing in the framework schedules
through @ember/runloop, and backburner and RSVP are gone.

- The default @ember/scheduler strategy is the renderer's clock: one
  render per tick, with bounded settle rounds, and a microtask,
  animation-frame, or task leg chosen by a microtask-window tick
  classifier.
- @ember/runloop is a compatibility layer without dependencies: run,
  join, and bind are plain calls, queues become microtasks, and
  timers are native timers. Its public types are the same as on main.
- Native promises replace RSVP in the router, the route managers,
  PromiseProxyMixin, and the tests.
- Render settledness is reported as edges (isRenderPending on
  @ember/renderer), for test waiters.
- Destroys drain synchronously when an engine instance is destroyed.

This commit holds only the scheduler parts of the closed spike emberjs#21520.
Its VM and allocation changes, and the "notify once per tick" latch,
are left out, so that each one can be measured on its own.

Known gap: the browser suite stops early. Tests that set state in
runTask and then assert the DOM in the same task expect a synchronous
render. The first ones are in the debug render tree tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A set walked the notify chain every time: the renderer callback and
the settledness sample. A 100k-set loop now notifies once and then
checks one boolean per set, until the end of the tick re-arms it.

The flag only latches when a renderer heard the notification, so
dirt during app boot cannot swallow later invalidations. Registering
a renderer re-arms it too.

From the spike emberjs#21520 (c93641e, d6be883).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A tracking frame allocated one tracker, one Set, and one array from
that Set. In CPU profiles of three createCache graphs, the tracker, its
Set and beginTrackFrame took 27% to 45% of the samples.

Frames are strictly nested, so one tracker for each depth is enough.
beginTrackFrame takes the tracker of its depth from a pool.

A tracker keeps its tags in an array. Each frame has a number, and the
tracker writes that number on each tag that it takes, so a tag that the
frame consumes again costs one comparison and no Set.

endTrackFrame now takes the tag that the same frame produced the last
time. If the frame consumed the same tags again, that tag is the
result, with its memoized revision and no allocation. getValue passes
the tag of its cache.

The loop over the subtags of a combined tag uses an index in place of
an iterator.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A TrackedValue made four arrow functions and one options object for
each instance. A value that the code reads and writes through `value`
did not use any of them.

`get`, `set`, `update` and `freeze` are now accessors that make the
bound function on its first read, and keep it. They are bound as
before, and each read gives the same function.

`trackedValue(value)` with no options shares one options object.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The tracker wrote the number of its frame on each tag. That counter
leaves the small-integer range of V8 after about one billion frames,
and the benchmark is then 1.1 to 1.3 times slower. A nested frame also
wrote its own number on a shared tag, so the outer frame took that tag
again each time: a computed that reads five computeds of one tag had a
combined tag with five entries.

A tag now keeps the index at which a tracker took it. A tracker has the
tag if its entry at that index is the tag. This needs no counter, and a
stale index cannot hide a tag, because the entry is then another tag or
no tag. A nested frame that takes the tag at the same index keeps the
check true, so the five computeds give one tag again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
`get`, `set`, `update` and `freeze` were own properties before, so
code could assign to them. Each accessor now has a setter. A write
through `value` uses an assigned `set`, as it did before.

The shared default options are frozen.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A tracked read went through three map hops: the central tag
registry (a WeakMap to a Map per object) and a separate values
WeakMap. Now each field has one WeakMap from the instance to a cell
with its value and tag, so a read is one hop and consumeTag, and a
write is one hop and DIRTY_TAG.

The cell registers its tag in the central registry when it is made,
so notifyPropertyChange and computed chains dirty and read the same
tag as the field.

From the spike emberjs#21520 (c93641e, d6be883).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a fresh iterator yields the same keys in the same order, as a
derived array in a getter does on every render, the list block now
updates each item's refs in place. It skips the diff bookkeeping, the
marker DOM, and rebuilding the children. On the first mismatch, a
PrefixedIterator replays the consumed items through the full sync.

From the spike emberjs#21520 (the list part of d9f1250).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
combine() now flattens nested combinators and drops constant tags, up
to 64 sub-tags, so validating a combined tag is one flat loop instead
of a tree walk.

A combinator keeps the tags it was made from in inputs. The tracker
from emberjs#21650 compares a frame's tags with those to reuse the previous
combinator, which the flattened list would no longer match.

From the spike emberjs#21520 (the combine part of d9f1250). The inputs
field is new.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every template property read consumed a tag for (object, key), so a
set() on a plain object re-renders. Array values also consumed the
'[]' tag, and each read checked for unknownProperty. This removes all
three: a plain-data read no longer joins the reactivity graph, so
only @Tracked fields, tracked collections, and replacing the object
trigger a re-render.

This breaks apps that set() plain objects shown in templates, rely
on EmberArray '[]' invalidation, or render ObjectProxy content. It
needs an RFC and a deprecation before it can land. It is in this
branch to measure the ceiling.

From the spike emberjs#21520 (2607466, 908f68b).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An untrack frame takes a depth but no tracker. A track frame that begins
inside one, at a depth that never had a tracker, leaves a hole below it.
resetTracking then called clear() on undefined. The QUnit setup calls
resetTracking after each test, so a filtered run of the tracking module
stopped at the first test after the hole. The full suite fills every
depth before that, so it did not show the error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants