Repository navigation
Old-gen garbage from untraced-promoted transient JSON trees is never reclaimed inside a parse/scan loop (8m:scan peak RSS 1.5–1.7× Node) #10182
Description
Activity
Measured the candidate design (branch
gc/promoted-cohort-bound, local, not pushed): old-reclaim due whenpromoted_since_full >= max(64 MB, 2 × old_live_at_last_full), credited incredit_promoted_bytes_to_old_baseline, reset infinish_full_old_reclaim_baseline, with the[gc-trigger]diag extended bypromoted_since_full=/cohort_bound=. The arm is live (counter 48 → 84 MB, then a full and a reset with the bound re-scaled to 2 × verified live), but it does not lower peak RSS on the target rows:cell main peak with bound why records_array_8m:scan(10 parses)188 MiB 200 MiB the full fires only after the 3rd promotion (84 MB > 64 MB), at the alloc-point arm near the end of the run; the peak was already set records_array_20m:parse256 276 same; one late full records_object_8m:parse187 210 same Two things the diag makes explicit:
- The trigger timelines are byte-identical to
mainup to the last minor:nursery cap scale 1x → 2x → 4xbecause Eden is fully live at each minor, so from-space grows 12 → 24 → 36 MB while the trees are promoted in place anyway. The young generation's residency is a third of the peak on these cells, and the cap scaling exists to avoid re-copying survivors — pointless for a cohort that is promoted untraced. - Even with the bound tightened so a full runs every two iterations, the arithmetic is: old ≈ 23 live + 46 dead, young ≈ 36, source 8 → ~120 MiB (1.1× Node), at a full marking ~31 MB every other iteration ≈ +5 ms on a ~22 ms iteration — about +20 % CPU on a cell that is 6 % ahead. That trades compute for RSS, which is out.
So on the current collector this row is a real trade-off, not a pacing bug. What would move it without the trade: (a) not scaling the nursery cap for a cohort the policy already promotes in place (the young third of the peak), (b) a materially cheaper full mark for pointer-heavy JSON trees (#10169's design 3), and (c) returning swept old blocks to the OS at the full so a full actually lowers RSS rather than only freeing arena space. Leaving the branch unpushed; the counters and diag fields are worth keeping when (a)–(c) are attempted.
- The trigger timelines are byte-identical to
Also measured lever (a) from the previous comment — not growing the nursery cap on a cycle that promoted in place (
retune_nursery_cap_scaleskips the ×2 step whencollector.stats.in_place_promotion; local branchgc/nursery-cap-in-place, not pushed). The guard is live (nursery cap scalelines: 2 → 0 on8m:scan, 1 → 0 on20m:parseandobject_8m:parse) and changes nothing else: peak RSS 188 / 255 / 188 MiB and the minor counts are identical tomain.The diag says why the cap is not the binding quantity here:
JSON.parseallocates the whole tree inside one suppressed window, so the young generation goes from ~0 to one tree between safepoints and every minor sees exactly one tree, whatever the cap says (the cap only decides when a minor becomes due, and it is due the moment the parse returns). The peak is thereforeold-gen dead promoted trees + one live tree + source, and only reclaiming those dead trees (a full) or not promoting them can lower it — which is the CPU trade the previous comment priced.So for parse/scan loops over document-sized inputs the honest state is: CPU ahead of Node/Bun, RSS 1.1–1.7× on the current collector, and the two cheap pacing levers do not move it. Remaining candidates are the structural ones: a materially cheaper full mark for JSON trees, and returning swept old blocks to the OS at a full.
Profile of the regime the RSS rows want, and the mechanism that would give it to them.
The fulls-only regime is the best one on both axes. With the (rejected) blanket knob that births every stringify result in the arena,
records_array_20m:roundtripruns 61 fulls and 0 minors over 61 iterations, at 64 ms/iteration on a loaded host and 151 MiB peak — better than Node on CPU (81) and RSS (261) at once, because every tree dies in Eden and is swept there; nothing is ever promoted. #10177's regime (untraced in-place promotion, then fulls when old-gen grows) lands at 57 ms/iteration and 266 MiB: the same CPU, +115 MiB, all of it dead promoted trees waiting for a full.Where a full's time goes (
sample, 8 s over that run, symbolized; GC = 28 % of all samples, parse 38 %, stringify 15 %):leaf samples share of all arena::walk::ArenaObjectCursor::next_budgeted250 11.1 % oldgen::ArenaSweepObjectsState::reclaim_dead_object73 3.3 % array::element_shape::forget_element_shape+clear_element_shape102 4.5 % oldgen::IncrementalSweepState::step46 2.0 % trace::ValidPointerSetBuilder::step45 2.0 % addr_class::try_read_tracked_gc_header29 1.3 % gc_type_finalize_unmarked_payload25 1.1 % verify::OldToYoungRememberedRebuildState::step+remember_retained_old_to_young_slots48 2.1 % The mark itself (
mark_addr, the slot visitors) is not in the top 26 leaves. So "a cheaper full mark" was the wrong lever: a full is dominated by walking every object in the arena (the object cursor, used by the valid-pointer-set pre-pass and by the sweep) and by per-dead-object reclamation — and the element-shape side table is forgotten per dead array (every record'stags), which looks like a hash operation per dead array and is the one cheap item here (check a header bit before touching the table).Why parse/scan loops do not get this regime today. The trigger order is OldReclaim first, then the nursery. In the knob run OldReclaim was due at every safepoint only by accident (each 20 MB born-old result tripped the absolute 48 MB arm), so the full always pre-empted the minor and the trees never left Eden. A parse/scan loop has no born-old results, OldReclaim is never due, the nursery minor runs, and the in-place promotion (correctly: the tree is live at that minor) moves the tree to old-gen, where only a full can reclaim it — the shape this issue opened with.
Mechanism that would make the regime deliberate rather than accidental:
- A full promotes its Eden survivors in place when the young generation is essentially all live after the mark (the same 95 % test the copying minor uses): retag the in-use young blocks to old via the
arena/promote.rsmachinery the copying minor already uses (retag_young_for_in_place_promotion/finish_in_place_promotion), after the sweep has reclaimed the dead objects in them; clear the remembered set (exact, no young left); credit the baseline; feed the survival measurement. Today a full is non-moving and promotes nothing (gc/mod.rs, perf: json_pipeline at 500k records is 97.6x bun (60.4s vs 618ms) while the same workload at 100 records BEATS bun — a scaling cliff, not a constant factor #7592's latch comment), which is why alternating a full with a minor ends in an evacuating minor that copies the whole tree (the failure the cohort-bound experiment above ran into). - The promoted-cohort bound from the earlier comment, with a floor of one nursery rather than 64 MB. With (1) in place the steady state of a parse/scan loop is one full per iteration and no minors: mark T_i (cheap), sweep and reclaim T_{i−1}, promote T_i in place — the knob run's cost structure, ~18 ms per iteration at 20 MB. Expected:
20m parse≈ 54 ms/iteration (vs the 208 ms/4 = 52 ms best, i.e. parity within noise) at ~150 MiB (0.7× Node);8m scan≈ 24 ms/iteration (vs 189/8 = 23.6 best) at ~90–110 MiB (parity);roundtripas measured above. - The sweep-side items in the table are the CPU headroom that makes (2) free instead of marginal: skip
forget_element_shapefor arrays that never carried a proof, and look at whether the valid-pointer-set pre-pass and the sweep can share one object walk.
This changes gated gc-ratchet counters on the retain probes (fulls that promote), so it needs a baseline regeneration on the quiet mini alongside the corpus run, and the cc rows on perrymaster. Not started; the two branches from the earlier comments (
gc/promoted-cohort-bound,gc/nursery-cap-in-place) are local and contain the counters/diag fields (2) reuses.- A full promotes its Eden survivors in place when the young generation is essentially all live after the mark (the same 95 % test the copying minor uses): retag the in-use young blocks to old via the
Two measurements that re-scope this issue after #10204 (full promotes Eden survivors; draft, does not meet its bar: CPU roughly doubles on the 20 MB parse/scan/sparse rows and
8m:scan).1. Perry's tree is not bigger than V8's. Heap census (
PERRY_GC_CENSUS,gc()after one retainedJSON.parseofrecords_array_8m.json, 59,000 records, currentmain), minus the 7.1 MB source string that is live only in the tree run:type live bytes objects bytes/record string ( name,email;tagsentries are SSO)6.13 MB 121,358 104 object (records) 4.25 MB 59,000 72 array ( tags+ the records array)2.48 MB 59,009 42 tree ≈ 12.4 MiB 239k ≈ 210 Node 26.5.1, same document,
--expose-gcheap delta: 13.0 MiB (231 B/record). So the RSS gap on the parse/scan rows is not representation size; it is how many dead trees are resident. On8m:scanold gen ends at 105 MB ≈ 8 trees: every tree parsed is still there, because no full runs.2. A full is expensive because the sweep is per object, not because of the mark.
ArenaSweepObjectsStatewalks every object in the arena withArenaObjectCursor; each dead object getsold_page_account_swept_object,finalize_dead_arena_payload(element-shape forget, layout-mask and overflow side tables), and a deferred page-index unregister; blocks with no live object are reset only after that walk (block_has_live). After a parse/scan loop nearly every old block holds only dead objects, so a full's cost scales with dead objects (#10204's head: ~23 ms per full on8m:scanfor a ~150 MB arena with one or two live 12 MB trees), which is what turns "reclaim every few iterations" into a CPU regression.The lever that fits the CPU headroom: block-granular reclamation. Record during the mark which blocks received a marked object (a per-block bit or count set where
GC_FLAG_MARKEDis set). In the sweep, a block with no marked object is reclaimed without visiting its objects, provided its dead population needs no per-object finalization: page-index entries dropped by block range, address-keyed side tables purged by range (or proven empty for the block), then the existing block reset. Blocks with any live object keep today's per-object path. With that, a full that runs every 2–3 iterations costs roughly the mark of the live tree (~240k objects on 8 MB, ~1.1M on 20 MB) plus O(blocks), i.e. about 1–3 ms per iteration amortized — inside the headroom of8m:scan(~1.5 ms/iteration) and the 20 MB rows (~15 ms/iteration) — and old gen holds at most 2–3 trees, which is the RSS parity point on both.Per-object obligations that must stay exact under a block skip, each to be settled before relying on it: finalizer hook types (
GcFinalizeHookKind: Map/Set side allocations, promises, native arena owners/views — a block containing any such object keeps the per-object path), element-shape records (array/element_shape.rs), layout slot masks / typed layouts, overflow fields and closure dynamic props, weak refs and finalization registries, the remembered set and dirty-page coverage, the old-page free-hole list (#7437), forwarding stubs (#6228), and pinned objects.Checked one more premise before the next step, with no build:
PERRY_GC_FORCE_EVACUATE=1on #10220's head worker (it disables in-place and untraced promotion, so every minor evacuates and measures).cell default CPU / peak forced evacuation CPU / peak survival at every minor records_array_8m:scan178 ms / 187 MiB (3 minors, 2 untraced) 504 ms / 284 MiB (8 minors) 998–1000 ‰ records_array_20m:parse141 ms / 241 MiB (2 minors, 1 untraced) 566 ms / 417 MiB (3 minors) 999–1000 ‰ So the untraced promotions are not acting on a stale reading: at minor time the young generation really is all live (the previous tree is still held by
lastwhile the current one is being scanned), and each promoted tree dies only after the minor. #10204's 333–500 ‰ was measured at the fulls it inserted, not at these minors. Evacuating instead copies live trees at ~72 ns/object and doubles the footprint — ruled out.Consequences for the remaining rows:
8m:scan(189 vs 110 MiB, 13.7 ms CPU lead over 8 iterations): any collection that reclaims its dead trees must trace the live 240k-object tree; at gc: make a synchronous full cheaper per live object (#10182) [pacing retry does not meet acceptance] #10220's ~42 ns/object mark that is ~10 ms per collection with every other phase at zero, and reaching 110 MiB needs about three. It does not fit the lead without roughly halving the per-object mark again.- 20 MB parse/scan/sparse and
object_20m:parse(240 vs 220–225 MiB, ~64 ms lead):k=1pacing already reaches 179 MiB with two fulls; the mark of the 585k-object tree is ~25 ms each, but a pacing full costs 62–68 ms today: census over the dead promoted trees (27 %), page-run re-materialization for blocks the sweep then releases whole (15 %), and the forced conservative stack scan because pacing fulls fire at an allocation point. Bringing a pacing full down to roughly its mark puts these four rows inside the lead. That is the next task, ongc/full-throughput. object_8m:parse(118 vs 112): different regime (5 evacuating minors, no fulls); needs its own look.small_record:parse(80 vs 60): the fix(gc): price the tiny-parse pressure guard by the productivity backoff (#9831) #9838 tiny-parse floor, a cc-rig decision.
- added 10 commits that reference this issue
on Sep 13, 2026 40 remaining items
- added 14 commits that reference this issue
on Sep 14, 2026 - added a commit that references this issue
on Sep 27, 2026
Symptom
benchmarks/json_performancerecords_array_8m:scan(parse 8 MB, iterate every record, drop; 10 parses) peaks at 188 MiB RSS against Node/Bun's 110 on currentmain(1.7×; 163 MiB = 1.48× on the combined parity branch).PERRY_GC_DIAG=1over the whole run:untraced=truein-place promotions (the cheap path perf(gc): promote a fully-live young generation without tracing it — retain −33.6%, deeplist −43% #7888 built for exactly this: a fully-live young generation);old_in_use=105 MB,arena_total=146 MB— four to five 23 MB trees in old gen of which one is live.Same shape on
records_array_20mparse/scan/sparse (peak 256 vs 220–225, 1.14×) and onrecords_object_8m:parse.Cause
Each iteration's tree is promoted whole and dies one iteration later, in old gen. Nothing then reclaims it:
credit_promoted_bytes_to_old_baselineadds every promoted byte toGC_LAST_OLD_RECLAIM_IN_USE_BYTES(perf(gc): main regressed the retain cluster 2.2-4.8x — retain now runs 2 full collections where it ran none (suspect #7901/#7902) #7965, deliberately — a pinned baseline degenerates the proportional band into a quadratic constant-band pacer onretain). Soold_in_use − baselinestays ≈ 0 and the proportional arm ofold_reclaim_pressure_duenever fires.old_in_use ≥ 48 MB && baseline < 48 MB) is exempted whileGC_MAJOR_PACING_RETAININGis set, and a young generation that survives at 999 ‰ sets it on every minor.untraced_promotion_budget_bytes(128 MB floor) forces a measuring minor, which measures the current young generation (again ~100 % live, so no contradiction);implied_dead_byteschargespromoted × (1000 − 980)/1000= 2 % against the 32 MBPROMOTED_DEAD_BUDGET_BYTES, i.e. a full after 1.6 GB of promotions;request_old_reclaim_for_untraced_promotionsneeds a contradicting measurement.So the predictor is right about survival (the tree is live at the minor) and wrong about lifetime (it dies right after), and no instrument observes lifetime. The garbage is reclaimed only when arena-growth escalation trips.
Candidate design (GC policy; gc-ratchet corpus before/after, then the cc rows on perrymaster, and both CPU and peak RSS on the JSON matrix — never trade the CPU lead for RSS)
Bound the unverified old-gen cohort: track
promoted_since_last_full(bytes promoted by untraced or in-place promotions since the last full) and make old-reclaim due when it exceedsmax(floor, k × old_live_at_last_full)with, e.g., floor 64 MB and k = 2. Rationale: the cohort's liveness is assumed, not measured; a full re-establishes it at a cost proportional to the verified live set, so bounding the unverified part by a multiple of the verified part keeps total major work linear (the #7592 argument) while capping residency.Expected on the rows (from the diag numbers):
8m:scanlive after a full ≈ 31 MB → a full every ~2.8 iterations → peak ≈ 31 + 64 + 23 ≈ 118 MiB (Node 110) at roughly +5 ms per iteration on a cell that is currently 6 % ahead on CPU.20mfamily: live ≈ 78 MB → band 156 MB → a full every ~2.7 iterations, CPU still ahead of the 208 ms best, RSS roughly unchanged — the 20 MB rows cannot reach RSS parity this way (one dead 58 MB tree resident is already the gap), only a materially cheaper full mark can (a JSON tree marks at ~47 ns/object today; see #10169's design 3).Not in scope, recorded here so it is not re-diagnosed:
small_record:parsepeaks at 80 vs 60 MiB because minors fire every ~50 MB — the tiny-parse guard's 48 MB floor from #9831/#9838, which exists to stop the cc minor storms; lowering it is a cc-rig decision, not a JSON-row one.Reproduce
Related: #10169 / #10177 (the stringify-result half of the same picture), #10123, #7965, #7902, #7888.