Skip to content

Old-gen garbage from untraced-promoted transient JSON trees is never reclaimed inside a parse/scan loop (8m:scan peak RSS 1.5–1.7× Node) #10182

Description

@proggeramlug

Symptom

benchmarks/json_performance records_array_8m:scan (parse 8 MB, iterate every record, drop; 10 parses) peaks at 188 MiB RSS against Node/Bun's 110 on current main (1.7×; 163 MiB = 1.48× on the combined parity branch). PERRY_GC_DIAG=1 over the whole run:

Same shape on records_array_20m parse/scan/sparse (peak 256 vs 220–225, 1.14×) and on records_object_8m:parse.

Cause

Each iteration's tree is promoted whole and dies one iteration later, in old gen. Nothing then reclaims it:

  • credit_promoted_bytes_to_old_baseline adds every promoted byte to GC_LAST_OLD_RECLAIM_IN_USE_BYTES (perf(gc): main regressed the retain cluster 2.2-4.8x — retain now runs 2 full collections where it ran none (suspect #7901/#7902) #7965, deliberately — a pinned baseline degenerates the proportional band into a quadratic constant-band pacer on retain). So old_in_use − baseline stays ≈ 0 and the proportional arm of old_reclaim_pressure_due never fires.
  • The absolute first-crossing arm (old_in_use ≥ 48 MB && baseline < 48 MB) is exempted while GC_MAJOR_PACING_RETAINING is set, and a young generation that survives at 999 ‰ sets it on every minor.
  • The three instruments that bound the untraced cohort act on evidence this workload never produces: untraced_promotion_budget_bytes (128 MB floor) forces a measuring minor, which measures the current young generation (again ~100 % live, so no contradiction); implied_dead_bytes charges promoted × (1000 − 980)/1000 = 2 % against the 32 MB PROMOTED_DEAD_BUDGET_BYTES, i.e. a full after 1.6 GB of promotions; request_old_reclaim_for_untraced_promotions needs a contradicting measurement.

So the predictor is right about survival (the tree is live at the minor) and wrong about lifetime (it dies right after), and no instrument observes lifetime. The garbage is reclaimed only when arena-growth escalation trips.

Candidate design (GC policy; gc-ratchet corpus before/after, then the cc rows on perrymaster, and both CPU and peak RSS on the JSON matrix — never trade the CPU lead for RSS)

Bound the unverified old-gen cohort: track promoted_since_last_full (bytes promoted by untraced or in-place promotions since the last full) and make old-reclaim due when it exceeds max(floor, k × old_live_at_last_full) with, e.g., floor 64 MB and k = 2. Rationale: the cohort's liveness is assumed, not measured; a full re-establishes it at a cost proportional to the verified live set, so bounding the unverified part by a multiple of the verified part keeps total major work linear (the #7592 argument) while capping residency.

Expected on the rows (from the diag numbers): 8m:scan live after a full ≈ 31 MB → a full every ~2.8 iterations → peak ≈ 31 + 64 + 23 ≈ 118 MiB (Node 110) at roughly +5 ms per iteration on a cell that is currently 6 % ahead on CPU. 20m family: live ≈ 78 MB → band 156 MB → a full every ~2.7 iterations, CPU still ahead of the 208 ms best, RSS roughly unchanged — the 20 MB rows cannot reach RSS parity this way (one dead 58 MB tree resident is already the gap), only a materially cheaper full mark can (a JSON tree marks at ~47 ns/object today; see #10169's design 3).

Not in scope, recorded here so it is not re-diagnosed: small_record:parse peaks at 80 vs 60 MiB because minors fire every ~50 MB — the tiny-parse guard's 48 MB floor from #9831/#9838, which exists to stop the cc minor storms; lowering it is a cc-rig decision, not a JSON-row one.

Reproduce

env PERRY_GC_DIAG=1 /usr/bin/time -l <worker> records_array_8m.json scan 8 2 2> diag.txt
grep -cE "ran pause_us" diag.txt; grep -c "kind=old_reclaim" diag.txt; grep -oE "old_in_use=[0-9]+" diag.txt | tail -1

Related: #10169 / #10177 (the stringify-result half of the same picture), #10123, #7965, #7902, #7888.

Activity

  1. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Measured the candidate design (branch gc/promoted-cohort-bound, local, not pushed): old-reclaim due when promoted_since_full >= max(64 MB, 2 × old_live_at_last_full), credited in credit_promoted_bytes_to_old_baseline, reset in finish_full_old_reclaim_baseline, with the [gc-trigger] diag extended by promoted_since_full=/cohort_bound=. The arm is live (counter 48 → 84 MB, then a full and a reset with the bound re-scaled to 2 × verified live), but it does not lower peak RSS on the target rows:

    cell main peak with bound why
    records_array_8m:scan (10 parses) 188 MiB 200 MiB the full fires only after the 3rd promotion (84 MB > 64 MB), at the alloc-point arm near the end of the run; the peak was already set
    records_array_20m:parse 256 276 same; one late full
    records_object_8m:parse 187 210 same

    Two things the diag makes explicit:

    1. The trigger timelines are byte-identical to main up to the last minor: nursery cap scale 1x → 2x → 4x because Eden is fully live at each minor, so from-space grows 12 → 24 → 36 MB while the trees are promoted in place anyway. The young generation's residency is a third of the peak on these cells, and the cap scaling exists to avoid re-copying survivors — pointless for a cohort that is promoted untraced.
    2. Even with the bound tightened so a full runs every two iterations, the arithmetic is: old ≈ 23 live + 46 dead, young ≈ 36, source 8 → ~120 MiB (1.1× Node), at a full marking ~31 MB every other iteration ≈ +5 ms on a ~22 ms iteration — about +20 % CPU on a cell that is 6 % ahead. That trades compute for RSS, which is out.

    So on the current collector this row is a real trade-off, not a pacing bug. What would move it without the trade: (a) not scaling the nursery cap for a cohort the policy already promotes in place (the young third of the peak), (b) a materially cheaper full mark for pointer-heavy JSON trees (#10169's design 3), and (c) returning swept old blocks to the OS at the full so a full actually lowers RSS rather than only freeing arena space. Leaving the branch unpushed; the counters and diag fields are worth keeping when (a)–(c) are attempted.

  2. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Also measured lever (a) from the previous comment — not growing the nursery cap on a cycle that promoted in place (retune_nursery_cap_scale skips the ×2 step when collector.stats.in_place_promotion; local branch gc/nursery-cap-in-place, not pushed). The guard is live (nursery cap scale lines: 2 → 0 on 8m:scan, 1 → 0 on 20m:parse and object_8m:parse) and changes nothing else: peak RSS 188 / 255 / 188 MiB and the minor counts are identical to main.

    The diag says why the cap is not the binding quantity here: JSON.parse allocates the whole tree inside one suppressed window, so the young generation goes from ~0 to one tree between safepoints and every minor sees exactly one tree, whatever the cap says (the cap only decides when a minor becomes due, and it is due the moment the parse returns). The peak is therefore old-gen dead promoted trees + one live tree + source, and only reclaiming those dead trees (a full) or not promoting them can lower it — which is the CPU trade the previous comment priced.

    So for parse/scan loops over document-sized inputs the honest state is: CPU ahead of Node/Bun, RSS 1.1–1.7× on the current collector, and the two cheap pacing levers do not move it. Remaining candidates are the structural ones: a materially cheaper full mark for JSON trees, and returning swept old blocks to the OS at a full.

  3. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Profile of the regime the RSS rows want, and the mechanism that would give it to them.

    The fulls-only regime is the best one on both axes. With the (rejected) blanket knob that births every stringify result in the arena, records_array_20m:roundtrip runs 61 fulls and 0 minors over 61 iterations, at 64 ms/iteration on a loaded host and 151 MiB peak — better than Node on CPU (81) and RSS (261) at once, because every tree dies in Eden and is swept there; nothing is ever promoted. #10177's regime (untraced in-place promotion, then fulls when old-gen grows) lands at 57 ms/iteration and 266 MiB: the same CPU, +115 MiB, all of it dead promoted trees waiting for a full.

    Where a full's time goes (sample, 8 s over that run, symbolized; GC = 28 % of all samples, parse 38 %, stringify 15 %):

    leaf samples share of all
    arena::walk::ArenaObjectCursor::next_budgeted 250 11.1 %
    oldgen::ArenaSweepObjectsState::reclaim_dead_object 73 3.3 %
    array::element_shape::forget_element_shape + clear_element_shape 102 4.5 %
    oldgen::IncrementalSweepState::step 46 2.0 %
    trace::ValidPointerSetBuilder::step 45 2.0 %
    addr_class::try_read_tracked_gc_header 29 1.3 %
    gc_type_finalize_unmarked_payload 25 1.1 %
    verify::OldToYoungRememberedRebuildState::step + remember_retained_old_to_young_slots 48 2.1 %

    The mark itself (mark_addr, the slot visitors) is not in the top 26 leaves. So "a cheaper full mark" was the wrong lever: a full is dominated by walking every object in the arena (the object cursor, used by the valid-pointer-set pre-pass and by the sweep) and by per-dead-object reclamation — and the element-shape side table is forgotten per dead array (every record's tags), which looks like a hash operation per dead array and is the one cheap item here (check a header bit before touching the table).

    Why parse/scan loops do not get this regime today. The trigger order is OldReclaim first, then the nursery. In the knob run OldReclaim was due at every safepoint only by accident (each 20 MB born-old result tripped the absolute 48 MB arm), so the full always pre-empted the minor and the trees never left Eden. A parse/scan loop has no born-old results, OldReclaim is never due, the nursery minor runs, and the in-place promotion (correctly: the tree is live at that minor) moves the tree to old-gen, where only a full can reclaim it — the shape this issue opened with.

    Mechanism that would make the regime deliberate rather than accidental:

    1. A full promotes its Eden survivors in place when the young generation is essentially all live after the mark (the same 95 % test the copying minor uses): retag the in-use young blocks to old via the arena/promote.rs machinery the copying minor already uses (retag_young_for_in_place_promotion / finish_in_place_promotion), after the sweep has reclaimed the dead objects in them; clear the remembered set (exact, no young left); credit the baseline; feed the survival measurement. Today a full is non-moving and promotes nothing (gc/mod.rs, perf: json_pipeline at 500k records is 97.6x bun (60.4s vs 618ms) while the same workload at 100 records BEATS bun — a scaling cliff, not a constant factor #7592's latch comment), which is why alternating a full with a minor ends in an evacuating minor that copies the whole tree (the failure the cohort-bound experiment above ran into).
    2. The promoted-cohort bound from the earlier comment, with a floor of one nursery rather than 64 MB. With (1) in place the steady state of a parse/scan loop is one full per iteration and no minors: mark T_i (cheap), sweep and reclaim T_{i−1}, promote T_i in place — the knob run's cost structure, ~18 ms per iteration at 20 MB. Expected: 20m parse ≈ 54 ms/iteration (vs the 208 ms/4 = 52 ms best, i.e. parity within noise) at ~150 MiB (0.7× Node); 8m scan ≈ 24 ms/iteration (vs 189/8 = 23.6 best) at ~90–110 MiB (parity); roundtrip as measured above.
    3. The sweep-side items in the table are the CPU headroom that makes (2) free instead of marginal: skip forget_element_shape for arrays that never carried a proof, and look at whether the valid-pointer-set pre-pass and the sweep can share one object walk.

    This changes gated gc-ratchet counters on the retain probes (fulls that promote), so it needs a baseline regeneration on the quiet mini alongside the corpus run, and the cc rows on perrymaster. Not started; the two branches from the earlier comments (gc/promoted-cohort-bound, gc/nursery-cap-in-place) are local and contain the counters/diag fields (2) reuses.

  4. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Two measurements that re-scope this issue after #10204 (full promotes Eden survivors; draft, does not meet its bar: CPU roughly doubles on the 20 MB parse/scan/sparse rows and 8m:scan).

    1. Perry's tree is not bigger than V8's. Heap census (PERRY_GC_CENSUS, gc() after one retained JSON.parse of records_array_8m.json, 59,000 records, current main), minus the 7.1 MB source string that is live only in the tree run:

    type live bytes objects bytes/record
    string (name, email; tags entries are SSO) 6.13 MB 121,358 104
    object (records) 4.25 MB 59,000 72
    array (tags + the records array) 2.48 MB 59,009 42
    tree ≈ 12.4 MiB 239k ≈ 210

    Node 26.5.1, same document, --expose-gc heap delta: 13.0 MiB (231 B/record). So the RSS gap on the parse/scan rows is not representation size; it is how many dead trees are resident. On 8m:scan old gen ends at 105 MB ≈ 8 trees: every tree parsed is still there, because no full runs.

    2. A full is expensive because the sweep is per object, not because of the mark. ArenaSweepObjectsState walks every object in the arena with ArenaObjectCursor; each dead object gets old_page_account_swept_object, finalize_dead_arena_payload (element-shape forget, layout-mask and overflow side tables), and a deferred page-index unregister; blocks with no live object are reset only after that walk (block_has_live). After a parse/scan loop nearly every old block holds only dead objects, so a full's cost scales with dead objects (#10204's head: ~23 ms per full on 8m:scan for a ~150 MB arena with one or two live 12 MB trees), which is what turns "reclaim every few iterations" into a CPU regression.

    The lever that fits the CPU headroom: block-granular reclamation. Record during the mark which blocks received a marked object (a per-block bit or count set where GC_FLAG_MARKED is set). In the sweep, a block with no marked object is reclaimed without visiting its objects, provided its dead population needs no per-object finalization: page-index entries dropped by block range, address-keyed side tables purged by range (or proven empty for the block), then the existing block reset. Blocks with any live object keep today's per-object path. With that, a full that runs every 2–3 iterations costs roughly the mark of the live tree (~240k objects on 8 MB, ~1.1M on 20 MB) plus O(blocks), i.e. about 1–3 ms per iteration amortized — inside the headroom of 8m:scan (~1.5 ms/iteration) and the 20 MB rows (~15 ms/iteration) — and old gen holds at most 2–3 trees, which is the RSS parity point on both.

    Per-object obligations that must stay exact under a block skip, each to be settled before relying on it: finalizer hook types (GcFinalizeHookKind: Map/Set side allocations, promises, native arena owners/views — a block containing any such object keeps the per-object path), element-shape records (array/element_shape.rs), layout slot masks / typed layouts, overflow fields and closure dynamic props, weak refs and finalization registries, the remembered set and dirty-page coverage, the old-page free-hole list (#7437), forwarding stubs (#6228), and pinned objects.

  5. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Checked one more premise before the next step, with no build: PERRY_GC_FORCE_EVACUATE=1 on #10220's head worker (it disables in-place and untraced promotion, so every minor evacuates and measures).

    cell default CPU / peak forced evacuation CPU / peak survival at every minor
    records_array_8m:scan 178 ms / 187 MiB (3 minors, 2 untraced) 504 ms / 284 MiB (8 minors) 998–1000 ‰
    records_array_20m:parse 141 ms / 241 MiB (2 minors, 1 untraced) 566 ms / 417 MiB (3 minors) 999–1000 ‰

    So the untraced promotions are not acting on a stale reading: at minor time the young generation really is all live (the previous tree is still held by last while the current one is being scanned), and each promoted tree dies only after the minor. #10204's 333–500 ‰ was measured at the fulls it inserted, not at these minors. Evacuating instead copies live trees at ~72 ns/object and doubles the footprint — ruled out.

    Consequences for the remaining rows:

    • 8m:scan (189 vs 110 MiB, 13.7 ms CPU lead over 8 iterations): any collection that reclaims its dead trees must trace the live 240k-object tree; at gc: make a synchronous full cheaper per live object (#10182) [pacing retry does not meet acceptance] #10220's ~42 ns/object mark that is ~10 ms per collection with every other phase at zero, and reaching 110 MiB needs about three. It does not fit the lead without roughly halving the per-object mark again.
    • 20 MB parse/scan/sparse and object_20m:parse (240 vs 220–225 MiB, ~64 ms lead): k=1 pacing already reaches 179 MiB with two fulls; the mark of the 585k-object tree is ~25 ms each, but a pacing full costs 62–68 ms today: census over the dead promoted trees (27 %), page-run re-materialization for blocks the sweep then releases whole (15 %), and the forced conservative stack scan because pacing fulls fire at an allocation point. Bringing a pacing full down to roughly its mark puts these four rows inside the lead. That is the next task, on gc/full-throughput.
    • object_8m:parse (118 vs 112): different regime (5 evacuating minors, no fulls); needs its own look.
    • small_record:parse (80 vs 60): the fix(gc): price the tiny-parse pressure guard by the productivity backoff (#9831) #9838 tiny-parse floor, a cc-rig decision.
  6. 40 remaining items

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions