Skip to content

perf(box): struct Box { value: u64 } has no header, so identity and capture counts live in address-keyed side tables — ~9.8% of native tsc, and likely a memory loss too #10703

Description

@proggeramlug

Summary

/// A box is simply a heap-allocated JSValue bit slot.
#[repr(C)]
pub struct Box { pub value: u64 }

A box has no header — eight bytes of payload and nothing else. It therefore cannot answer any question about itself, so every question is answered by a thread-local side table keyed by its address:

  • BOX_REGISTRY (PtrHashSet<usize>) — "is this pointer a box?", because the object carries no identity;
  • BOX_CAPTURE_COUNTS (RefCell<PtrHashMap<usize, usize>>) — the capture-edge count, because the object carries no count field.

Giving Box one header word would let both live in the object. On the profile below that is ~9.8% of a real run, and — because a hash entry costs more than the word it would replace — it also looks like a net memory win.

Measured

tsc --noEmit demo.ts (two-line input) compiled with Perry 0.5.1596 + #10656/#10685, sample leaf histogram, 4,917 attributed samples:

symbol samples share what it is
is_registered_box_ptr 254 5.2% BOX_REGISTRY probe (+ plausibility heuristic + 8-slot cache)
js_closure_set_box_capture_ptr 171 3.5% capture-slot publish
increment_cell_capture_count 107 2.2% BOX_CAPTURE_COUNTS hash
decrement_cell_capture_count 75 1.5% BOX_CAPTURE_COUNTS hash
publish_box_cell 44 0.9%

is_registered_box_ptr is the single largest symbol in the whole profile, which is otherwise flat (next is 4.0%).

Both counting paths are the same shape:

fn increment_cell_capture_count(cell: usize, amount: usize) {
    BOX_CAPTURE_COUNTS.with(|counts| {        // thread-local resolution
        let mut counts = counts.borrow_mut(); // RefCell borrow
        let record = counts.entry(cell).or_default();   // hash lookup + insert
        ...

Why the current design cannot be optimised further in place

box.rs's own comment explains the ceiling, and it is correct:

A negative cache would NOT be sound — an address that is not a box today can be minted as one tomorrow — so a miss always falls through to the hash set, and only a confirmed positive is recorded.

The same comment records that the positive cache already took this family from 8.2% + 5.9% + 5.5% down to today's numbers. That was the available win. What remains is structural: a slot holding a raw box pointer is indistinguishable from a slot holding a NaN-boxed value, so the runtime must ask an external authority, and negatives — the overwhelming majority — can never be cached because the question is about an address rather than about a value.

A header answers the question at the object, where "not a box today, box tomorrow" cannot arise: the bytes either are a box or are not, and the check is a load and a compare.

Proposal: give Box a header word

#[repr(C)]
pub struct Box { pub header: u64, pub value: u64 }   // 8 -> 16 bytes

What that removes:

  • BOX_REGISTRY entirely — is_registered_box_ptr becomes a tag compare on a header the caller is about to touch anyway;
  • the three 8-slot BoxPtrCachees in HotTls and their eviction/quarantine interplay;
  • is_plausible_box_ptr's heuristic, and with it the perry#4898 false-positive class (a read-only __TEXT.__cstring address that passes every structural check) — a tagged header cannot be forged by an unrelated allocation;
  • BOX_CAPTURE_COUNTS — the count becomes a field, so increment/decrement is a load/add/store instead of TLS + RefCell + hash.

Memory (estimate, labelled as such). Per live box today: 8 B of Box, plus a PtrHashSet entry (~10 B with control bytes and load factor), plus — when captured — a PtrHashMap<usize, usize> entry (~19 B). That is up to ~37 B. With a header: 16 B and no tables. So this should reduce footprint as well as work, which is unusual for a "make it faster" change and is worth measuring before and after. It also fits the GC census finding of 147 MB of side tables against a 23 MB live heap.

Scope and risks

  • Box, I32Box and BoolBox are all headerless (I32Box and BoolBox are align(8) wrappers around a scalar), so all three would change shape. Codegen emits box allocations and capture-slot reads/writes directly, so the field offset change reaches perry-codegen, not just the runtime.
  • The GC traces boxes via the registry today; a header makes tracing more precise rather than less, but the tracing path must be moved over in the same change.
  • Worth checking whether the header can be narrower than a word (a tag byte plus a 32-bit count would fit in one u32 pair) before committing to +8 B.

Relationship to the wider pattern

Predicates answering "what kind of pointer is this?" are 18.5% of this profile — is_registered_box_ptr 5.2%, is_registered_buffer_slow 2.8%, classify_heap_generation_uncached 1.3%, resolve_strategy_slow 1.2%, get_parent_class_id 1.2%, is_class_object_ptr 1.1%, keys_find_slot_by_bytes 1.0%, plus is_registered_map and classify_arena. Add the capture-count tables and it is roughly 27% of the run spent on metadata that is stored beside objects instead of in them.

Arrays already do this the right way and it is measurably cheaper: receiver_may_be_registered_exotic reads the GC header's obj_type and is one warm byte plus a compare — see #10694, where wiring the indexing path to it halved 79.7 M probes. Boxes have no equivalent because they have no header to read.

Related: #10694, #10688.

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    on Sep 19, 2026
  2. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Two findings that unblock this — I was too cautious above

    1. It is NOT an ABI break. Codegen never dereferences a box.

    I claimed this reaches perry-codegen and the wire format. It does not. Codegen treats a box pointer as an opaque i64 and goes through runtime calls for every access — runtime_decls/strings.rs:1112-1117:

    js_box_alloc_bits(I64) -> I64
    js_box_get_bits(I64)   -> I64
    js_box_set_bits(I64, I64) -> VOID
    

    and the capture-read lowering says so explicitly (expr/literals_vars.rs:331-335):

    If the captured id is a boxed var, the capture slot holds a raw box pointer. Read the capture, extract the box pointer, and deref via js_box_get_bits.

    There is no GEP, no field offset, no inlined load anywhere in perry-codegen for a box. The layout is entirely a runtime concern, so widening Box is a runtime-local change.

    2. The real constraint is safe dereferencing — and the primitive already exists

    A header alone does not solve it, which is worth stating because it is the non-obvious part. is_registered_box_ptr is handed an arbitrary u64 out of a capture slot; to read a header you must dereference it, and if it is a NaN-boxed double instead you have a wild read. That is what the registry actually guards, and why is_plausible_box_ptr's structural checks are not enough — perry#4898 is a read-only __TEXT.__cstring address that passes all of them.

    But arena::page_meta::classify_heap_generation(addr) already answers "is this address inside a perry arena page?", returning HeapGeneration::Unknown for anything foreign. It is with a one-entry hot-TLS cache on the hit arm (the out-of-line miss arm is classify_heap_generation_uncached, itself 1.3% of the profile), and it already runs on every heap store the write barrier sees.

    So the sound sequence is:

    1. is_plausible_box_ptr(p)                       // existing, free
    2. classify_heap_generation(p) != Unknown        // existing, O(1) cached -> safe to deref
    3. (*p).header == BOX_MAGIC                      // definitive, one load + compare
    

    No registry, no per-box hash entry, no 8-slot caches, and no "not a box today, box tomorrow" problem — the bytes either carry the magic or they do not. Step 3 also closes the #4898 class: a __cstring address fails step 2 outright.

    The capture count then lives in the same header word (or a second field), so increment_cell_capture_count / decrement_cell_capture_count become a load/add/store instead of TLS + RefCell + hash.

    Scope, restated

    Runtime-local, but not small: Box, I32Box and BoolBox all change shape, and the allocation, quarantine/free-pool, scope_release and GC-tracing paths move with them (27 call sites in box.rs, 17 in box/scope_release.rs). The thing that made me hedge — codegen and the tag ABI — turns out not to apply, and the safe-deref primitive it needs is already written and already hot.

    Worth measuring before and after: per live box today is 8 B of Box plus a PtrHashSet entry plus, when captured, a PtrHashMap entry. A 16 B header'd box with no tables should win on both axes.

  3. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Correction to my previous comment: classify_heap_generation does NOT work here

    I proposed gating the header read on arena::page_meta::classify_heap_generation(addr) != Unknown. That is wrong. Boxes are not arena-allocated:

    // js_box_alloc_bits, box.rs:718-720
    let layout = Layout::new::<Box>();
    let ptr = alloc(layout) as *mut Box;

    Straight from the system allocator, which BOX_REGISTRY's own doc comment already says — "boxes come from std::alloc::alloc directly, not the GC arena". So classify_heap_generation would answer Unknown for every box and the guard would reject all of them. Anyone implementing from my last comment would hit that immediately. Apologies.

    Finding #1 from that comment still stands and is verified: codegen treats box pointers as opaque i64 and never dereferences one, so the layout is runtime-local.

    What the registry is actually doing, and why this is bigger than a header

    BOX_REGISTRY carries two responsibilities, and a header only addresses the first:

    1. Identity — "is this arbitrary u64 from a capture slot a box?" It is the authoritative answer because there is no other way to know an address came from js_box_alloc_bits.
    2. GC root enumeration — the collector traces the JSValue bits inside each live box. Per the doc comment, without it a box-captured heap object "is never marked AND the JSValue bits inside are never scanned — heap objects referenced only through box-captures can be swept mid-await."

    A header cannot supply (2): you cannot enumerate boxes from their headers without scanning the heap. So "add a header, delete the registry" is not a viable shape.

    The design that does work: allocate boxes from a dedicated slab

    Move box allocation out of std::alloc into a dedicated slab/region. Then:

    • Identity becomes an address-range or page test against the slab — O(1), no hash, and safe to dereference, which is the precondition a header needs. A header magic then confirms, closing the perry#4898 __TEXT.__cstring class.
    • GC enumeration becomes a walk of the slab, which replaces (2) rather than dropping it.
    • The capture count moves into the box, so increment_cell_capture_count / decrement_cell_capture_count stop being TLS + RefCell + hash.
    • BOX_REGISTRY, I32_BOX_REGISTRY, BOOL_BOX_REGISTRY, the three 8-slot caches and BOX_CAPTURE_COUNTS all go away — that is the ~9.8% plus their memory.

    It should also help locality: the registry is pre-sized to 128 k buckets (~2 MB) because promise_all_chains allocates ~150 k boxes per run, and those boxes are currently scattered across the malloc heap.

    Honest scoping

    This is an allocator change, not a struct change — larger than I implied twice now. It touches allocation, the free pool and quarantine, scope_release, and GC tracing (27 call sites in box.rs, 17 in box/scope_release.rs), and the slab needs its own lifetime story for released cells.

    The measurements motivating it are unchanged and stand: is_registered_box_ptr is the single largest symbol in the profile at 5.2%, the capture-count pair adds 3.7%, and the wider "what kind of pointer is this?" family is 18.5%.

  4. proggeramlug commented on Sep 30, 2026

    @proggeramlug
    ContributorAuthor

    Current target (2026-09-30): the box still has no header; a boxed hot read/write costs +186 instructions per call over an object holder

    Package impact (#11464)

    Bucket row 6, closures / boxed captures / arguments, is 4.6% of equal-weight excess. It is ≥5% in 9 packages. Share of each package's excess: dayjs 12, qs 9, moment 8, axios 7, decimal.js 7, fastify 7, rate-limiter-flexible 6, ioredis 5, node-forge 5.

    Most of that bucket is not box cells or arguments. Most of it enters under the method dispatcher (js_typed_feedback_native_call_method_by_id) and the IC miss. Its top chains end in function-object property leaves (closure_get_dynamic_prop, is_closure_ptr, closure_is_key_deleted). The part that enters through box and arguments entry points is small:

    workload entry → leaf instr/iter % of excess
    dayjs/diff_startof js_arguments_object_alloc 239k 2.6
    dayjs/parse_format js_arguments_object_alloc 58k 2.9
    dayjs/parse_format js_closure_set_box_capture_ptr → is_registered_box_ptr 46k 2.3
    qs/stringify_nested js_closure_set_box_capture_ptr (side-channel factories) 263k 1.7
    qs/parse_nested js_closure_set_box_capture_ptr 37k 1.0
    mongodb/batch_query box_get_bits_named 3.07M 1.4
    rate-limiter-flexible/get_penalty box_get_bits_named + js_box_set_bits + set_box_capture_ptr 4.5k 3.2
    moment/parse_format js_arguments_object_alloc 36k 1.0
    redis/set_get set_box_capture_ptr + js_box_release 86k 1.3

    Boxes also add root-scan work to the copying minor (the gc_minor bucket). The data cannot separate that share.

    For this issue specifically, the leaf is_registered_box_ptr under js_closure_set_box_capture_ptr is the BOX_REGISTRY probe this issue wants to delete (dayjs/parse_format, 46k instr/iter).

    What landed since the issue was filed

    Reproducer, fresh numbers (instructions per call)

    function mkBoxed() { let c = 0; return (x: number) => { c = (c + x) & 0xffff; return c; }; }
    function mkObj()   { const st = { c: 0 }; return (x: number) => { st.c = (st.c + x) & 0xffff; return st.c; }; }
    const f = variant === "boxed" ? mkBoxed() : mkObj();
    let s = 0; for (let i = 0; i < n; i++) s = (s + f(i & 7)) | 0;
    variant Perry Node
    state in an object (control) 217 16
    state in a captured let (boxed) 403 44
    boxed − control +186 +27

    The whole-program cost (box creation and capture-count maintenance per closure) is measured on #10520: arrow8 is 36.5k vs 10.2k instructions per call.

    Acceptance target

    • boxed within 1.2× of the object control (≤ ~260 instr per call). No is_registered_box_ptr or increment/decrement_cell_capture_count in a profile of the reproducer.
    • Peak RSS no higher on tsc and Zod. The issue expects a memory win, so measure it.
    • No regression on dayjs, qs, redis and rate-limiter-flexible.

    Package check (Linux, needs perf). Build with cargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static, then run:

    (cd benchmarks/packages && npm ci --ignore-scripts)
    python3 scripts/package_bench.py compile --perry-bin-dir /tmp/pb --filter dayjs --filter qs --filter redis --filter rate-limiter-flexible
    python3 scripts/package_bench.py run --perry-bin-dir /tmp/pb --arms node,perry --modes instr --filter dayjs --filter qs --filter redis --filter rate-limiter-flexible --out /tmp/pb/instr.json

    Run it once on a base-commit build and once on the branch, and compare instructions per iteration. --filter is a workload-id substring, and control/* always runs. For attribution, run profile --callgraph on PERRY_KEEP_SYMBOLS=1 binaries (see benchmarks/packages/PROFILE.md).

    Fresh numbers were measured on origin/main 5fbc2c3 (v0.5.1654) with cargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static. Binaries were compiled with PERRY_NO_AUTO_OPTIMIZE=1 and compared with Node 26.5.1 on Linux x86-64 using perf stat -e instructions:u. Each figure is the median of 3 runs at two sizes of N, with per-op = ΔI/ΔN. N1 is at least 200k operations (20k calls for the factory bench), which keeps most of Node's JIT warm-up inside the constant term. Output was byte-identical to Node on every row. The #11464 figures come from benchmarks/packages/profile/callgraph.{md,json}, measured at Perry 36420d2 with auto-optimize. node-forge was measured from 2febf42 binaries. "% of excess" means the share of a package's (Perry − Node) instructions per iteration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions