Skip to content

perf(buffer): 79.7M is_registered_buffer probes for 9 buffers (90 true positives) — 2.8% of native tsc; the diag itself sizes a 1024-bit Bloom at 0% FP #10694

Description

@proggeramlug

Summary

Compiling a two-line file with a natively compiled tsc issues 79,691,777 is_registered_buffer probes in a 7.8 s run — roughly 10 million per second — to answer a question about a set that never holds more than 9 buffers. 90 of those probes are true positives.

is_registered_buffer_slow is 2.8% of leaf samples in that run.

Measured

Perry 0.5.1596 (fix/10656-codepointat-linear, i.e. with #10656/#10685 applied), macOS arm64, using the runtime's own PERRY_BUFFER_DIAG instrument:

[buffer-diag] probes=79691777 admits=26198956 (32.88 %) rejected=53492821 (67.12 %)
              true_positives=90 (0.000344 % of admits)
  window   [0x5abfc600008, 0x5ac0204f938] span  90.3 MB
  probed   [0x5abf7a30030, 0x5ac216c9f48] span 668.6 MB -- window covers 13.5 % of the probed range
  registrations=9 unregistrations=9 live_max=9
  => a 1024-bit/3-hash Bloom holding all admissions would be 0.0 % false-positive

Workload: ./tsc --noEmit demo.ts where demo.ts is two lines. Input file:

type Config = { port: number };
const cfg: Config = { port: "8080" };

Two distinct problems

1. The gate lets 26.2 M probes through to a 9-element hash set. BUFFER_LIKE_ADDR_WINDOW is doing real work — it rejects 67.12% inline, exactly as designed — but the surviving 32.88% each pay a function call, a thread-local resolution, a RefCell borrow and a hash lookup to discover, 99.999656% of the time, that the answer is no. The window spans 90.3 MB inside a 668.6 MB probed range, so a min/max bounding box simply cannot separate 9 addresses from the heap around them.

The instrument already names the fix and sizes it: a 1024-bit, 3-hash Bloom filter over registered addresses would be 0.0% false-positive on this workload — 128 bytes of state replacing 26.2 M hash probes. It fits the same soundness contract as the window (every writer arms it before publishing; a negative is authoritative, a positive falls through to the existing set).

2. The deeper question: why 79.7 M probes for 9 buffers? Nothing in this workload is buffer-shaped — it is a type-check of two lines of TypeScript. At ~10 M probes/second something on a very hot path is asking "is this a registered buffer?" about values that are overwhelmingly not buffers and never could be. is_registered_buffer has ~200 call sites per its own comment; the probe volume suggests one of them sits in a generic value/type ladder that runs per property access or per value touched.

Fixing (1) makes each probe ~free. Fixing (2) removes them. (2) is the larger win and is likely a small change once the caller is identified — a PERRY_BUFFER_DIAG variant that samples call-site backtraces would find it immediately.

Why this is worth doing

is_registered_buffer_slow is 2.8% of the run on its own, but it is one member of a family: predicates answering "what kind of pointer is this?" are 18.5% of the post-fix profile —

symbol share
is_registered_box_ptr 5.2%
is_registered_buffer_slow 2.8%
classify_heap_generation_uncached 1.3%
resolve_strategy_slow 1.2%
get_parent_class_id 1.2%
is_class_object_ptr 1.1%
keys_find_slot_by_bytes 1.0%
is_registered_map, classify_arena, … rest

This one is the most clear-cut instance because the ratio is so extreme (90 hits in 79.7 M probes) and the runtime already ships the instrument that proves it.

Reproduction

PERRY_BUFFER_DIAG=/tmp/bufdiag.log ./tsc --noEmit demo.ts
tail -6 /tmp/bufdiag.log

Related: the same probe family is discussed in #10688's context; GC census recorded 147 MB of side tables against a 23 MB live heap, which is the memory face of this design.

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    on Sep 19, 2026
  2. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Partial fix pushed to #10664; the other half is still unattributed

    array/indexing.rs was probing the buffer and typed-array registries on every element access. receiver_may_be_registered_exotic (a GC_TYPE_ARRAY header test) already existed for exactly this and was used by 13 call sites in array/iter_methods.rs and none in array/indexing.rs. Gating js_array_get_f64 / js_array_set_f64 / js_array_set_f64_extend and the iteration-exotic helpers:

    before: probes=79,691,777  admits=26,198,956  true_positives=90
    after:  probes=39,845,889  admits=21,646,032  true_positives=90
    

    Exactly halved, correctness unchanged (443 tests pass; a differential test against Node covering every typed-array kind, Uint8Array wrapping, Uint8ClampedArray clamping, f32 precision, Buffer, subarray aliasing and an Array subclass is byte-for-byte identical).

    Two honest caveats

    1. Not measurable on tsc. Five interleaved rounds: median -1.2%, ranges overlap (A 6.85-8.02 s, B 6.76-7.86 s). That is consistent with the attribution rather than contradicting it — is_registered_buffer_slow was 2.8% of leaf samples, so halving its calls predicts ~1.4%, which is below this workload's run-to-run variance. The change stands on removing provably-wasted work and on matching the gate the iteration helpers already use, not on a demonstrated speedup.

    2. The other ~39.8 M probes are NOT in array/indexing.rs. I gated the four remaining ||-shaped sites in that file and the probe count moved by zero. So they come from elsewhere — typed_feedback.rs, typedarray/access.rs, json/stringify.rs, object/instanceof.rs and node_stream_readwrite.rs all call is_registered_buffer, and I have not attributed which dominates. Reading callers out of a sample call tree collapsed to a single compiled tsc closure, so that needs a call-site-counting diag rather than archaeology.

    The Bloom filter the diagnostic already sizes (1024-bit/3-hash, 0.0% false-positive here) remains the better fix, because it makes every surviving probe cheap regardless of which caller issues it.

  3. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Validated on a real workload: this matters MORE than the two-line measurement suggested

    Every number in this issue came from tsc --noEmit demo.ts — a two-line file. I re-profiled the same binary against a genuine project (kimi-code: 3,449 TypeScript files, 27 MB of source), and the picture changes.

    is_registered_buffer_slow is the #3 symbol at 3.7% of 28,991 leaf samples — higher than the 2.8% measured on demo.ts, and that is with the array-indexing gate already applied.

    This confirms the caveat I recorded when landing the gate: the remaining ~39.8 M probes originate outside array/indexing.rs, and on real multi-file checking they are a larger share of the run, not a smaller one. The call-site gate was the cheap half; it is not the fix.

    The Bloom filter the diagnostic sizes itself remains the right answer — 1024 bits, 3 hashes, 0.0% false-positive on the measured admission set, 128 bytes of state replacing tens of millions of hash probes. It makes every surviving probe cheap regardless of which caller issues it, which is exactly what is needed when the callers are spread across typed_feedback.rs, typedarray/access.rs, json/stringify.rs, object/instanceof.rs and node_stream_readwrite.rs.

    For context, the full bucket comparison between the two workloads:

    bucket 3,449 files two-line file
    pointer-kind predicates 16.9% 18.5%
    GC / layout 12.4% 23.7%
    boxes / closures 12.2% 25.4%
    property access 11.3% 18.3%
    strings 9.7% 3.6%

    The predicate family is the one priority that transfers cleanly across both workloads.

  4. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Bloom filter landed in #10664 — 21.6M hash probes eliminated

    The diagnostic sized this filter itself and it holds exactly:

    before after
    probes 39,845,889 39,845,889
    admits (reach TLS + RefCell + hash) 21,646,032 (54.32%) 862 (0.00%)
    true_positives 90 90
    precision 0.000416% 10.44%

    25,100× fewer hash probes, answers unchanged. true_positives holding at exactly 90 is the soundness proof — a Bloom that rejected one real buffer would drop it.

    Timing, honestly: median −1.3% on tsc --noEmit demo.ts with ranges overlapping (A 8.520–8.920 s, B 8.400–8.660 s) — but B is faster in 7 of 7 interleaved rounds, p ≈ 0.008 under the null. The paired comparison is the meaningful one here; it controls for the machine drift that inflates the ranges. I measured on the demo input deliberately: the same 39.8M probes land in an ~8 s run there rather than a ~140 s one, so effect density is ~17× higher. On the 3,449-file project the same change reads +1.7% with ±10 s of variance, i.e. below that workload's resolution.

    A near-miss worth recording

    My first version armed only one of the two writers. The file documents the contract — every writer arms before publishing — and I quoted it in my own comment while violating it, because a blind single-occurrence text replacement hit note_buffer_like_registered instead of register_buffer. Ordinary buffers were therefore never added to the filter and every one was rejected.

    It failed loudly (tsc could not parse its own lib.d.ts), but a Bloom false negative makes instanceof Uint8Array answer wrong, and a different bit pattern could have made that subtle instead of obvious. It was caught immediately only because true_positives == 90 was fixed in advance as the acceptance test rather than checking timings first.

    Still open on this issue

    The probe count is unchanged at 39.8M — the Bloom makes each one cheap but does not remove them. Those still originate outside array/indexing.rs (typed_feedback.rs, typedarray/access.rs, json/stringify.rs, object/instanceof.rs, node_stream_readwrite.rs) and remain unattributed. 39.8M probes for a 9-element set is still 39.8M probes.

  5. proggeramlug commented on Sep 27, 2026

    @proggeramlug
    ContributorAuthor

    Package-level measurement from the Phase 3 attribution (#11464, benchmarks/packages/PROFILE.md): origin/main 36420d2, release compiler, auto-optimize, Linux x86-64, perf record -e instructions:u --call-graph dwarf at two N (per-iteration, startup cancels), Node 26.5.1 oracle with outputs checked equal. "% of excess" = share of (Perry − Node) instructions/iter, equal-weight over the 23 packages with Perry/Node ≥ 2×.

    is_registered_buffer_slow also runs on plain Array element stores. In node-forge/rsa_sign (jsbn big-integer loops, 20.6G instr/iter, 57× Node), 14.2% of all instructions are js_dyn_index_set_strict → is_registered_buffer_slow for arr[i] = v on an ordinary untyped Array (binaries at 2febf42, because node-forge does not compile on main: #11450).

    Across packages, frames matching is_registered_buffer account for 2.1% of total excess: nanoid/generate 19% (js_uint8array_get/set, js_uint8array_index_get_value), node-forge/rsa_sign 19%, uuid/v7 7%, axios 6% (isBuffer via val.constructor.isBuffer), uuid/v5_parse 6%. The parents are js_dyn_index_get, js_uint8array_get, js_uint8array_set and js_uint8array_index_get_value, plus object_static_prototype on the dispatcher path. A Bloom pre-filter that rejects non-buffers before the slow probe would remove most of this.

  6. proggeramlug commented on Sep 28, 2026

    @proggeramlug
    ContributorAuthor

    Progress: #11589 landed. Uint8Array/Buffer element access now goes through the byte-view admission cache: u8[i] on typed params 876 → 76 instructions per op, untyped params 2695 → 763, closure-captured nanoid table 1336 → 561. Package workloads: nanoid/generate −34%, uuid/v7 −13%, uuid/v5_parse −10% (big.js +1.2% from the admission test on the runtime Array paths).

    Still open: the untyped Buffer store has no inline arm yet (it uses the runtime fast path); that is a follow-up now that #11588's untyped-Array store structure has landed. The untyped read is still ~10× the typed path.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions