Repository navigation
perf(buffer): 79.7M is_registered_buffer probes for 9 buffers (90 true positives) — 2.8% of native tsc; the diag itself sizes a 1024-bit Bloom at 0% FP #10694
Description
Activity
- addedperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performance
on Sep 19, 2026 - added a commit that references this issue
on Sep 19, 2026 Partial fix pushed to #10664; the other half is still unattributed
array/indexing.rswas probing the buffer and typed-array registries on every element access.receiver_may_be_registered_exotic(aGC_TYPE_ARRAYheader test) already existed for exactly this and was used by 13 call sites inarray/iter_methods.rsand none inarray/indexing.rs. Gatingjs_array_get_f64/js_array_set_f64/js_array_set_f64_extendand the iteration-exotic helpers:before: probes=79,691,777 admits=26,198,956 true_positives=90 after: probes=39,845,889 admits=21,646,032 true_positives=90Exactly halved, correctness unchanged (443 tests pass; a differential test against Node covering every typed-array kind, Uint8Array wrapping, Uint8ClampedArray clamping, f32 precision, Buffer, subarray aliasing and an Array subclass is byte-for-byte identical).
Two honest caveats
1. Not measurable on tsc. Five interleaved rounds: median -1.2%, ranges overlap (A 6.85-8.02 s, B 6.76-7.86 s). That is consistent with the attribution rather than contradicting it —
is_registered_buffer_slowwas 2.8% of leaf samples, so halving its calls predicts ~1.4%, which is below this workload's run-to-run variance. The change stands on removing provably-wasted work and on matching the gate the iteration helpers already use, not on a demonstrated speedup.2. The other ~39.8 M probes are NOT in
array/indexing.rs. I gated the four remaining||-shaped sites in that file and the probe count moved by zero. So they come from elsewhere —typed_feedback.rs,typedarray/access.rs,json/stringify.rs,object/instanceof.rsandnode_stream_readwrite.rsall callis_registered_buffer, and I have not attributed which dominates. Reading callers out of asamplecall tree collapsed to a single compiled tsc closure, so that needs a call-site-counting diag rather than archaeology.The Bloom filter the diagnostic already sizes (1024-bit/3-hash, 0.0% false-positive here) remains the better fix, because it makes every surviving probe cheap regardless of which caller issues it.
Validated on a real workload: this matters MORE than the two-line measurement suggested
Every number in this issue came from
tsc --noEmit demo.ts— a two-line file. I re-profiled the same binary against a genuine project (kimi-code: 3,449 TypeScript files, 27 MB of source), and the picture changes.is_registered_buffer_slowis the #3 symbol at 3.7% of 28,991 leaf samples — higher than the 2.8% measured ondemo.ts, and that is with the array-indexing gate already applied.This confirms the caveat I recorded when landing the gate: the remaining ~39.8 M probes originate outside
array/indexing.rs, and on real multi-file checking they are a larger share of the run, not a smaller one. The call-site gate was the cheap half; it is not the fix.The Bloom filter the diagnostic sizes itself remains the right answer — 1024 bits, 3 hashes, 0.0% false-positive on the measured admission set, 128 bytes of state replacing tens of millions of hash probes. It makes every surviving probe cheap regardless of which caller issues it, which is exactly what is needed when the callers are spread across
typed_feedback.rs,typedarray/access.rs,json/stringify.rs,object/instanceof.rsandnode_stream_readwrite.rs.For context, the full bucket comparison between the two workloads:
bucket 3,449 files two-line file pointer-kind predicates 16.9% 18.5% GC / layout 12.4% 23.7% boxes / closures 12.2% 25.4% property access 11.3% 18.3% strings 9.7% 3.6% The predicate family is the one priority that transfers cleanly across both workloads.
Bloom filter landed in #10664 — 21.6M hash probes eliminated
The diagnostic sized this filter itself and it holds exactly:
before after probes 39,845,889 39,845,889 admits (reach TLS + RefCell+ hash)21,646,032 (54.32%) 862 (0.00%) true_positives 90 90 precision 0.000416% 10.44% 25,100× fewer hash probes, answers unchanged.
true_positivesholding at exactly 90 is the soundness proof — a Bloom that rejected one real buffer would drop it.Timing, honestly: median −1.3% on
tsc --noEmit demo.tswith ranges overlapping (A 8.520–8.920 s, B 8.400–8.660 s) — but B is faster in 7 of 7 interleaved rounds, p ≈ 0.008 under the null. The paired comparison is the meaningful one here; it controls for the machine drift that inflates the ranges. I measured on the demo input deliberately: the same 39.8M probes land in an ~8 s run there rather than a ~140 s one, so effect density is ~17× higher. On the 3,449-file project the same change reads +1.7% with ±10 s of variance, i.e. below that workload's resolution.A near-miss worth recording
My first version armed only one of the two writers. The file documents the contract — every writer arms before publishing — and I quoted it in my own comment while violating it, because a blind single-occurrence text replacement hit
note_buffer_like_registeredinstead ofregister_buffer. Ordinary buffers were therefore never added to the filter and every one was rejected.It failed loudly (tsc could not parse its own
lib.d.ts), but a Bloom false negative makesinstanceof Uint8Arrayanswer wrong, and a different bit pattern could have made that subtle instead of obvious. It was caught immediately only becausetrue_positives == 90was fixed in advance as the acceptance test rather than checking timings first.Still open on this issue
The probe count is unchanged at 39.8M — the Bloom makes each one cheap but does not remove them. Those still originate outside
array/indexing.rs(typed_feedback.rs,typedarray/access.rs,json/stringify.rs,object/instanceof.rs,node_stream_readwrite.rs) and remain unattributed. 39.8M probes for a 9-element set is still 39.8M probes.Package-level measurement from the Phase 3 attribution (#11464,
benchmarks/packages/PROFILE.md):origin/main36420d2, release compiler, auto-optimize, Linux x86-64,perf record -e instructions:u --call-graph dwarfat two N (per-iteration, startup cancels), Node 26.5.1 oracle with outputs checked equal. "% of excess" = share of (Perry − Node) instructions/iter, equal-weight over the 23 packages with Perry/Node ≥ 2×.is_registered_buffer_slowalso runs on plain Array element stores. In node-forge/rsa_sign (jsbn big-integer loops, 20.6G instr/iter, 57× Node), 14.2% of all instructions arejs_dyn_index_set_strict→is_registered_buffer_slowforarr[i] = von an ordinary untyped Array (binaries at 2febf42, because node-forge does not compile on main: #11450).Across packages, frames matching
is_registered_bufferaccount for 2.1% of total excess: nanoid/generate 19% (js_uint8array_get/set,js_uint8array_index_get_value), node-forge/rsa_sign 19%, uuid/v7 7%, axios 6% (isBufferviaval.constructor.isBuffer), uuid/v5_parse 6%. The parents arejs_dyn_index_get,js_uint8array_get,js_uint8array_setandjs_uint8array_index_get_value, plusobject_static_prototypeon the dispatcher path. A Bloom pre-filter that rejects non-buffers before the slow probe would remove most of this.- added a commit that references this issue
on Sep 28, 2026 Progress: #11589 landed.
Uint8Array/Bufferelement access now goes through the byte-view admission cache:u8[i]on typed params 876 → 76 instructions per op, untyped params 2695 → 763, closure-captured nanoid table 1336 → 561. Package workloads: nanoid/generate −34%, uuid/v7 −13%, uuid/v5_parse −10% (big.js +1.2% from the admission test on the runtime Array paths).Still open: the untyped Buffer store has no inline arm yet (it uses the runtime fast path); that is a follow-up now that #11588's untyped-Array store structure has landed. The untyped read is still ~10× the typed path.
- added a commit that references this issue
on Oct 4, 2026
Summary
Compiling a two-line file with a natively compiled
tscissues 79,691,777is_registered_bufferprobes in a 7.8 s run — roughly 10 million per second — to answer a question about a set that never holds more than 9 buffers. 90 of those probes are true positives.is_registered_buffer_slowis 2.8% of leaf samples in that run.Measured
Perry 0.5.1596 (
fix/10656-codepointat-linear, i.e. with #10656/#10685 applied), macOS arm64, using the runtime's ownPERRY_BUFFER_DIAGinstrument:Workload:
./tsc --noEmit demo.tswheredemo.tsis two lines. Input file:Two distinct problems
1. The gate lets 26.2 M probes through to a 9-element hash set.
BUFFER_LIKE_ADDR_WINDOWis doing real work — it rejects 67.12% inline, exactly as designed — but the surviving 32.88% each pay a function call, a thread-local resolution, aRefCellborrow and a hash lookup to discover, 99.999656% of the time, that the answer is no. The window spans 90.3 MB inside a 668.6 MB probed range, so a min/max bounding box simply cannot separate 9 addresses from the heap around them.The instrument already names the fix and sizes it: a 1024-bit, 3-hash Bloom filter over registered addresses would be 0.0% false-positive on this workload — 128 bytes of state replacing 26.2 M hash probes. It fits the same soundness contract as the window (every writer arms it before publishing; a negative is authoritative, a positive falls through to the existing set).
2. The deeper question: why 79.7 M probes for 9 buffers? Nothing in this workload is buffer-shaped — it is a type-check of two lines of TypeScript. At ~10 M probes/second something on a very hot path is asking "is this a registered buffer?" about values that are overwhelmingly not buffers and never could be.
is_registered_bufferhas ~200 call sites per its own comment; the probe volume suggests one of them sits in a generic value/type ladder that runs per property access or per value touched.Fixing (1) makes each probe ~free. Fixing (2) removes them. (2) is the larger win and is likely a small change once the caller is identified — a
PERRY_BUFFER_DIAGvariant that samples call-site backtraces would find it immediately.Why this is worth doing
is_registered_buffer_slowis 2.8% of the run on its own, but it is one member of a family: predicates answering "what kind of pointer is this?" are 18.5% of the post-fix profile —is_registered_box_ptris_registered_buffer_slowclassify_heap_generation_uncachedresolve_strategy_slowget_parent_class_idis_class_object_ptrkeys_find_slot_by_bytesis_registered_map,classify_arena, …This one is the most clear-cut instance because the ratio is so extreme (90 hits in 79.7 M probes) and the runtime already ships the instrument that proves it.
Reproduction
Related: the same probe family is discussed in #10688's context; GC census recorded 147 MB of side tables against a 23 MB live heap, which is the memory face of this design.