Repository navigation
perf(json): JSON.stringify costs ~3,000 instructions per object visited (~20x node); array elements and primitives are fine #10696
Description
Activity
Confirmed — the
toJSONprobe is 66% of the per-object cost. Plus two corrections to the numbers above, both mine.Callgrind, deltas between N=20,000 and N=2,000 divided by 18,000, so #10686's one-time realm cost (a single 24.2M-Ir
js_get_global_this_builtin_valuecall, visible in the raw profile) cancels exactly.perf statat N=1e5→1e6 reproduces the same per-op cost within 2%, and the deltas reproduce this issue's table to 0.03% (7,649.4 / 16,490.0 against the 7,651 / 16,496 reported above).to_json_definitely_absent_after_own_keys(stringify_tojson_probe.rs:350) is 1,493 of 2,256 instructions per object — 66%. Its whole subtree is 47.6% ofJSON.stringifyon the nesting shape.Corrections to the body above
- "~2,950 per object" was per object plus one property — the nesting pair moves both together. Splitting with the property pair gives 2,256.3 per object and 690.5 per property, constant 1,761.3. That model fits all four programs to 0.07%.
- node's per-object term is 78, not 144. At N=2,000→20,000, node's startup (105–118M Ir, varying by ±10M run to run) is 3–5× the work being measured, so that fit was noise-dominated. At N=1e5→1e6 node is stable to 2% over three reps. The real gap is 29×, not 20×.
The dominant work has no per-instance input
to_json_definitely_absent_after_own_keysdoes three things, and two of them take nothing about the instance:class_chain_may_have_to_json(class_id)— 448 Ir/object, five by-name/by-id lookups across five separate registries. Its only argument is the class id.object_proto_may_have_to_json()— 584 Ir/object, and it takes no arguments at all. It is a process-global question, and almost all of the 584 is revalidating its own memo by rebuilding a six-field signature with two tracked-GC-header reads.
1,032 of the 1,493 — 69% of the probe, 46% of the entire per-object cost — is re-deriving answers that do not depend on the object being serialised.
Two structural findings behind that:
- Every object is probed twice, once by its parent before descending and once by its own
stringify_object_inner. Call counts: 40,000object_get_to_jsonfor 20,000 single-object iterations; 200,000 for 20,000 five-object iterations. compute_object_proto_tojson_stateruns once perJSON.stringifycall — the memo is force-invalidated at every top-level entry, so the first probe of every call always recomputes and the rest each pay 292 Ir to revalidate.
Two witnesses
- Fresh vs hoisted literal. Allocating a new
{a:{b:1}}every iteration costs byte-identical probe cost — 1,493.0 either way. Class ids are per shape, so a per-shape memo would hit 100% of the time. JSON.parse-built vs literal, same shape. The five class-registry functions drop to exactly zero (−448) and every remaining probe function halves; −33.4% per object overall.
That second witness exposes a separate defect.
to_json_definitely_absent_without_gcanddata_record_global_to_json_absent_without_gcboth open withif class_id != 0 { return false; }— and the file's own comment records that HIR gives object literals anonymous shape classes. So the fast plain-data record emitter is dead for every object literal in every program. It only ever runs onJSON.parseoutput.Fix direction
- Memoise
class_chain_may_have_to_jsonperclass_idunder an epoch — −448/object, high confidence. - Replace the
Object.prototypesignature revalidation with an epoch compare, and stop invalidating at every top-level entry — −584/object, −390/call. High confidence in the measurement; medium that it is this simple, since the moving-GC token trick and mid-stringifymonkey-patching are the correctness risks. - Stop double-probing.
- Re-gate the plain-data emitter on a cached ABSENT verdict rather than
class_id == 0.
(1) and (2) alone take the five-object shape from 16,490 to ~9,150 Ir/op, −45%.
Falsifier: if a change claims to remove those functions and the same two-N callgrind delta does not move by ~1,032 Ir/object, this attribution is wrong.
- added 8 commits that reference this issue
on Sep 19, 2026
Summary
JSONis one of the two primitives where perry loses to both node and bun (#10695 has the crossover map). The cost is not per-element — arrays scale well. It is a ~3,000-instruction per-object overhead, about 20× node's, paid on every object visited.perry crosses below node just under N=10,000 round-trips, and below bun before that.
Attribution
One program per shape, size from
process.argv[2], cost per operation fitted across N=2,000→20,000 so startup drops out.perf stat -e instructions:uon perrymaster. perry from main68a545439+ #10611; node v26.8.1; bun 1.3.14. Output checked equal to node on every row.{a:1}{a:1,b:2,c:3,d:4}{a:{b:1}}— 2 objects{a:{b:{c:{d:{e:1}}}}}— 5 objects[1][{a:1},{a:2},{a:3},{a:4}]Fitting the nesting pair (2 objects → 5 objects) isolates the per-object term cleanly:
Array elements are fine. Properties are moderately expensive. The per-object term dominates everything and is the defect.
Where I would look
The bisect on #10686 found
JSON.stringifyon an object reachingcompute_object_proto_tojson_state(stringify_tojson_probe.rs:165→prototype_chain.rs:570) — thetoJSONprobe. That is currently notable as one of the operations that force the whole globalThis realm to populate the first time.A by-name prototype-chain walk to ask "does this object or its prototype have
toJSON?", run once per object visited, would account for a per-object constant of this size. That is a hypothesis from the call path, not a profile of the steady state — it should be confirmed with callgrind on the steady-state loop before anyone changes code.If it is the probe, the fix direction is to answer the question without a by-name walk per object — the overwhelmingly common case is an object whose prototype is the pristine
Object.prototypewith notoJSON, which is a property of the shape and not of the instance.Notes on measurement
JSON.stringifyandJSON.parseof primitives are healthy — perry beats node on both (stringify(12345)858 vs 1,118;parse("12345")363 vs 695). The regression is specific to containers.node's per-op figure for the smallest shapes is unstable across runs (
{a:1}measured 124/op in one run and 1,047/op in another, same program and method), so small-shape ratios against node should not be quoted tightly. The per-object term above is fitted from the nesting pair, where both runtimes were stable.Related: #10695 (crossover map), #10686 (the one-time realm constant sharing this call path), #10689.