Skip to content

perf(json): JSON.stringify costs ~3,000 instructions per object visited (~20x node); array elements and primitives are fine #10696

Description

@proggeramlug

Summary

JSON is one of the two primitives where perry loses to both node and bun (#10695 has the crossover map). The cost is not per-element — arrays scale well. It is a ~3,000-instruction per-object overhead, about 20× node's, paid on every object visited.

perry crosses below node just under N=10,000 round-trips, and below bun before that.

Attribution

One program per shape, size from process.argv[2], cost per operation fitted across N=2,000→20,000 so startup drops out. perf stat -e instructions:u on perrymaster. perry from main 68a545439 + #10611; node v26.8.1; bun 1.3.14. Output checked equal to node on every row.

shape perry node bun
{a:1} 4,708 1,047 786
{a:1,b:2,c:3,d:4} 6,780 989 1,200
16 properties 15,990 3,457 3,332
{a:{b:1}} — 2 objects 7,651 1,491 947
{a:{b:{c:{d:{e:1}}}}} — 5 objects 16,496 1,924 1,586
[1] 2,815 511 1,288
16-element array 4,872 1,793 1,555
[{a:1},{a:2},{a:3},{a:4}] 13,025 1,455 1,884

Fitting the nesting pair (2 objects → 5 objects) isolates the per-object term cleanly:

term perry node ratio
per object visited ~2,950 ~144 ~20×
per property ~750 ~161 ~4.7×
per array element ~137 ~85 ~1.6×

Array elements are fine. Properties are moderately expensive. The per-object term dominates everything and is the defect.

Where I would look

The bisect on #10686 found JSON.stringify on an object reaching compute_object_proto_tojson_state (stringify_tojson_probe.rs:165 → prototype_chain.rs:570) — the toJSON probe. That is currently notable as one of the operations that force the whole globalThis realm to populate the first time.

A by-name prototype-chain walk to ask "does this object or its prototype have toJSON?", run once per object visited, would account for a per-object constant of this size. That is a hypothesis from the call path, not a profile of the steady state — it should be confirmed with callgrind on the steady-state loop before anyone changes code.

If it is the probe, the fix direction is to answer the question without a by-name walk per object — the overwhelmingly common case is an object whose prototype is the pristine Object.prototype with no toJSON, which is a property of the shape and not of the instance.

Notes on measurement

JSON.stringify and JSON.parse of primitives are healthy — perry beats node on both (stringify(12345) 858 vs 1,118; parse("12345") 363 vs 695). The regression is specific to containers.

node's per-op figure for the smallest shapes is unstable across runs ({a:1} measured 124/op in one run and 1,047/op in another, same program and method), so small-shape ratios against node should not be quoted tightly. The per-object term above is fitted from the nesting pair, where both runtimes were stable.

Related: #10695 (crossover map), #10686 (the one-time realm constant sharing this call path), #10689.

Activity

  1. proggeramlug commented on Sep 19, 2026

    @proggeramlug
    ContributorAuthor

    Confirmed — the toJSON probe is 66% of the per-object cost. Plus two corrections to the numbers above, both mine.

    Callgrind, deltas between N=20,000 and N=2,000 divided by 18,000, so #10686's one-time realm cost (a single 24.2M-Ir js_get_global_this_builtin_value call, visible in the raw profile) cancels exactly. perf stat at N=1e5→1e6 reproduces the same per-op cost within 2%, and the deltas reproduce this issue's table to 0.03% (7,649.4 / 16,490.0 against the 7,651 / 16,496 reported above).

    to_json_definitely_absent_after_own_keys (stringify_tojson_probe.rs:350) is 1,493 of 2,256 instructions per object — 66%. Its whole subtree is 47.6% of JSON.stringify on the nesting shape.

    Corrections to the body above

    1. "~2,950 per object" was per object plus one property — the nesting pair moves both together. Splitting with the property pair gives 2,256.3 per object and 690.5 per property, constant 1,761.3. That model fits all four programs to 0.07%.
    2. node's per-object term is 78, not 144. At N=2,000→20,000, node's startup (105–118M Ir, varying by ±10M run to run) is 3–5× the work being measured, so that fit was noise-dominated. At N=1e5→1e6 node is stable to 2% over three reps. The real gap is 29×, not 20×.

    The dominant work has no per-instance input

    to_json_definitely_absent_after_own_keys does three things, and two of them take nothing about the instance:

    • class_chain_may_have_to_json(class_id) — 448 Ir/object, five by-name/by-id lookups across five separate registries. Its only argument is the class id.
    • object_proto_may_have_to_json() — 584 Ir/object, and it takes no arguments at all. It is a process-global question, and almost all of the 584 is revalidating its own memo by rebuilding a six-field signature with two tracked-GC-header reads.

    1,032 of the 1,493 — 69% of the probe, 46% of the entire per-object cost — is re-deriving answers that do not depend on the object being serialised.

    Two structural findings behind that:

    • Every object is probed twice, once by its parent before descending and once by its own stringify_object_inner. Call counts: 40,000 object_get_to_json for 20,000 single-object iterations; 200,000 for 20,000 five-object iterations.
    • compute_object_proto_tojson_state runs once per JSON.stringify call — the memo is force-invalidated at every top-level entry, so the first probe of every call always recomputes and the rest each pay 292 Ir to revalidate.

    Two witnesses

    • Fresh vs hoisted literal. Allocating a new {a:{b:1}} every iteration costs byte-identical probe cost — 1,493.0 either way. Class ids are per shape, so a per-shape memo would hit 100% of the time.
    • JSON.parse-built vs literal, same shape. The five class-registry functions drop to exactly zero (−448) and every remaining probe function halves; −33.4% per object overall.

    That second witness exposes a separate defect. to_json_definitely_absent_without_gc and data_record_global_to_json_absent_without_gc both open with if class_id != 0 { return false; } — and the file's own comment records that HIR gives object literals anonymous shape classes. So the fast plain-data record emitter is dead for every object literal in every program. It only ever runs on JSON.parse output.

    Fix direction

    1. Memoise class_chain_may_have_to_json per class_id under an epoch — −448/object, high confidence.
    2. Replace the Object.prototype signature revalidation with an epoch compare, and stop invalidating at every top-level entry — −584/object, −390/call. High confidence in the measurement; medium that it is this simple, since the moving-GC token trick and mid-stringify monkey-patching are the correctness risks.
    3. Stop double-probing.
    4. Re-gate the plain-data emitter on a cached ABSENT verdict rather than class_id == 0.

    (1) and (2) alone take the five-object shape from 16,490 to ~9,150 Ir/op, −45%.

    Falsifier: if a change claims to remove those functions and the same two-N callgrind delta does not move by ~1,032 Ir/object, this attribution is wrong.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions