Skip to content

perf: stores to name/type/value/length/url/E… are 47–680× slower than Node (key literal not flagged interned, so the write IC never primes) #10500

Description

@proggeramlug

Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. The audit reported an "IC cliff at ≥ 5 subclass fields" in
@noble/hashes; reduction shows the cliff is not the field count but the field NAME: SHA-256's fifth state word is
called E. A static-key store o.<k> = v on a receiver whose layout is not statically known overwrites an existing
slot through the write PIC, but for keys the runtime has already interned (E, PI, now, NaN, name, type,
value, length, message, url, method, headers, body, …) the PIC never primes: every store costs ≈ 5,000
instructions more than the same store to an unaffected name.

Reproduction

bench.ts (20 lines):

// Static-key store `o.<name> = v` on a receiver whose class is not statically known
const variant = process.argv[2] || "sha_fields_E"; const N = Number(process.argv[3] || "2000000");
abstract class ShaE { protected abstract A: number; protected abstract B: number; protected abstract C: number; protected abstract D: number; protected abstract E: number;
  protected set(A: number, B: number, C: number, D: number, E: number): void { this.A = A; this.B = B; this.C = C; this.D = D; this.E = E; }
  step(i: number): number { const { A, B, C, D, E } = this; this.set((E + i) | 0, A, B, C, D); return this.A; } }
class Sha256E extends ShaE { protected A = 1; protected B = 2; protected C = 3; protected D = 4; protected E = 5; constructor() { super(); } }
abstract class ShaX { protected abstract A: number; protected abstract B: number; protected abstract C: number; protected abstract D: number; protected abstract X: number;
  protected set(A: number, B: number, C: number, D: number, X: number): void { this.A = A; this.B = B; this.C = C; this.D = D; this.X = X; }
  step(i: number): number { const { A, B, C, D, X } = this; this.set((X + i) | 0, A, B, C, D); return this.A; } }
class Sha256X extends ShaX { protected A = 1; protected B = 2; protected C = 3; protected D = 4; protected X = 5; constructor() { super(); } }
const sE = new Sha256E(), sX = new Sha256X();
const objs: any[] = [{ name: 0, type: 0, nam: 0, typ: 0 }, { name: 1, type: 1, nam: 1, typ: 1 }];
const V: Record<string, (n: number) => number> = {
  sha_fields_E(n) { let a = 0; for (let i = 0; i < n; i++) a = (a + sE.step(i)) % 1000000007; return a; },
  sha_fields_X(n) { let a = 0; for (let i = 0; i < n; i++) a = (a + sX.step(i)) % 1000000007; return a; },
  store_name_type(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 1]; o.name = i; o.type = i; a += o.name & 1; } return a; },
  store_nam_typ(n) { let a = 0; for (let i = 0; i < n; i++) { const o = objs[i & 1]; o.nam = i; o.typ = i; a += o.nam & 1; } return a; },
};
V[variant](N / 5 | 0); const t0 = performance.now(); const cs = V[variant](N);
console.log(`variant=${variant} checksum=${cs} ms=${(performance.now() - t0).toFixed(2)}`);
PERRY_NO_AUTO_OPTIMIZE=1 perry compile bench.ts -o bench
for v in sha_fields_X sha_fields_E store_nam_typ store_name_type; do node bench.ts $v 2000000; ./bench $v 2000000; done

Measurements

Median of 3, shared host (loaded; instruction counts are the load-independent figure). N = 2,000,000; Perry
instructions per iteration = (whole-process instructions:u − 38 M startup) / 2.4 M.

variant Node loop ms Perry loop ms ratio Perry instructions (per iter) Node wall ms Perry wall ms
sha_fields_X (control: 5th state word named X) 37.7 202.9 5.4× 2.61 G (1,070) 164 299
sha_fields_E (noble shape: 5th state word named E) 37.6 1,757.2 47× 14.73 G (6,120) 185 2,181
store_nam_typ (control: o.nam = i; o.typ = i) 2.9 59.0 20× 1.03 G (410) 86 94
store_name_type (o.name = i; o.type = i) 3.0 2,042.7 680× 24.00 G (9,980) 116 2,622

Checksums identical. The two SHA variants differ only in one identifier; the one extra slow store is ≈ 5,050
instructions. Renaming name/type to nam/typ removes ≈ 4,800 instructions per store.

Name sweep (one function o.<k> = i per name on an object literal that already has every key, N = 300,000;
instructions per call including ≈ 350 of loop/call overhead):

  • slow (≈ 5,100–5,500): length, name, value, type, message, key, next, size, headers, url,
    method, body, start, flags, text, source, buffer, offset, E, PI, LN2, SQRT2, NaN,
    Infinity, now.
  • fast (≈ 350): x, done, data, id, status, code, error, result, index, count, end, pos,
    kind, parent, input, callback, options, ttl, state, target, A–D, F–H, e, MAX_VALUE,
    EPSILON, UTC.
  • Reads are not affected (a read of this.E from the base method was as fast as this.A), and neither are stores where the receiver's class is statically known (a method on
    the class that declares the field lowers to a direct slot store).

Impact

  • @noble/hashes 2.2.0 SHA-256 (sha2.js:43: SHA2_32B.set(A, …, H) writes this.E = E | 0 on the _SHA256
    subclass instance, once per compression block; legacy.js:40 has the same store): the numeric-group report (v0.5.1587) attributed ≈ 27 % of sha256's
    Perry CPU to this property bucket (set() 14.7 % inclusive) and reproduced the gap as "4 fields 2.7×, 5 fields 28×".
  • Every package that stores to one of the slow names on a dynamically-typed receiver: err.message = …,
    this.name = …, config.headers = …, config.url/method/body = …, node.flags = …, this.type/value/length = …
    (axios, pg, mysql2, typescript all do this in hot paths; not separately attributed by the audit).

Mechanism

  • js_put_value_set_ic_miss (crates/perry-runtime/src/proxy/put_value.rs:374) runs the store and then declines to
    prime unless the key string's GcHeader has GC_FLAG_INTERNED set and GC_FLAG_FORWARDED clear
    (put_value.rs:444-445) (verified). Without a prime, the codegen PIC (crates/perry-codegen/src/expr/proxy_reflect.rs:465,
    identical IR for this.E and this.X, checked with --trace llvm) misses on every execution.
  • gdb at js_put_value_set_ic_miss (verified): the key pointer passed for E and now has gc_flags = 0x02; for x
    and UTC it has gc_flags = 0x12 (GC_FLAG_INTERNED = 0x10, crates/perry-runtime/src/gc/types.rs:1145).
  • Codegen materializes every string-pool literal with js_string_from_bytes
    (crates/perry-codegen/src/codegen/string_pool.rs:442-447), which does not intern (verified). The runtime interns
    the names it spells itself through canonical_key / intern_ascii_literal → intern_dispatch_bytes
    (crates/perry-runtime/src/string/mod.rs:210, mod.rs:446, crates/perry-runtime/src/string/intern.rs), which
    flags the runtime's own copy (verified). The slow set matches names the runtime installs or dispatches on — builtin
    property names, Math constants, now (inferred).
  • (inferred) When the program's pooled copy of such a name is later canonicalized, the intern table already holds the
    runtime's copy, so the lookup returns that pointer and never flags the pooled one (js_string_intern,
    string/intern.rs:58-99, sets GC_FLAG_INTERNED only on a table miss). The PIC call site keeps passing the
    pooled, unflagged copy, so the prime check fails forever. Names the runtime has not pre-interned are inserted and
    flagged on first use, which is why x/nam/X hit.

What fast looks like

The write PIC should prime on any key whose content is canonical, not on the identity of the pooled copy: either
intern string-pool literals at module init and store the canonical pointer in the handle global, or have the miss
handler prime with (and codegen compare against) the canonical pointer. Targets: sha_fields_E equal to
sha_fields_X (≈ 1,070 instructions per iteration, from 6,120); store_name_type equal to store_nam_typ; every name
in the sweep at ≈ 350 instructions per call. A regression test can assert the prime for a store to name.

Notes

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    package-auditFound by the 2026 package audit: compiling real npm packages from source instead of native bindings
    on Sep 17, 2026
  2. proggeramlug commented on Sep 30, 2026

    @proggeramlug
    ContributorAuthor

    Current target (2026-09-30): no longer reproduces on main; needs a regression test and a close decision

    Package impact (#11464)

    #11464 files this issue under bucket row 5, string ops / key transcoding: 5.1% of equal-weight excess, ≥5% in 12 packages (per-package shares are on #10753). The data does not isolate this mechanism, a write PIC that never primes for runtime-interned key names. Only two js_put_value_set chains in that bucket name a site, and both have unrelated leaves:

    • big.js big.mjs:393 → own_set_descriptor
    • validator merge.js:11 → js_to_primitive

    The issue's own impact case, @noble/hashes, is not in the package set. So the package attribution holds at bucket level only.

    What landed since the issue was filed

    No PR references this issue. The check this issue describes is still in put_value.rs (GC_FLAG_INTERNED, lines 579–580). Even so, the slow store does not reproduce on main, so the store no longer depends on that prime. I did not bisect which change fixed it.

    Reproducer, fresh numbers (the issue's bench.ts, instructions per iteration)

    variant Perry, 7661bc0 (issue) Perry, main Node
    sha_fields_X (control) 1,070 469 112
    sha_fields_E 6,120 469 124
    store_nam_typ (control) 410 82 46
    store_name_type 9,980 82 40

    The issue's column was computed as whole-process instructions ÷ 1.2·N, and main's column uses two-N. The collapse of the name gap does not depend on that method difference.

    The name effect is gone: E equals X, and name/type equals nam/typ, to the instruction. The remaining 2–4× over Node is generic store/call cost. It is the same for both names, so it is not this issue.

    Acceptance target

    • The issue's own target is met: sha_fields_E = sha_fields_X and store_name_type = store_nam_typ.
    • Still to do before closing: the regression test the issue asked for, which pins that a store to name (and E, length, now) costs the same as a store to nam, so the pooled-literal vs interned-copy split cannot come back silently.
    • As a package guard, run the hash-shaped workloads: --filter node-forge --filter jsonwebtoken.

    Package check (Linux, needs perf). Build with cargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static, then run:

    (cd benchmarks/packages && npm ci --ignore-scripts)
    python3 scripts/package_bench.py compile --perry-bin-dir /tmp/pb --filter node-forge --filter jsonwebtoken
    python3 scripts/package_bench.py run --perry-bin-dir /tmp/pb --arms node,perry --modes instr --filter node-forge --filter jsonwebtoken --out /tmp/pb/instr.json

    Run it once on a base-commit build and once on the branch, and compare instructions per iteration. --filter is a workload-id substring, and control/* always runs. For attribution, run profile --callgraph on PERRY_KEEP_SYMBOLS=1 binaries (see benchmarks/packages/PROFILE.md).

    Fresh numbers were measured on origin/main 5fbc2c3 (v0.5.1654) with cargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static. Binaries were compiled with PERRY_NO_AUTO_OPTIMIZE=1 and compared with Node 26.5.1 on Linux x86-64 using perf stat -e instructions:u. Each figure is the median of 3 runs at two sizes of N, with per-op = ΔI/ΔN. N1 is at least 200k operations (20k calls for the factory bench), which keeps most of Node's JIT warm-up inside the constant term. Output was byte-identical to Node on every row. The #11464 figures come from benchmarks/packages/profile/callgraph.{md,json}, measured at Perry 36420d2 with auto-optimize. node-forge was measured from 2febf42 binaries. "% of excess" means the share of a package's (Perry − Node) instructions per iteration.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    package-auditFound by the 2026 package audit: compiling real npm packages from source instead of native bindingsperformanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions