Repository navigation
perf: stores to name/type/value/length/url/E… are 47–680× slower than Node (key literal not flagged interned, so the write IC never primes) #10500
Description
Activity
- addedperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performancepackage-auditFound by the 2026 package audit: compiling real npm packages from source instead of native bindingsFound by the 2026 package audit: compiling real npm packages from source instead of native bindings
on Sep 17, 2026 Current target (2026-09-30): no longer reproduces on main; needs a regression test and a close decision
Package impact (#11464)
#11464 files this issue under bucket row 5, string ops / key transcoding: 5.1% of equal-weight excess, ≥5% in 12 packages (per-package shares are on #10753). The data does not isolate this mechanism, a write PIC that never primes for runtime-interned key names. Only two
js_put_value_setchains in that bucket name a site, and both have unrelated leaves:- big.js
big.mjs:393→own_set_descriptor - validator
merge.js:11→js_to_primitive
The issue's own impact case, @noble/hashes, is not in the package set. So the package attribution holds at bucket level only.
What landed since the issue was filed
No PR references this issue. The check this issue describes is still in
put_value.rs(GC_FLAG_INTERNED, lines 579–580). Even so, the slow store does not reproduce on main, so the store no longer depends on that prime. I did not bisect which change fixed it.Reproducer, fresh numbers (the issue's
bench.ts, instructions per iteration)variant Perry, 7661bc0 (issue) Perry, main Node sha_fields_X(control)1,070 469 112 sha_fields_E6,120 469 124 store_nam_typ(control)410 82 46 store_name_type9,980 82 40 The issue's column was computed as whole-process instructions ÷ 1.2·N, and main's column uses two-N. The collapse of the name gap does not depend on that method difference.
The name effect is gone:
EequalsX, andname/typeequalsnam/typ, to the instruction. The remaining 2–4× over Node is generic store/call cost. It is the same for both names, so it is not this issue.Acceptance target
- The issue's own target is met:
sha_fields_E=sha_fields_Xandstore_name_type=store_nam_typ. - Still to do before closing: the regression test the issue asked for, which pins that a store to
name(andE,length,now) costs the same as a store tonam, so the pooled-literal vs interned-copy split cannot come back silently. - As a package guard, run the hash-shaped workloads:
--filter node-forge --filter jsonwebtoken.
Package check (Linux, needs
perf). Build withcargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static, then run:(cd benchmarks/packages && npm ci --ignore-scripts) python3 scripts/package_bench.py compile --perry-bin-dir /tmp/pb --filter node-forge --filter jsonwebtoken python3 scripts/package_bench.py run --perry-bin-dir /tmp/pb --arms node,perry --modes instr --filter node-forge --filter jsonwebtoken --out /tmp/pb/instr.jsonRun it once on a base-commit build and once on the branch, and compare instructions per iteration.
--filteris a workload-id substring, andcontrol/*always runs. For attribution, runprofile --callgraphonPERRY_KEEP_SYMBOLS=1binaries (seebenchmarks/packages/PROFILE.md).Fresh numbers were measured on
origin/main5fbc2c3 (v0.5.1654) withcargo build --release -p perry -p perry-runtime-static -p perry-stdlib-static. Binaries were compiled withPERRY_NO_AUTO_OPTIMIZE=1and compared with Node 26.5.1 on Linux x86-64 usingperf stat -e instructions:u. Each figure is the median of 3 runs at two sizes of N, with per-op = ΔI/ΔN. N1 is at least 200k operations (20k calls for the factory bench), which keeps most of Node's JIT warm-up inside the constant term. Output was byte-identical to Node on every row. The #11464 figures come frombenchmarks/packages/profile/callgraph.{md,json}, measured at Perry 36420d2 with auto-optimize. node-forge was measured from 2febf42 binaries. "% of excess" means the share of a package's (Perry − Node) instructions per iteration.- big.js
- added a commit that references this issue
on Sep 30, 2026 - added a commit that references this issue
on Sep 30, 2026
Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. The audit reported an "IC cliff at ≥ 5 subclass fields" in
@noble/hashes; reduction shows the cliff is not the field count but the field NAME: SHA-256's fifth state word is
called
E. A static-key storeo.<k> = von a receiver whose layout is not statically known overwrites an existingslot through the write PIC, but for keys the runtime has already interned (
E,PI,now,NaN,name,type,value,length,message,url,method,headers,body, …) the PIC never primes: every store costs ≈ 5,000instructions more than the same store to an unaffected name.
Reproduction
bench.ts(20 lines):Measurements
Median of 3, shared host (loaded; instruction counts are the load-independent figure). N = 2,000,000; Perry
instructions per iteration = (whole-process
instructions:u− 38 M startup) / 2.4 M.sha_fields_X(control: 5th state word namedX)sha_fields_E(noble shape: 5th state word namedE)store_nam_typ(control:o.nam = i; o.typ = i)store_name_type(o.name = i; o.type = i)Checksums identical. The two SHA variants differ only in one identifier; the one extra slow store is ≈ 5,050
instructions. Renaming
name/typetonam/typremoves ≈ 4,800 instructions per store.Name sweep (one function
o.<k> = iper name on an object literal that already has every key, N = 300,000;instructions per call including ≈ 350 of loop/call overhead):
length,name,value,type,message,key,next,size,headers,url,method,body,start,flags,text,source,buffer,offset,E,PI,LN2,SQRT2,NaN,Infinity,now.x,done,data,id,status,code,error,result,index,count,end,pos,kind,parent,input,callback,options,ttl,state,target,A–D,F–H,e,MAX_VALUE,EPSILON,UTC.this.Efrom the base method was as fast asthis.A), and neither are stores where the receiver's class is statically known (a method onthe class that declares the field lowers to a direct slot store).
Impact
sha2.js:43:SHA2_32B.set(A, …, H)writesthis.E = E | 0on the_SHA256subclass instance, once per compression block;
legacy.js:40has the same store): the numeric-group report (v0.5.1587) attributed ≈ 27 % of sha256'sPerry CPU to this property bucket (
set()14.7 % inclusive) and reproduced the gap as "4 fields 2.7×, 5 fields 28×".err.message = …,this.name = …,config.headers = …,config.url/method/body = …,node.flags = …,this.type/value/length = …(axios, pg, mysql2, typescript all do this in hot paths; not separately attributed by the audit).
Mechanism
js_put_value_set_ic_miss(crates/perry-runtime/src/proxy/put_value.rs:374) runs the store and then declines toprime unless the key string's
GcHeaderhasGC_FLAG_INTERNEDset andGC_FLAG_FORWARDEDclear(
put_value.rs:444-445) (verified). Without a prime, the codegen PIC (crates/perry-codegen/src/expr/proxy_reflect.rs:465,identical IR for
this.Eandthis.X, checked with--trace llvm) misses on every execution.js_put_value_set_ic_miss(verified): the key pointer passed forEandnowhasgc_flags = 0x02; forxand
UTCit hasgc_flags = 0x12(GC_FLAG_INTERNED = 0x10,crates/perry-runtime/src/gc/types.rs:1145).js_string_from_bytes(
crates/perry-codegen/src/codegen/string_pool.rs:442-447), which does not intern (verified). The runtime internsthe names it spells itself through
canonical_key/intern_ascii_literal→intern_dispatch_bytes(
crates/perry-runtime/src/string/mod.rs:210,mod.rs:446,crates/perry-runtime/src/string/intern.rs), whichflags the runtime's own copy (verified). The slow set matches names the runtime installs or dispatches on — builtin
property names,
Mathconstants,now(inferred).runtime's copy, so the lookup returns that pointer and never flags the pooled one (
js_string_intern,string/intern.rs:58-99, setsGC_FLAG_INTERNEDonly on a table miss). The PIC call site keeps passing thepooled, unflagged copy, so the prime check fails forever. Names the runtime has not pre-interned are inserted and
flagged on first use, which is why
x/nam/Xhit.What fast looks like
The write PIC should prime on any key whose content is canonical, not on the identity of the pooled copy: either
intern string-pool literals at module init and store the canonical pointer in the handle global, or have the miss
handler prime with (and codegen compare against) the canonical pointer. Targets:
sha_fields_Eequal tosha_fields_X(≈ 1,070 instructions per iteration, from 6,120);store_name_typeequal tostore_nam_typ; every namein the sweep at ≈ 350 instructions per call. A regression test can assert the prime for a store to
name.Notes
A..H); a 6-field subclass writingFis fast, a 5-field subclass writing
Qis fast, and a 6-field subclass writingEat index 5 is slow.o.k = vis ~100× slower than Node (static-key write IC primes only overwrites; no add-transition cache) #10496).