Repository navigation
perf: obj.method() through the runtime dispatcher is ~3,000× slower than Node (two O(own-keys) string-compare scans per call before any cache) #10502
Description
Activity
- addedperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performancepackage-auditFound by the 2026 package audit: compiling real npm packages from source instead of native bindingsFound by the 2026 package audit: compiling real npm packages from source instead of native bindings
on Sep 17, 2026 The acceptance matrix puts a number on the split: 4737 on an object literal, 92 on a class
From the ONE PATH acceptance matrix (
OBJECT_MODEL_SINGLE_PATH_DESIGN_2026-09-20.md§C1.4),
234 cells on v0.5.1631, perrymaster x86-64. Marginalinstructions:u, min of 3, fitted
500k→5M, every cell output-identical to node, no| 0anywhere (#10897), every receiver
escaping into a module-level array the program reads, and every cell disassembled under
PERRY_KEEP_SYMBOLS=1to confirm the call is really inside the loop.Body is
o.d = k; h += o.m()withm() { return this.a }— the keeper store is there so the
call cannot be hoisted. Columns are how the receiver was constructed, rows are where its
binding lives.provenance object literal class field init class ctor-param Object.createfactory function-local const4739.00 152.00 153.00 5071.28 4739.00 local, escapes by return 4739.00 92.00 99.00 5072.28 4739.00 module-level const4874.00 153.00 153.00 5064.28 4874.00 function parameter 4876.00 1311.00 2058.00 5067.22 4876.00 thisin its own method4982.00 156.00 156.00 5187.32 4982.00 varying ( objs[k & 7])5073.00 1372.00 2119.00 5132.83 4938.00 closure-captured 4746.00 156.00 156.00 5072.30 4746.00 min 92.00 · max 5187.32 · max/min 56.38× · median 4739.00 · node median 13.22 ·
median/node 358.47×Two separable defects, and the matrix separates them
1. The construction decides, by 51×. The same call, on the same field, with the same
body, costs 92–156 when the method came from aclassand 4739–5073 when it came
from an object literal or a factory returning one.Object.createis worse again at
5064–5187. Verified linear by hand rather than trusted to the fit — 500k → 2.369e9,
1M → 4.738e9, 5M → 2.369e10 onmethod/local/lit— so it is a real per-call cost, not a
fixture artefact. This is consistent with this issue's "two O(own-keys) string-compare
scans": what the matrix adds is that a class receiver escapes those scans entirely while
a literal never does, which makes the dispatcher cost a construction property.2. On a class receiver, the provenance decides, by 14× — and only for this operation.
92–156 for local / module-const /this/ captured, then 1311 (field init) and 2058
(ctor-param assignment) for a function parameter, 1372 / 2119 for a varying receiver. Every
other operation in the matrix (read1,read4,overwrite,addkey,inherited) moves
2–5× across provenance; this one moves 14×, and it splits exactly on whether the compiler saw
the allocation — i.e. it is the direct-call proof falling off, not the shape.Notably
cctoris consistently worse thancfieldin that column (2058 vs 1311, 2119 vs
1372) for the identical class and field set, which matches #10777's landed fix firing on
a: number = 1.5and not onconstructor(a){ this.a = a }.Why this is worth a row here rather than a new issue
The 4737 figure does not appear in any
ops4row, inmainwatch, or in any fixture this
campaign has quoted — every method-call fixture on the campaign used a class. It is the
largest single number in the matrix by an order of magnitude and it is this issue's
mechanism, measured across the axis that shows when it fires.Harness and raw rows:
secret-tests/matrix/(gen.py,run.py,validate.py,
report.py,results-1631.tsv,MATRIX-v0.5.1631.md); on perrymaster at/root/matrix.
Reproduce one cell withpython3 gen.py fx && perry build fx/method__local_esc__cfield__num.ts.Reproducibility note: pin the runtime link mode
The figures above were measured with the fixtures linking an auto-optimized runtime. If
you reproduce withPERRY_NO_AUTO_OPTIMIZE=1(which the acceptance harness now pins, so the
numbers in/root/matrix/results-pinned-*.tsvare the prebuilt-runtime ones) every cell is
+0 to +7.6% higher, uniformly and additively:cell auto-optimized prebuilt methodon an object literal4739.00 4944.00 methodon a class, escaping local92.00 97.00 read1on anObject.createreceiver322.00 327.00 read1on a literal, module-const binding112.00 112.00 this.ain an object-literal method224.00 230.00 this.ain a class method118.00 123.00 The ratios this issue is about are unchanged (51× / 2.9× / 1.87×) — the shift is a
constant few instructions, not a change in the effect.Worth recording separately, because it cost this lane a false alarm: an unpinned A/B across
two worktrees produced 132 "regressed" cells between v0.5.1631 and v0.5.1632 purely
because one tree had atarget/perry-auto-*archive lying around and the other did not
(fixture binaries 12.7 MB / 12,249 symbols against 19.8 MB / 17,037 symbols). Re-run
like-for-like with the mode pinned, the same two compilers are 0.00% on all 234 cells.
Any A/B on object operations has to pin this or it is measuring the runtime build mode.Isolating the matrix's largest delta: 2183 → 177 from a type annotation, and it is two different cliffs, not one
Follow-up to the method grid above. All arms built
PERRY_KEEP_SYMBOLS=1+
PERRY_NO_AUTO_OPTIMIZE=1on841b605c9(v0.5.1632); marginalinstructions:u, min of 3,
fitted 500k→5M;perf record -e instructions:ufor the attribution. Body is
o.d = k; h += o.m()withm() { return this.a }.cell marginal delta method / param / cctor / **any**177.00 — method / param / cctor / **num**2183.00 12.3× method / param / cfield / num1344.00 7.6× method / local / cctor / num158.00 (provenance present) method / local / cctor / any94.00 method / param / lit / num5081.00 The static test refutes the predicted form
The prediction was that the
: numberarm's loop carries a by-id dispatch helper while the
: anyarm carries a direct call. The two loops are structurally identical — same length
(457 instructions), same call set, call for call:both: js_native_call_method_by_id, js_native_call_value, js_object_get_own_field_or_undef, js_implicit_this_set ×2, js_put_value_set_ic_{miss,overflow_store,poly_tail}, js_gc_loop_safepoint ×2, js_rel_lt, js_write_barrier_root_heap_word ×4, perry_method_…_C__mNothing is added and nothing is swapped. The 2,006-instruction difference is the same code
doing 12× the work at runtime, so it cannot be found by diffing disassembly — which is why
this one needed a profile.Where the 2,006 instructions go (
method / param / cctor / num)The
: anyarm spends 82% in its own compiledrun$spec_b_band 11% in the compiled
method — it stays in compiled code. The: numberarm spends 1.88% in its own compiled
code. Everything else is runtime, and it is one chain, per call:stage % of total ≈ instr/iter symbols typed-layout guard, checked and failing 14.0% ~306 class_field_fast_contract,verify_typed_intact_enabled,class_field_raw_f64_layout_contract,js_typed_feedback_class_field_get_guard"what kind of object is this?" probe cascade 32.3% ~705 is_async_resource_handle,is_closure_ptr,is_date_cell_addr,is_anon_shape_class_id,is_temporal_cell_addr,is_arguments_object,is_registered_buffer,lookup_typed_array_kind,typed_array_addr_from_value,closure_get_dynamic_prop,closure_dynamic_prop_by_key,try_async_resource_property_dispatch,async_resource_propertygeneric by-name field read 43.4% ~948 get_field_by_name_past_inherited_cache,get_field_by_name_object_tail,try_data_get_bytes,OwnedStringBytes::copy_of_key(8.41%),core::str::from_utf8(2.49%),js_object_get_field,inherited_read_cache_lookupown-keys materialisation 2.4% ~52 object_keys_arrayrecording the miss 2.0% ~43 typed_feedback::record_fallback_callrecord_fallback_callat ~43 instructions per iteration is the tell: the guard does not
miss once and settle — it misses every call, forever. And ~184 instructions per iteration
go tocopy_of_key+from_utf8: the key string is copied and UTF-8-validated on every
single field access.So the answer to "one call or a scan" is neither: it is a cliff off the inline
class-field path into the complete generic by-name read, and the single largest term is the
receiver-kind probe cascade (~705), not the lookup itself.It is the field read inside the method body, not the dispatch.
js_native_call_method_by_id
is in both arms, so the call is by-id either way; what degrades isthis.aoncethiscame
from a receiver the caller could not prove.The
litarm is a different cliff, and this is the one that is literally this issuemethod / param / lit / num(5081) profiles completely differently:18.11% __memcmp_evex_movbe 14.75% js_array_get_f64 7.10% string::compare::js_string_key_matches_bytes 5.98% native_call_method::handle_methods::dispatch_handle 5.51% js_native_call_method 3.10% native_call_method::class_vtable_fast_guardThat is method-name string comparison — this issue's "two O(own-keys) string-compare
scans", 25% of the run inmemcmp+js_string_key_matches_bytes. The typed-class-via-parameter
case never reaches it; its dispatch is fine and its field read is what collapses.Two mechanisms, both landing here: an object-literal receiver pays a by-name scan on the
method; a typed class reached through a parameter pays the generic by-name path on the
field inside the method. A fix for one will not move the other.The control behaves as the source reading predicts
method / local / cctor— same class, same method, receiver has provenance — emits 2
js_typed_feedback_*guard calls in thenumarm and 0 in theanyarm, and costs
158 vs 94. So the guard is emitted only for the typed class, and where the receiver is
proven it succeeds: +64, of which ~31 is the store-side
js_array_numeric_value_to_raw_f64already isolated on #10907. The annotation only becomes
catastrophic when the guard it introduces cannot be satisfied — through a parameter, 12.3×.Why this is posted here rather than on #10907
#10907 is the store-side
js_array_numeric_value_to_raw_f64cost (+31 per typed store),
which is a different mechanism and a bounded one. The method row is this issue's generic
dispatch/by-name path; the type annotation is only the trigger that routes an ordinary
constructor(a: number) { this.a = a }class onto it. Recording it here so the method work
starts from the attribution rather than rediscovering it. Diagnosis only — no fix proposed.Artefacts:
/root/matrix/probe/on perrymaster (all arms,PERRY_KEEP_SYMBOLS=1), perf data
at/tmp/p_method__param__cctor__{num,any}.dataand/tmp/p_lit.data.Pricing the receiver-kind cascade across the matrix: it is not a general tax — it is confined to the method cliff
Asked whether the ~705-instruction probe cascade is flat across every generic access (which
would price charter step 1 for the whole language). It is not.perf record -e instructions:uon841b605c9,PERRY_KEEP_SYMBOLS=1+PERRY_NO_AUTO_OPTIMIZE=1; the
percentage columns are shares of that cell's own profile, the instruction columns are
share × the cell's marginal.cell marginal cascade cascade instr key-copy instr by-name instr guard instr in compiled code read1 / varying / lit180.00 0.3% 1 1 1 0 93.8% read1 / varying / cctor180.00 0.3% 1 0 1 0 93.8% read1 / param / lit114.00 0.0% 0 1 1 0 90.5% read1 / param / cctor114.00 0.5% 1 0 1 0 90.5% overwrite / varying / lit144.00 0.8% 1 1 1 0 92.3% read1 / local / ocreate327.00 0.0% 0 0 3 0 53.8% inherited / local / ocreate497.00 0.0% 0 1 176 0 17.4% inherited / varying / ocreate568.00 0.0% 0 1 163 0 39.5% addkey / local / lit448.00 0.4% 2 0 0 0 45.5% addkey / varying / lit513.00 0.4% 2 2 0 0 23.4% method / param / cctor2183.00 36.4% 794 294 584 405 1.2% method / param / lit5081.00 7.8% 399 845 164 365 0.4% 1. The cascade is 0–2 instructions on every non-method cell
Ten cells spanning own reads, own writes, varying receivers, parameter receivers,
Object.createreceivers, inherited reads and key addition: ≤ 2 instructions each, ≤ 0.8%.
The two method cells are the only place it appears, at 794 and 399.So step 1 (native handles as real objects) does not buy ~700 on every generic access. On
this evidence it buys ~0 on ordinary property access and ~400–800 per method call on the
two shapes that fall off the compiled path. That is a real payoff, but a narrow one, and it
should be carried into the charter as per method call on a degraded receiver, not per
access. Worth noting the cells that look "generic" mostly are not:read1/varyingand
read1/paramare 90–94% in compiled code — they are inline-cached reads, not runtime
calls, and there is no cascade to remove.2. The key copy + UTF-8 validation is not general either
copy_of_key/from_utf8/js_string_key_matches_bytes/memcmptotal 0–2
instructions on every ordinary cell — includinginherited, which is a genuine by-name
prototype-chain read doing 163–176 instructions of by-name work with essentially zero key
copying. It is 294 and 845 instructions on the two method cells.So the per-access key copy is a property of the method path, not of miss-to-generic reads.
Checked #10761 and #10753 before concluding; neither is the right home and no new issue is
warranted — it is already reported above as part of this issue's method attribution.Bonus: what the
inheritedandaddkeycells actually spend, since it is not the cascadeBoth are dominated by overflow / spill storage, which the cascade question surfaced by
elimination:inherited/local/ocreate (497) 35.3% inherited_read_cache_lookup 20.6% object::spill::overflow_set 12.2% gc::layout::layout_note_slot 6.5% js_put_value_set_ic_overflow_store 17.4% (its own compiled code) addkey/local/lit (448) 45.5% (its own compiled code) 18.0% object::spill::overflow_set 11.4% object::spill::overflow_get 9.2% js_put_value_set_ic_overflow_store 6.0% gc::layout::layout_note_slot 3.0% field_get_set::ic_slow::overflow_armAn
Object.create()receiver's own properties and a newly added key both land in overflow
storage rather than inline slots, so every subsequent access to them is
spill::overflow_{get,set}plus alayout_note_slot. That names the mechanism behind
#10905 (the flat 3–4×Object.createtax on own reads and writes) and behind the
addkeyrow, and it is a side table — i.e. it prices charter step 3, not step 1. Posting
theObject.createhalf on #10905 as well.Diagnosis only; nothing built, nothing proposed.
Package-level measurement from the Phase 3 attribution (#11464,
benchmarks/packages/PROFILE.md):origin/main36420d2, release compiler, auto-optimize, Linux x86-64,perf record -e instructions:u --call-graph dwarfat two N (per-iteration, startup cancels), Node 26.5.1 oracle with outputs checked equal. "% of excess" = share of (Perry − Node) instructions/iter, equal-weight over the 23 packages with Perry/Node ≥ 2×.The dispatcher is the largest construct-level cost in real packages. The runtime work under
js_typed_feedback_native_call_method_by_id, up to the next JS frame (so excluding the JS it calls), is 18.7% of total excess across 16 packages: node-forge/hmac 82%, node-forge/aes_cbc 72%, qs/stringify_nested 57%, qs/parse_nested 56%, axios/get_json 47%, dayjs/diff_startof 46%.Inside it, the largest mechanism is not the own-key scan described here. It is the own-override check:
resolve_own_user_method(8.3% of excess) →js_object_has_own(7.3%) →is_function_prototype_object_value(5.4%) →builtin_prototype_value("Function"), which looks upglobalThis.Functionby name on every dispatched call (the #10497 mechanism). The rest is the usualshape_descriptor_by_id/dispatch_handle/from_utf8leaves.Top sites: node-forge
util.jsByteStringBuffer methods;side-channel/index.js:31$channelData.get(key)andqs/lib/stringify.js:99tmpSc.get(object)(object-literal methods named get/set/has, #10506);axios/lib/utils.js:15hasOwnProperty.call(obj, prop).- added a commit that references this issue
on Sep 27, 2026 #11489 moves
recv.m(args)off the runtime dispatcher for object-literal,Object.create, ES5-prototype, factory and class receivers (≈3,000–10,000 → ≈90–97 instr/call; tsc −12.8%, Zod −1.2%). Remaining for #10502:
(1) Function-object receivers (F.m(), ≈2,500/call) still use the dispatcher. They are handled in the every-shape follow-up (site entry bit 61 is reserved for them).
(2) Primitive receivers ("s".slice(),n.toFixed()) intentionally stay on the dispatcher for now.
(3) A hit costs ≈90 instructions vs node's ≈12. Thethissave/restore through TLS goes away with the this-as-parameter lane; the rest is call spills.
(4) Class methods do not yet become real prototype slots (decision D4, after the class-constructor work).
(5) A site that sees one shape with many different method bodies latches to the dispatcher (matrixvaryingcell ≈6,700).- added a commit that references this issue
on Oct 3, 2026 - added 11 commits that reference this issue
on Oct 4, 2026
Found by the package performance audit (real npm packages compiled from source, profiled against Node 26.5.1) and
re-measured on Perry 7661bc0 (v0.5.1589), Linux x64. Every method call that reaches
js_native_call_methodscansevery own key of the receiver with a byte compare in
class_vtable_fast_guard, and, when that does not dispatch,scans them all again in
dispatch_handle— so a call on a 64-property object literal costs ~22,800 instructions(~2 µs, ~3,000× Node), growing ~270 instructions per own key.
Reproduction
ns.ts(thens64literal is written out in full so it keeps its static shape):Measurements
Median of 3, N = 2,000,000, shared host (load average 100–300 on 64 threads; instruction counts are the
load-independent figure). Node inlines these calls (~0.7 ns each), so the loop ratios mostly measure Perry's absolute
per-call cost.
ns64.f40(x)— 64 own propsns4.f40(x)— 4 own propsf40(x)through a local (control)k64.m(x)class, 64 fields, direct call (control)k2.m(x)class, 2 fields, direct call (control)Checksums identical. The same closure costs 215 instructions called through a local and 22,799 called as
ns64.f40(); going from 4 to 64 own keys adds ~17,100 instructions (~285 per key). Class instances with 64 fields arefine when codegen emits the guarded direct call (161 instructions); the cost appears whenever the call falls to
the runtime tower (see #10503 for one codegen trigger).
perf recordofns64:js_array_get_f6452.7 % self,js_string_key_matches_bytes15.1 %,__memcmp6.7 %;inclusive
dispatch_handle45.9 %,class_vtable_fast_guard38.7 %.Impact
From the audit profiles (v0.5.1587):
axios.getagainst a local server, 9.7× Node): dynamic method dispatch onutils$1.*— a63-key object literal (
utils$1.isUndefined(v),utils$1.forEach(...), …) — is 22.6 % of Perry CPU, ~25 % of thePerry-minus-Node time (io-group report).
js_array_get_f64was the Support custom menu bar items #1 self symbol (7.8 %) and is not array code.cheerio.load+ 3 selector queries, 58× Node): 50 % of CPU in the dispatch bucket,js_array_get_f6420 % self, 11 points of it fromclass_vtable_fast_guard(large-group report). Re-profiled on7661bc0 (300 rows):
js_native_call_method93 % inclusive,try_class_vtable_fast_dispatch74 %,js_array_get_f6418 % self with ≥ 8 points called fromclass_vtable_fast_guard. Perry 424 ms/iteration vsNode 6.9 ms (≈ 61×). (On 7661bc0 this bench dies after 2–3 iterations with
TypeError: Cannot read properties of undefined (reading 'xmlMode'), the stale-pointer crash already seen by theaudit, so only short runs were measured.)
methods — see perf:
class X extends EventEmitteris 440× slower than Node to construct and 200–2,400× to use (super() copies 15 bound closures onto every instance) #10508); typescript 5.8transpileModule≈ 7 %, rate-limiter-flexible ≈ 7 %.Mechanism
Call path for a static-name call site that is not proven:
js_typed_feedback_native_call_method_by_id→js_native_call_method.crates/perry-runtime/src/typed_feedback/guards.rs:853-900(verified): the typed-feedback entry records anobservation (shape, name hash) and then always calls the full
js_native_call_method; nothing it records isused to skip the tower on the next call.
crates/perry-runtime/src/object/native_call_method.rs:1292(verified): first step of the tower istry_class_vtable_fast_dispatch→class_vtable_fast_guard(:115). To prove no own field shadows the method itloops over all
logical_key_countown keys, reading each key through the genericjs_array_get(the fullreceiver-classifying
js_array_get_f64) and comparing it withjs_string_key_matches_bytes(:179-184,verified). This runs on the cache hit path too — a hit on a 17-field parse5 Tokenizer still pays 17 key
compares per call.
None, so the tower continues throughthe native-module, disposal, TextDecoder, URLSearchParams (perf: methods named get/set/has/delete/keys/… on ordinary objects are 167–368× slower than Node, 2.3–3.8× slower than other names (URLSearchParams probe per call) #10506), AbortSignal and primitive probes to
handle_methods::dispatch_handle(native_call_method.rs:1980), which scans the same keys again with the samejs_array_get+ byte compare (crates/perry-runtime/src/object/native_call_method/handle_methods.rs:948-957,verified) and finally calls the closure through
js_native_call_value.scans repeat on every call.
What fast looks like
own-property method call on a known shape becomes a slot load + closure-magic check + call, and a prototype/vtable
hit is a ShapeId compare, not a key scan (a ShapeId with no own key of that name proves "no shadow" once per
shape, not once per call).
ns64.f40()andns4.f40()within 2× of thelocalcontrol (≤ 450instructions/call) and independent of own-key count; cheerio's
class_vtable_fast_guard+js_array_get_f64sharebelow 2 %.
Notes
__pshapeunder an inline keys check), Every dynamic method call pays 7 side-registry probes to exclude kinds the program never creates —is_registered_symbolalone takes a mutex + SipHash (6.5% ofpipeline) #7850 (probecascade), perf(runtime): classify native-call receivers from the tracked header; no Buffer/typed-array registry probe for known kinds #9937 (open draft: receiver classification from the tracked header — does not remove the scans).
this.m()on a class whose constructor adds fields inside anifis ~490× slower than Node (exact birth-ShapeId guard misses every call) #10503 and perf: oneObject.setPrototypeOf/util.inheritson a plain object makes every class method call in the process 5× slower, 155–189× Node (sticky global latch retires all direct-call guards) #10504 (two triggers that sendthis.m()to this tower), perf: method calls on untyped receivers (prototype methods,fn.call,pushonany) are 44–520× slower than Node (no call-site cache; name re-resolved per call) #10505 (prototype methods /.call/pushon untyped receivers), perf: methods named get/set/has/delete/keys/… on ordinary objects are 167–368× slower than Node, 2.3–3.8× slower than other names (URLSearchParams probe per call) #10506 (URLSearchParams probe in the same tower), perf:class X extends EventEmitteris 440× slower than Node to construct and 200–2,400× to use (super() copies 15 bound closures onto every instance) #10508.mb_this_v4.js(parse5-shaped Tokenizer) no longer reaches the tower on 7661bc0 — its calls aredirect now and its remaining 50× is by-name field access — but real parse5 in cheerio still does (profile above).
Which condition sends parse5's calls there was not isolated. Ruled out on 7661bc0: construction from another
module, a constructor closure capturing
this, fields assigned outside the constructor, and the prototype-surgerylatch from perf: one
Object.setPrototypeOf/util.inheritson a plain object makes every class method call in the process 5× slower, 155–189× Node (sticky global latch retires all direct-call guards) #10504 (gdb: not armed in the cheerio+undici build).