Repository navigation
perf(regex): RegExp.prototype.test costs ~1.5 µs per call under Perex — 44x slower than the old engine even with the regex hoisted #10166
Description
Activity
- addedbugConfirmed defect or regressionConfirmed defect or regressionperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performance
on Sep 13, 2026 Cost placement: per-call candidates identified, not yet ranked
Each
.test()call currently does all of the following. Reading source lists the work; it cannot rank it — that needs a profile, which is still the first acceptance item here:- a property
Getofexec(the spec's observable lookup) BoundProgram::new: full program revalidationBoundSubject::new: an O(n) subject decode (small for these short strings; see perf(regex): split, replace, matchAll and global exec do quadratic work under Perex — split regressed from ~7x Node to timing out #10165)MatchBuffersallocation- several GC safepoint polls
Search::new
Proposed direction (Perex and Perry together), for per-call loops written in JavaScript: reuse validation across calls. Perex would rebind from a witness of an earlier successful validation, checked by length and header (the same constant-time check
with_viewalready does). Perry would cache that witness on theRegExpand on the string, plus alastIndex→ byte-offset hint for non-ASCII subjects so a JS-levelexecloop does not re-seek either. The contract and its GC details need maintainer sign-off before any implementation.Status: placement from reading source, by the Perex maintainer and checked line by line against perry
mainb5a82cfeaeand the publishedperex0.1.0 crate. It is not yet measured. Ownership and scheduling are awaiting a maintainer decision, so please do not start source changes yet.Related: #10164 (seek charge), #10165 (per-operation rebinding).
- a property
Status: not resolved by what has landed. #10183 (train 180, #10195) makes a RegExp's program and subject rebind in constant work across calls. That removed the cross-call quadratic in JS
exec/matchAll/testloops on long subjects (see #10165), but it barely moves the short-string per-call cost this issue is about.Load-independent instruction counts (
perf stat -e instructions:u, spread under 0.1 %): 1,000,000 calls on strings like"record_123"/"!bad_123", minus the program that only builds the strings.per call before #10183 after change hoisted re.test(v),/^[a-z]+_[0-9]+$/24,571 23,634 −3.8 % literal in loop 25,967 25,022 −3.6 % re.exec(v)+m[2]41,184 39,835 −3.3 % Hoisted and literal stay within about 6 % of each other, so this remains a per-call cost, not construction. About 23.6k instructions per hoisted
testmeans subject binding and program validation were a small share. The first acceptance item still applies: record how the per-call time divides between setup, binding/positioning, matching and result construction. Wall clock on perrymaster is about 1.5–2.8 s per million calls, depending on load; Node takes 15–75 ms.Same-host rerun of the full reproducers, as the acceptance list asks.
Setup: perrymaster (Linux x86_64, 16 logical CPUs, shared with other lanes; load average 11–23 during the runs), Node v26.8.1,
hostile-benchmarks/runtime/run.py --filter regex, release builds from source.- Fix:
mainat5d3bf85f9(0.5.1554), which includes perf(regex): bind once per split/replace/match, and split searches forward (#10165) #10174, fix(regex): do not cap RegExp operations by work (#10164) #10176, perf(regex): resume searches and capture reads from the previous position (#10164) #10181, perf(regex): bind a RegExp's program and subject in constant work across calls (#10166) #10183 and perex 0.1.4 (deps(regex): take perex 0.1.4 — required-text search starts at the requested start #10201). - Pre-Perex:
9495bfc95. - Method: the two arms ran back to back. Each cell is the median of 7 samples after warm-up, with the ratio to Node in parentheses. The harness gives each size a 60 s process budget covering at least 5 warm-up runs plus 7 samples.
Verdict: open (per-call work pending)
The per-call cost is unchanged by the landed work: it is linear but 39–135× Node, against pre-Perex's 2.7–5.4×. The per-call breakdown (Perry side and Perex side) is with Ralph. Perex 0.1.5 (unpublished) makes an anchored miss return without trying later starts: 2,726 → 1,031 instructions per search on the
"!bad_…"case.regex-test-literaln Node pre-Perex 9495bfc95main5d3bf85f9checksums (main) 100 0.004 ms 0.01 ms (2.7×) 0.25 ms (65.2×) = Node 1,000 0.040 ms 0.11 ms (2.8×) 2.90 ms (72.9×) = Node 10,000 0.411 ms 1.12 ms (2.7×) 27.94 ms (68.0×) = Node 100,000 2.112 ms 11.25 ms (5.3×) 285.63 ms (135.2×) = Node 1,000,000 37.791 ms 112.32 ms (3.0×) 1,905.90 ms (50.4×) = Node Log-log slope over 1k–100k: Node 0.863,
main0.997 (delta +0.134), pre-Perex 1.003.regex-test-literal-unicoden Node pre-Perex 9495bfc95main5d3bf85f9checksums (main) 100 0.004 ms 0.01 ms (2.7×) 0.24 ms (61.1×) = Node 1,000 0.041 ms 0.11 ms (2.7×) 2.77 ms (66.8×) = Node 10,000 0.401 ms 1.15 ms (2.9×) 15.57 ms (38.9×) = Node 100,000 4.022 ms 11.63 ms (2.9×) 157.62 ms (39.2×) = Node 1,000,000 21.164 ms 114.78 ms (5.4×) 2,013.81 ms (95.2×) = Node Log-log slope over 1k–100k: Node 0.993,
main0.878 (delta -0.116), pre-Perex 1.007.- Fix:
- added 3 commits that reference this issue
on Sep 13, 2026 Remeasured on current
main— 2.9× faster than when this was filed, still 5–6× the pre-Perex ratioSame-host, back-to-back, all three arms built from source and run sequentially on perrymaster (AMD Ryzen 7 7700X, 16 logical cores), Node v26.8.1,
--no-auto-optimize,TZ=UTC LC_ALL=en_US.UTF-8.- main
92eadb77ab(v0.5.1569) — includes perf(regex): cache compiled programs and speed up test/replace #10193, perf(regex): RegExp.prototype.test costs ~1.5 µs per call under Perex — 44x slower than the old engine even with the regex hoisted #10166's own exec-lookup/scratch/capture work (12983d4ad5), and the five generic levers from the breakdown below (perf(runtime): fast-path ordinary native property Get #10248, perf(runtime): reduce handle scope overhead with cached TLS metadata #10252, perf(gc): make the "nothing due" GC check cheap on safepoint polls and trigger checks #10253, perf(runtime): reduce JS catch-frame setup overhead #10257, perf(runtime): store array named properties with the array (brief 4 of #10166) #10263). - pre-Perex
9495bfc95(perry 0.5.1531), rebuilt for this run. - Load average at start 4.3; Node's own numbers land within 2 % of the ones in the issue body, so the host was as quiet as the original sweep.
regex-test-literaln Node pre-Perex (× Node) main (× Node) main ÷ pre-Perex checksums 100 0.0021 ms 0.010 ms (4.8×) 0.054 ms (26.3×) 5.4× = Node 1,000 0.0208 ms 0.103 ms (4.9×) 0.584 ms (28.0×) 5.7× = Node 10,000 0.2039 ms 1.094 ms (5.4×) 5.687 ms (27.9×) 5.2× = Node 100,000 2.0128 ms 11.122 ms (5.5×) 56.639 ms (28.1×) 5.1× = Node 1,000,000 20.1780 ms 110.424 ms (5.5×) 587.641 ms (29.1×) 5.3× = Node Log-log slopes: Node 0.997, pre-Perex 1.013, main 1.006 — all linear.
regex-test-literal-unicoden Node pre-Perex (× Node) main (× Node) main ÷ pre-Perex checksums 100 0.0021 ms 0.010 ms (4.8×) 0.059 ms (28.5×) 5.9× = Node 1,000 0.0213 ms 0.110 ms (5.2×) 0.639 ms (30.1×) 5.8× = Node 10,000 0.2098 ms 1.113 ms (5.3×) 6.558 ms (31.3×) 5.9× = Node 100,000 2.0292 ms 11.092 ms (5.5×) 65.548 ms (32.3×) 5.9× = Node 1,000,000 19.8748 ms 112.931 ms (5.7×) 651.449 ms (32.8×) 5.8× = Node Log-log slopes: Node 0.994, pre-Perex 1.010, main 1.009.
Against the numbers in the issue body (72–79× Node), main is now 26–33× Node: 2.4–2.8× faster. Acceptance asks for the pre-Perex ratio (4.8–5.7×); main is 5.1–5.9× that. Not met.
Hoisted-versus-literal control (the standalone program, 1M calls, one timed pass)
probe Node pre-Perex main issue body (Perex, 0.5.1545) test literal-in-loop 1M73.2 ms 94.9 ms 548.3 ms 1,603.7 ms test hoisted regex 1M13.6 ms 32.5 ms 471.9 ms 1,509.6 ms Hoisted and literal are 16 % apart on main, so this is still per-call cost, not construction — acceptance item 3 holds. Same program, for context on what Perex bought:
exec global ascii n=10000is 50.1 ms on main vs 1,116.9 ms pre-Perex, andreplace cb unicode n=10000is 35.8 ms vs 715.8 ms. The regression is confined to fixed per-call overhead on short subjects; the scanning workloads are 20× better than the old engine.Acceptance item 1: where the per-call cost goes
Load-independent instruction counts,
perf stat -e instructions:u, three runs per arm, spread under 0.01 %. Each probe is 1,000,000 calls; the per-call figure subtracts that arm's own control program, which builds the same million strings and does no regex.per call Node pre-Perex main main ÷ Node main ÷ pre-Perex vs. 2026-09-13 hoisted re.test(v)475 592 8,315 17.5× 14.1× −64.8 % (was 23,634) literal in loop 578 2,055 9,631 16.6× 4.7× −61.5 % (was 25,022) re.exec(v)+m[2]824 4,638 11,479 13.9× 2.5× −71.2 % (was 39,835) Attribution for the hoisted
test, fromperf record -e instructions:uon a--debug-symbolsbuild of the same probe (verified to execute within 0.001 % of the stripped binary's instruction count). Percentages are of the whole probe; the instructions/call column scales them by the 9,629 instructions each loop iteration costs including string building.phase what is in it % ≈ instr/call per-call scratch setup MatchBuffers::new(7.9 %) + thememcpyit andfind_neardo (20.3 %)28.2 % 2,710 matching perex::executor::Vm::{run, initialize, trial, fixed_atom_charge, atom_scan, sought, atom_holds}23.8 % 2,295 binding / dispatch / entry find_nearself (7.2 %),execute_output,bind_program,perex_dispatch::{execute, test_string}, plusmemcmp(6.2 %) reached from dispatch andfind_near18.9 % 1,820 GC bookkeeping at safepoints from_space_in_use_bytes,influx_driven_nursery_cap_bytes,gc_budgeted_*,ValidPointerSet::contains,layout_note_slot, minus the share the control program shows for the same symbols~4.5 % ~440 exception frame setup exception::try_push_with_kind+perry_sjlj_try3.8 % 370 property lookup js_object_get_field,object::spill::overflow_get1.6 % 155 result construction nothing — testreturns a boolean and no capture-materialization symbol appears in the profile0 % 0 (string building) the probe's own control work: js_string_concat_value*,main13.7 % 1,314 Roughly 95 % of the program is accounted for, and the answer to the question this issue opened with is: only about a quarter of a
.test()call is matching. A third is per-call scratch construction, and there is no result-construction cost at all fortest.The largest single item is concrete.
find_nearbuilds a freshMatchBuffersper call and passes it by value intoSearch::new;MatchBuffersis 360 bytes (itsSlots<usize, INLINE_REGISTERS = 32>inline register array) andSearchis 664. The hottestmemcpycall site in the profile isfind_near+0x361, disassembling tomov $0x150,%edx; call memcpy— a 336-byte struct move on every call, with a second one insideMatchBuffers::new. Nothing about that scratch depends on the subject: for a given program it is the same shape on every call.Verdict: still open
- The gap is no longer dominated by any one thing. Even eliminating the whole scratch-setup bucket leaves ~5,600 instructions per call, against pre-Perex's 592 and Node's 475. Reaching the acceptance ratio needs both a reusable per-call search state and a cheaper engine entry (
Vm::initializealone is 5.8 %, ~560 instructions, for a 10-byte anchored match). - Next lever on the Perry side, if the maintainers want it: hold the
MatchBuffers(and ideally the wholeSearchshape) on the bound program and rebind it per call instead of constructing and moving it, so atestloop stops paying ~2,700 instructions of setup. This needs the Perex side's agreement on the borrow/reentrancy contract — a nested regex call inside a replacer callback must not see a scratch another search is using — which is why I am posting it rather than implementing it. - Nothing in this comment changed behaviour: all three arms are unmodified source, and every checksum matches Node at every size.
- main
- added a commit that references this issue
on Sep 15, 2026 Prototype: lending the thread's scratch removes 42 % of a
.test()callBranch
proggeramlug/perry:perf/10166-lent-scratch(93f49cf876), no PR — it pins perex through[patch.crates-io]at gitf0b1b0b(0.1.6). 0.1.6 is now on crates.io, but the 7-day soak window in.cargo/config.toml(global-min-publish-age = "7 days") does not admit it until 2026-09-22, so landing needs a one-time publish-age override, which is Ralph's call.perex 0.1.6 adds
impl ScratchOwner for &mut O, so a host can lend scratch to aSearchinstead of giving it up. That answers the 336-byte-per-call move in the attribution above directly:find_nearbuilt aMatchBuffersper call whose 32-register inline array was zeroed and then moved by value intoSearch. One cell per thread now serves every search, borrowed for the duration of one search.The cell also keeps whatever frames/undo length an earlier call needed, which turned out to be the other half of the cost —
PERRY_REGEX_DIAGcounted a scratch growth on nearly every search, so each call was rebuffering as well as constructing. In steady state a loop now grows nothing and constructs nothing.Instructions per call
1M-call probes minus their own string-building control, perrymaster, three runs per arm, spread under 0.01 %. Non-ASCII rows use
/^[ä中😀]+_[0-9]+$/uover"ä中😀_<i>".per call Node pre-Perex main 92eadb77abthis branch change × Node hoisted test, ASCII475 592 8,315 4,836 −41.8 % 17.5× → 10.2× literal in loop, ASCII 578 2,055 9,631 6,151 −36.1 % 16.6× → 10.6× exec+m[2], ASCII824 4,638 11,479 8,016 −30.2 % 13.9× → 9.7× hoisted test, non-ASCII295 604 9,119 5,715 −37.3 % 30.9× → 19.4× literal in loop, non-ASCII 520 2,067 10,435 7,031 −32.6 % 20.1× → 13.5× Full reproducers
Median ms per run, same host, back to back, checksums = Node at every size.
n Node main this branch regex-test-literal1,000,000 20.909 ms 573.805 ms (27.4×) 380.681 ms (18.2×) 100,000 2.073 ms 58.606 ms (28.3×) 36.866 ms (17.8×) regex-test-literal-unicode1,000,000 20.500 ms 662.350 ms (32.3×) 455.405 ms (22.2×) 100,000 2.078 ms 66.223 ms (31.9×) 45.353 ms (21.8×) The standalone control moves with it:
test hoisted regex 1M471.9 → 280.6 ms,test literal-in-loop 1M548.3 → 363.8 ms (Node 13.6 / 73.2).Semantics and safety
- Reentrancy is a runtime borrow rather than the compile-time one perex gives within a single frame: a nested regex — a replacer callback that matches, or a poll that re-enters — finds the cell borrowed and takes the owned path, so two searches never share slots.
- Work charging is unchanged per call. A search that asks for more frames or undo entries than the cell holds grows the cell for the next call and lets this call run the owned path from the budget it entered on, so the work a single call charges is what it charges today.
- The operation's memory limit still sees the slots. The lent path takes a
Chargefor the registers, frames and undo it lends, exactly as the owner it replaces did; the thread keeps the memory, the operation only borrows it. - No GC pointer is stored in the cell — registers are subject offsets, frames and undo are the engine's opaque scratch, as in the owned buffers.
INLINE_REGISTERSdrops 32 → 8 for the owned path, which is now only a fallback; the lent cell keeps 32 underLENT_REGISTERS, since it is allocated once per thread rather than moved per call.cargo test -p perry-runtime --lib -- --test-threads=1: 3916 passed, 1 failed —native_stack::tests::stack_top_respects_custom_thread_stack_sizes, which fails on bare main on this host too.cargo fmt --checkclean; clippy adds nothing beyond main's 12 pre-existingapprox_constanterrors.
Still open
This is one of the two levers perex 0.1.6 offers. The other,
Search::restart_at, applies to the three Perry loops that search one subject repeatedly —perex_split.rs:206,perex_remove.rs:82, and the replace fast path's span collection atperex_replace_direct.rs:213— and is not in this branch. Even with both, a hoistedtestwould sit around 10× Node against pre-Perex's 1.25×, so the acceptance bar here still needs a cheaper engine entry:Vm::initializeis ~560 instructions per call for a 10-byte anchored match.- added a commit that references this issue
on Sep 15, 2026 Correction: the −41.8 % above is two changes, not one
My previous comment attributed the whole drop to lending the scratch. The prototype branch pins perex 0.1.6 through
[patch.crates-io], andmainresolves perex 0.1.4, so those numbers also carry 0.1.5's start-anchored change (a program anchored at the subject's start tries only its first start). Splitting the two with a middle arm —mainwith the same perex pin and no lent scratch, same host, deterministic builds:per call main (perex 0.1.4) main + perex 0.1.6 + lent scratch perex 0.1.6's share lending's share hoisted test, ASCII8,315 7,356 4,835 −959 (−11.5 %) −2,521 (−34.3 %) literal in loop, ASCII 9,631 8,672 6,151 −959 (−10.0 %) −2,521 (−29.1 %) exec+m[2], ASCII11,479 10,524 8,016 −955 (−8.3 %) −2,508 (−23.8 %) hoisted test, non-ASCII9,119 8,251 5,715 −868 (−9.5 %) −2,536 (−30.7 %) literal in loop, non-ASCII 10,435 9,567 7,031 −868 (−8.3 %) −2,536 (−26.5 %) So the engine bump is worth a flat ~870–960 instructions per call and lending is worth a flat ~2,510–2,540 on top of it. Lending is still the larger lever by roughly three to one, but it is −34.3 % on the hoisted probe rather than −41.8 %; the combined figure is what the earlier table reported.
The flatness is itself informative: both savings are per-call constants, independent of the subject, which is what a fixed-overhead problem looks like.
- added 2 commits that reference this issue
on Sep 16, 2026 The GC safepoint poll is worth 10.5 % of a
.test()call — and that is the ceilingRecording a measurement so it does not live only in session chat. The per-call path runs one
gc_runtime_safepoint_pollbefore each search (find_near_lent, beforeSearch::new). A profile put it at ~18 % of samples; this prices it by removing it.Deliberately unsafe probe, never a PR: the
poll()?beforeSearch::new/new_neardeleted. Baseline and variant built from the same commit (e6dcb6274d) with the same flags, on the same host.Arm 1 — the hoisted
.test()loop (1,000,000 calls, string-building control subtracted,instructions:u):instructions per call baseline 4,792 poll removed 4,290 saving 502 (−10.5 %) Both arms print the same answer.
Arm 2 —
regex-replace-callbackat n=700,000, an allocating workload:baseline poll removed ms per run 6,084 6,264 cycle_starts189 189 steps1,474,773 1,474,817 share_permille674 680 mark_barrier_armed_us33,354,846 33,494,684 checksum 203210458 203210458 No win, every counter within half a percent. That workload allocates constantly, so the allocation-site assists already carry the GC stepping and the regex poll is redundant there. The poll earns its cost only in a loop that allocates nothing.
What this does NOT establish
Neither arm exercises the hazard. The
.test()probe reportscycle_starts=0on both arms — there is never an open budgeted cycle during it, so nothing could have been left unstepped. In a non-allocating loop the pre-search poll may be the only safepoint on the path, and an open incremental cycle that stops being stepped keeps its mark barrier armed, taxing every write in the program. On the replace workload above,mark_barrier_armed_usis 33 s against a 73 s wall, so that failure mode is expensive when it happens.A witness for it needs a third shape that neither arm has: build a large live set, get an old-generation cycle open, then run a long non-allocating regex loop and assert the cycle still completes and the barrier does not stay armed across it.
What it means for the options
- Removing the poll outright is off until that witness exists.
- A stride (poll every Nth search) recovers a fraction of the 502 proportional to the stride, and needs the witness to choose N.
- Memoising
gc_budgeted_due_trigger_evalacross polls can only ever recover part of the 502, and its cost is the invalidation funnel: every input must bump an epoch through one enforced path, with a per-input test that fails when its bump is removed. One of those inputs isexternal_side_old_reclaim_pressure_bytes, whose update path was wrong until perf(gc): drained-side-bytes pressure (#10268) triggers unproductive old-gen cycles — +33% on a regex loop, GC 0.3% -> 19.3% of wall #10376 closed.
So the whole question is worth at most 502 instructions per call against the 4,792 a hoisted
.test()costs today — set against Node's 475 and the pre-Perex engine's 592.Open-cycle witness: the pre-search poll does no stepping — and my safety objection does not hold up
I argued the pre-search poll was load-bearing: in a loop that allocates nothing it might be the only safepoint, so removing it could strand an open budgeted cycle with its mark barrier armed. The witness says otherwise.
Building it took three designs, and the first two proved nothing
Worth recording, because each failed its own subject-live assertion rather than producing a misleading green:
- Retain a large live set, then loop.
cycle_starts=0— no cycle ever opened. The old-reclaim trigger follows allocation churn, not a static heap. 8M retained strings still gavecycle_starts=0. - Churn first, loop after. 11 starts, 11 completions, and the two arms identical — every cycle finished inside the churn, before the loop began, because the churn's own allocation safepoints stepped them.
- Interleave. Each round does a global string replace (known to open cycles) and then a long chunk of
.test()calls that allocate nothing. A chunk that begins with a cycle open can only be stepped by the pre-search poll.
Result
Baseline versus the no-poll probe, same commit (
e6dcb6274d) and flags, 4 rounds of 12,000,000 non-allocating.test()calls each — several seconds of allocation-free work per round, 48,000,000 calls in total:cycle_startscompletionsstepsmark_barrier_armed_usbaseline 6 6 5,286 301,691 poll removed 6 6 5,286 304,911 Identical starts, identical completions, identical step counts, and armed time within 1 %. The 8-round × 2M variant reads the same way (15/15, 39,828 steps on both arms).
The step counts being exactly equal is the finding. If the poll were stepping cycles, removing it would change that number. It does not, so in these programs the poll performs no stepping at all: the
#10253due-check answers "nothing due" and returns, and every actual step comes from an allocation-site assist.What this does not establish
- That a cycle ever spanned a non-allocating window.
mark_barrier_armed_usis ~50 ms per cycle against chunks lasting seconds, so the cycles here complete inside the allocation phase. The hazard is therefore unproven, not disproven — I could not construct a program that holds a cycle open across a long allocation-free stretch, and without mid-run observation I cannot tell whether that is because the runtime prevents it or because my setup never achieved it. - Cancellation. The poll also returns
EngineError::Cancelled. Nothing here tests whether a long regex loop stays interruptible without it. That is a separate witness and it matters for any option that skips polls.
Where that leaves the options
The safety objection I raised against removing the poll is not supported by evidence. That makes a stride (option 1b) low-risk on the stepping axis and leaves cancellation as the open question. It also reframes the decision as purely economic: the whole poll is 502 instructions of the 4,792 a hoisted
.test()costs, and memoising the due-check can only recover part of that while requiring an invalidation funnel where every input bumps an epoch through one enforced path.Probe branch is local and was never pushed; the witness source is
open-cycle-witness.tsin my reproducer set, and it carries its assertions in comments so the next person does not repeat designs 1 and 2.- Retain a large live set, then loop.
Cancellation: the production poll cannot cancel, so it is not an argument for keeping it
The remaining safety question was whether a long regex loop stays interruptible without the pre-search poll. It does not need to: nothing in production ever cancels through it.
host::pollis the only poll the runtime passes into the engine:pub(crate) fn poll() -> Result<(), EngineError> { crate::gc::gc_runtime_safepoint_poll(); Ok(()) }
It returns
Ok(())unconditionally. AndEngineError::Cancelledhas no producer anywhere inperry-runtimeoutside tests — every construction of it is ingc/tests/runtime_roots/perex_{split,glob,strings,execution}.rs, where a test supplies its own cancelling closure. Production carries only the enum variant and its message mapping inperex_api.rs:71.That plumbing is not dead weight: perex's contract lets a host cancel, and those tests prove Perry's paths release scratch and preserve consumed work when one does. But since no production path cancels, cancellation cannot be a reason to keep the per-call poll.
Following
gc_runtime_safepoint_pollthrough confirms the rest of the contract is equally narrow:gc_runtime_safepoint_poll -> gc_runtime_safepoint_report -> gc_budgeted_step_work_units_inner_with_progress(...) // report discardedNo user code, no throw, no interrupt delivery. Stepping a budgeted GC is the entire effect.
The poll's complete account
can it cancel? No — no production producer of Cancelleddoes it step cycles in practice? No — identical step counts with and without it (5,286 on both arms over 48M allocation-free calls) what does it cost? 502 instructions per call, 10.5 % of the 4,792 a hoisted .test()costswhat does it buy? the option of stepping a due GC from a loop that allocates nothing That last row is the whole remaining case for it, and it is a real one — it is the mechanism by which a mutator that is not allocating still participates in an incremental collection. I could not build a program where it mattered (three designs, described in the previous comment), but "I could not construct it" is not "it cannot happen", and the person who owns the pacing should weigh that rather than take my failure to trigger it as proof.
What has changed is that both of my earlier objections are gone. Removing or striding the poll is not blocked by stepping and not blocked by cancellation; it is a judgement about whether 502 instructions per call is worth the loss of a participation point that nothing currently exercises.
Current number, so the parked decision is not made against a stale one
Not new analysis — the bar question on this issue still needs a decision rather than more measurement. But the figure in the title has moved by about half since it was written, and the decision reads differently at the new one.
A hoisted
.test()of/^[a-z]+_[0-9]+$/over 1,000,000 short subjects, instructions per call with a control binary subtracted (same array build, same loop, trivial predicate in place of the regex):instructions/call vs Node 26.5.1 Node 26.5.1 873.6 1.00x Perry, current main 4,382.2 5.02x Perry, with #10580 4,063.7 4.65x The title's "~1.5 µs per call, 44x slower than the old engine" predates perex 0.1.4 -> 0.1.9, the per-thread lent match scratch (#10372), the native replace pieces (#10412), the strided pre-search GC poll (#10494) and the one-shot search entry (#10580, open).
Instruction counts rather than wall time on purpose: the measuring host carried other sessions' builds throughout, and round-to-round spread reached 68-244% against effects of 0.5-25%, enough to invert the sign of a known-good result. Instruction counts are load-insensitive; the ratio above is the one I would stand behind, and a microsecond figure is not currently measurable here.
For the bar itself, my read is unchanged and still on this issue: it should be restated as a ratio to Node, because the pre-Perex engine it names was ~20x worse on scanning work, so "1.25x of the old engine" is a moving and mostly meaningless target. At 4.65x Node the gap is real but it is no longer the order-of-magnitude story the title tells.
Package-level measurement from the Phase 3 attribution (#11464,
benchmarks/packages/PROFILE.md):origin/main36420d2, release compiler, auto-optimize, Linux x86-64,perf record -e instructions:u --call-graph dwarfat two N (per-iteration, startup cancels), Node 26.5.1 oracle with outputs checked equal. "% of excess" = share of (Perry − Node) instructions/iter, equal-weight over the 23 packages with Perry/Node ≥ 2×.Regex is the second-largest bucket: 17.1% of total excess, ≥5% in 11 packages. Perex frames are on the stack for 18.4% of excess (≥2% in 14 packages). About 73% of the bucket is the matcher itself (
Vm::run,trial,atom_scan,begin_class,Cursor::next_unit); ~19% is result building; ~5% is spec-protocol property gets (#10518); ~3% is compilation.Per-package share of excess: dotenv 80%, node-cron 74%, uuid 53%, jsonwebtoken 41%, nanoid 35%, date-fns 29%, validator 24%, commander 17%, decimal.js 14%.
Top sites:
jws/lib/verify-stream.js:41JWS_REGEX.test(string)(60% of jsonwebtoken/decode); uuidvalidateREGEX.test;validator/lib/isByteLength.jsencodeURI(str).split(/%..|./)(11.5% of validator/batch, the #10165 split path);decimal.jsstr.search(/e/i), where 8.3% of decimal.js/parse_sum issymbol::is_registered_symbol_slow(the@@searchlookup, #10518 family); dayjsString.replacewith a callback, which sets up anexception::try_push_with_kindper callback call.Progress: #11548 landed perex 0.1.11 (the PerryTS/perex#3 matcher), under the owner-approved one-time publish-age override. Instructions: jsonwebtoken decode −61%, uuid v4 −48%, v7 −34%, v5_parse −28%, nanoid −30%, uuid/jws regex micro −65%/−84%; commander +0.4%; RSS flat. Together with #11543 this is the planned matcher + host work; leaving it open for the remaining gap to Node.
What happened
A million
.test()calls against short strings take about 1.5 seconds on currentmain, a flat per-call cost at every input size. On the pre-Perex runtime the same loop took about a tenth of a second. The regression is the same whether the regex literal sits inside the loop or is hoisted into a singleRegExpcreated once, so it is not explained by losing literal-site or compilation caching: the cost is in each match call. Checksums match Node at every size.Measured against Node
v26.8.1using Perryperry 0.5.1545at8a058e205385ec8ebb353ce0f231530871ce4a49. This is evidence from that pinned revision, not a claim that current main was remeasured. First reproduce on current main; if it is already fixed, identify the fixing commit and attach the comparison.Measurements
Times are median milliseconds per workload invocation. Ratios are Perry/Node. A correctness or timeout classification takes precedence over performance; successful smaller-size timings on those rows are diagnostic evidence.
regex-test-literal— SEVERELog(time)/log(n) least-squares slopes: Perry 1.002, Node 0.997, delta 0.005.
Workload: Regex baseline, excluded from the main ranking. Literal is intentionally inside the loop; mixed valid/invalid seeded records prevent a constant all-true checksum.
regex-test-literal-unicode— SEVERELog(time)/log(n) least-squares slopes: Perry 1.003, Node 0.985, delta 0.018.
Workload: Unicode regex baseline, excluded from the main ranking. Subject and pattern contain umlauts, CJK and/or emoji, using the Unicode flag. n counts records; non-ASCII /g indices are UTF-16 offsets, not UTF-8 byte offsets. Checksums consume test outcomes, actual match contents/indices, or replacement/split outputs. Literal is intentionally inside the loop; mixed valid/invalid seeded records prevent a constant all-true checksum.
What is expected / acceptance criteria
.test()on a short ASCII string. (Recorded 2026-09-14 on92eadb77ab: scratch setup ~2,710, matching ~2,295, binding/dispatch ~1,820, GC ~440, exception frame ~370, property lookup ~155, result construction 0, of 8,315 instructions per call.).test()measurement so a construction fix cannot be mistaken for a per-call fix.teston agoryregex reads and advanceslastIndexexactly asexecdoes, aRegExpsubclass or patchedexecis observed, and ToString coercion of a non-string argument still runs.Implementation to inspect
Same-host before/after. The "old engine" column is the pre-Perex runtime at
9495bfc95eand the "Perex" column ismainat8a058e2053, both built from source with the identical release command and measured sequentially on the same host with the same Node. Perex became the runtime's only regex engine in #10142 (landed via merge train #10146); #10149 then resolved it from crates.io asperex = "0.1".Before/after, full suite workloads (median ms per invocation, same host):
regex-test-literalregex-test-literalregex-test-literalregex-test-literalregex-test-literalregex-test-literal-unicoderegex-test-literal-unicoderegex-test-literal-unicoderegex-test-literal-unicoderegex-test-literal-unicoderegex-test-literal: Perry slope 1.01 → 1.00 (SLOW → SEVERE), Node 1.00regex-test-literal-unicode: Perry slope 1.01 → 1.00 (SLOW → SEVERE), Node 0.99Hoisted control, from a standalone program (one timed pass of 1M calls; source below). The first row is the benchmark's shape, a literal evaluated inside the loop; the second creates the regex once:
test literal-in-loop 1Mtest hoisted regex 1MThe hoisted row is the decisive one. If the regression were construction — for example the removal of the old engine's compile caches in #10142 — hoisting would recover most of it. Instead the hoisted loop is as slow as the literal loop. A literal-site path still exists at
site_test.rs(lookup/installkeyed by site), which is consistent with construction not being the cost.Where the per-call time goes has not been located. Candidates to measure, not conclusions: per-call setup of the work and memory budgets and a runtime handle scope, binding the subject and positioning the span reader, and result or capture materialization that a boolean
testdoes not need. As with the DataView setter work in #10089, the first step should be an attribution measurement recorded on this issue before any change.Standalone program
The benchmark metadata comments embedded below still name the pre-Perex runtime functions they were written against; #10142 deleted those files. "Implementation to inspect" lists the code paths that exist at the measured revision.
js_regexp_testlookupSource reading narrows the investigation; it does not establish exclusive runtime/compiler attribution. No compiler or runtime changes were made to obtain these measurements.
Agent scope and coordination
These three issues are owned by the Perex maintainers, who have taken them and are confirming cost attribution before splitting ownership. Do not start a source fix without coordinating on the issue first, so work does not collide. Changes belong in the Perex crate and/or the runtime adapter under
crates/perry-runtime/src/regex/perex_*. The acceptance numbers should be re-measured against the same pre-Perex revision on one host, as above, not against the original September 11 sweep, which ran on a different machine.Implementation work can proceed in separate branches. Serialize benchmark runs on any shared host; parallel timing runs invalidate small performance comparisons. Preserve language semantics and moving-GC safety.
Coordinate with this benchmark task: bug(regex): split and replace throw "Regular expression work limit exceeded" on 32,000-unit strings Node handles in under a millisecond #10164
Coordinate with this benchmark task: perf(regex): split, replace, matchAll and global exec do quadratic work under Perex — split regressed from ~7x Node to timing out #10165
Related history/context: Make Perex the runtime's only regular-expression engine #10142
Related history/context: Merge train 170: #10142 #10146
Related history/context: Merge train 172: #10148 #10149
Reproduce and remeasure
Everything needed for the workload is embedded below; no private repository, fixture, npm package, or shared prelude is required. Save a complete benchmark block under its indicated filename in
/tmp/perry-builtin-repro/. Use Node 26.8.1 to match this measurement; it runs these TypeScript files directly.From the Perry checkout/branch being evaluated:
To reproduce the historical baseline, use the pinned commit above in a separate checkout and build the compiler and both libraries there. Repeat compilation for each additional benchmark below. Run this small driver from the same checkout, changing
nameandsizesfor that benchmark:Record before/after results from the same unchanged source, engine versions and host. The measured driver uses seeded setup outside timers, at least 200 ms AND five warmup runs, then seven samples with at least 20 ms measured work each. Fresh input is prepared before each timer for mutating workloads. The median per-run time is reported, with checksum consistency checked on every invocation. Timeouts cover setup, warmup and sampling, not just one builtin call.
Environment and limits
Linux-6.17.0-23-generic-x86_64-with-glibc2.39; target: native host.v26.8.1; Perry:perry 0.5.1545; build: release from source.--no-auto-optimize; compiler and both matching runtime archives were rebuilt together.[3.537109375, 4.07470703125, 5.009765625]; end:[3.58251953125, 3.47265625, 4.1953125].Minimal correctness reductions
This issue is a performance workload; the complete checksum-gated reproducer follows.
Complete standalone benchmark sources
regex-test-literal.ts — sizes [100, 1000, 10000, 100000, 1000000]
Size meanings and fresh-input policy are in the leading metadata.
result_on_stderrfor this file:False.regex-test-literal-unicode.ts — sizes [100, 1000, 10000, 100000, 1000000]
Size meanings and fresh-input policy are in the leading metadata.
result_on_stderrfor this file:False.