Skip to content

perf(regex): RegExp.prototype.test costs ~1.5 µs per call under Perex — 44x slower than the old engine even with the regex hoisted #10166

Description

@proggeramlug

What happened

A million .test() calls against short strings take about 1.5 seconds on current main, a flat per-call cost at every input size. On the pre-Perex runtime the same loop took about a tenth of a second. The regression is the same whether the regex literal sits inside the loop or is hoisted into a single RegExp created once, so it is not explained by losing literal-site or compilation caching: the cost is in each match call. Checksums match Node at every size.

Measured against Node v26.8.1 using Perry perry 0.5.1545 at 8a058e205385ec8ebb353ce0f231530871ce4a49. This is evidence from that pinned revision, not a claim that current main was remeasured. First reproduce on current main; if it is already fixed, identify the fixing commit and attach the comparison.

Measurements

Times are median milliseconds per workload invocation. Ratios are Perry/Node. A correctness or timeout classification takes precedence over performance; successful smaller-size timings on those rows are diagnostic evidence.

regex-test-literal — SEVERE

n Node ms / status Perry ms / status ratio Node checksum Perry checksum
100 0.002062 0.153508 74.46× 952009575 952009575
1000 0.021067 1.525797 72.42× 515425129 515425129
10000 0.206593 15.459840 74.83× 248696985 248696985
100000 2.027071 156.154489 77.03× 603020940 603020940
1000000 20.348576 1560.366017 76.68× 36901593 36901593

Log(time)/log(n) least-squares slopes: Perry 1.002, Node 0.997, delta 0.005.

Workload: Regex baseline, excluded from the main ranking. Literal is intentionally inside the loop; mixed valid/invalid seeded records prevent a constant all-true checksum.

regex-test-literal-unicode — SEVERE

n Node ms / status Perry ms / status ratio Node checksum Perry checksum
100 0.002134 0.155277 72.77× 952009575 952009575
1000 0.026292 1.573017 59.83× 515425129 515425129
10000 0.211933 15.949114 75.26× 248696985 248696985
100000 2.067177 160.227477 77.51× 603020940 603020940
1000000 20.259069 1597.342334 78.85× 36901593 36901593

Log(time)/log(n) least-squares slopes: Perry 1.003, Node 0.985, delta 0.018.

Workload: Unicode regex baseline, excluded from the main ranking. Subject and pattern contain umlauts, CJK and/or emoji, using the Unicode flag. n counts records; non-ASCII /g indices are UTF-16 offsets, not UTF-8 byte offsets. Checksums consume test outcomes, actual match contents/indices, or replacement/split outputs. Literal is intentionally inside the loop; mixed valid/invalid seeded records prevent a constant all-true checksum.

What is expected / acceptance criteria

  • Before changing behaviour, measure and record on this issue how the per-call time divides between setup, subject binding/positioning, matching, and result construction for a hoisted .test() on a short ASCII string. (Recorded 2026-09-14 on 92eadb77ab: scratch setup ~2,710, matching ~2,295, binding/dispatch ~1,820, GC ~440, exception frame ~370, property lookup ~155, result construction 0, of 8,315 instructions per call.)
  • Rerun both full reproducers and the standalone hoisted-versus-literal program on the fix and on the pre-Perex revision on one host at every size. Checksums must match Node, and the Perex/Node ratio must return at least to the pre-Perex engine's ratio for both the literal-in-loop and hoisted forms.
  • Keep the hoisted and literal forms within a small constant of each other, and add the hoisted form to any future .test() measurement so a construction fix cannot be mistaken for a per-call fix.
  • Preserve semantics: test on a g or y regex reads and advances lastIndex exactly as exec does, a RegExp subclass or patched exec is observed, and ToString coercion of a non-string argument still runs.

Implementation to inspect

Same-host before/after. The "old engine" column is the pre-Perex runtime at 9495bfc95e and the "Perex" column is main at 8a058e2053, both built from source with the identical release command and measured sequentially on the same host with the same Node. Perex became the runtime's only regex engine in #10142 (landed via merge train #10146); #10149 then resolved it from crates.io as perex = "0.1".

Before/after, full suite workloads (median ms per invocation, same host):

workload n Node old engine Perex old ÷ Node Perex ÷ Node Perex ÷ old
regex-test-literal 100 0.002 0.010 0.154 4.8× 74.5× 15.5×
regex-test-literal 1,000 0.021 0.108 1.526 5.1× 72.4× 14.2×
regex-test-literal 10,000 0.207 1.104 15.460 5.3× 74.8× 14.0×
regex-test-literal 100,000 2.027 11.021 156.154 5.5× 77.0× 14.2×
regex-test-literal 1,000,000 20.349 111.067 1,560.366 5.4× 76.7× 14.0×
regex-test-literal-unicode 100 0.002 0.010 0.155 4.8× 72.8× 15.4×
regex-test-literal-unicode 1,000 0.026 0.117 1.573 5.5× 59.8× 13.5×
regex-test-literal-unicode 10,000 0.212 1.104 15.949 5.3× 75.3× 14.5×
regex-test-literal-unicode 100,000 2.067 11.073 160.227 5.4× 77.5× 14.5×
regex-test-literal-unicode 1,000,000 20.259 111.796 1,597.342 5.5× 78.8× 14.3×
  • regex-test-literal: Perry slope 1.01 → 1.00 (SLOW → SEVERE), Node 1.00
  • regex-test-literal-unicode: Perry slope 1.01 → 1.00 (SLOW → SEVERE), Node 0.99

Hoisted control, from a standalone program (one timed pass of 1M calls; source below). The first row is the benchmark's shape, a literal evaluated inside the loop; the second creates the regex once:

probe Node old engine Perex
test literal-in-loop 1M 67.6 ms 99.1 ms 1,603.7 ms
test hoisted regex 1M 12.8 ms 33.6 ms 1,509.6 ms

The hoisted row is the decisive one. If the regression were construction — for example the removal of the old engine's compile caches in #10142 — hoisting would recover most of it. Instead the hoisted loop is as slow as the literal loop. A literal-site path still exists at site_test.rs (lookup / install keyed by site), which is consistent with construction not being the cost.

Where the per-call time goes has not been located. Candidates to measure, not conclusions: per-call setup of the work and memory budgets and a runtime handle scope, binding the subject and positioning the span reader, and result or capture materialization that a boolean test does not need. As with the DataView setter work in #10089, the first step should be an attribution measurement recorded on this issue before any change.

Standalone program
function t(label: string, f: () => number): void {
  const s = performance.now();
  try { const r = f(); console.log(`${label.padEnd(34)} ok     ${(performance.now() - s).toFixed(1).padStart(9)} ms  result=${r}`); }
  catch (e) { console.log(`${label.padEnd(34)} THREW  ${(performance.now() - s).toFixed(1).padStart(9)} ms  ${String(e)}`); }
}
for (const n of [1000, 2000, 4000, 10000]) {
  t(`split ascii        n=${n}`, () => "ab12,cd345;ef6 ".repeat(n).split(/[,; ]+/).length);
  t(`split unicode      n=${n}`, () => "ä中12,Ö漢345;ef6😀".repeat(n).split(/[,;😀]+/u).length);
  t(`replace cb unicode n=${n}`, () => "ä中😀12 Ö漢🦊345;".repeat(n).replace(/[ä中😀Ö漢🦊]+/gu, (m) => "[" + m + "]").length);
  t(`exec global ascii  n=${n}`, () => { const s = "ab12 cd345;".repeat(n), re = /([a-z]+)([0-9]+)/g; let c = 0; while (re.exec(s) !== null) c++; return c; });
}
const vals: string[] = []; for (let i = 0; i < 1000000; i++) vals.push((i % 2 ? "record_" : "!bad_") + i);
t("test literal-in-loop 1M", () => { let c = 0; for (let i = 0; i < vals.length; i++) if (/^[a-z]+_[0-9]+$/.test(vals[i])) c++; return c; });
const hoisted = /^[a-z]+_[0-9]+$/;
t("test hoisted regex  1M", () => { let c = 0; for (let i = 0; i < vals.length; i++) if (hoisted.test(vals[i])) c++; return c; });

The benchmark metadata comments embedded below still name the pre-Perex runtime functions they were written against; #10142 deleted those files. "Implementation to inspect" lists the code paths that exist at the measured revision.

Source reading narrows the investigation; it does not establish exclusive runtime/compiler attribution. No compiler or runtime changes were made to obtain these measurements.

Agent scope and coordination

These three issues are owned by the Perex maintainers, who have taken them and are confirming cost attribution before splitting ownership. Do not start a source fix without coordinating on the issue first, so work does not collide. Changes belong in the Perex crate and/or the runtime adapter under crates/perry-runtime/src/regex/perex_*. The acceptance numbers should be re-measured against the same pre-Perex revision on one host, as above, not against the original September 11 sweep, which ran on a different machine.

Implementation work can proceed in separate branches. Serialize benchmark runs on any shared host; parallel timing runs invalidate small performance comparisons. Preserve language semantics and moving-GC safety.

Reproduce and remeasure

Everything needed for the workload is embedded below; no private repository, fixture, npm package, or shared prelude is required. Save a complete benchmark block under its indicated filename in /tmp/perry-builtin-repro/. Use Node 26.8.1 to match this measurement; it runs these TypeScript files directly.

From the Perry checkout/branch being evaluated:

mkdir -p /tmp/perry-builtin-repro
cargo build --release --locked -p perry -p perry-runtime-static -p perry-stdlib-static
export PERRY_RUNTIME_DIR="$PWD/target/release"
export TZ=UTC LC_ALL=en_US.UTF-8
git rev-parse HEAD
node --version
target/release/perry compile /tmp/perry-builtin-repro/regex-test-literal.ts --no-auto-optimize -o /tmp/perry-builtin-repro/app

To reproduce the historical baseline, use the pinned commit above in a separate checkout and build the compiler and both libraries there. Repeat compilation for each additional benchmark below. Run this small driver from the same checkout, changing name and sizes for that benchmark:

import json, math, subprocess
name = 'regex-test-literal'
sizes = [100, 1000, 10000, 100000, 1000000]
result_on_stderr = False
stopped = set()
points = {"node": [], "perry": []}
for n in sizes:
    pair = {}
    for engine, cmd in [("node", ["node", f"/tmp/perry-builtin-repro/{name}.ts"]),
                        ("perry", ["/tmp/perry-builtin-repro/app"])]:
        if engine in stopped: continue
        try:
            p = subprocess.run(cmd + [str(n)], capture_output=True, text=True, timeout=60)
        except subprocess.TimeoutExpired:
            print(engine, n, "TIMEOUT"); stopped.add(engine); continue
        if p.returncode:
            print(engine, n, "ERROR", p.returncode, p.stderr, p.stdout); continue
        r = json.loads(p.stderr if result_on_stderr else p.stdout)
        pair[engine] = r
        points[engine].append((n, r["ms_per_run"]))
        print(engine, r)
    if len(pair) == 2:
        print("ratio", n, pair["perry"]["ms_per_run"] / pair["node"]["ms_per_run"],
              "checksum_match", pair["perry"]["checksum"] == pair["node"]["checksum"])
def slope(rows):
    if len(rows) < 2: return None
    x = [math.log(n) for n, t in rows]; y = [math.log(t) for n, t in rows]
    mx = sum(x)/len(x); my = sum(y)/len(y)
    return sum((a-mx)*(b-my) for a,b in zip(x,y))/sum((a-mx)**2 for a in x)
print("slopes", {engine: slope(rows) for engine, rows in points.items()})

Record before/after results from the same unchanged source, engine versions and host. The measured driver uses seeded setup outside timers, at least 200 ms AND five warmup runs, then seven samples with at least 20 ms measured work each. Fresh input is prepared before each timer for mutating workloads. The median per-run time is reported, with checksum consistency checked on every invocation. Timeouts cover setup, warmup and sampling, not just one builtin call.

Environment and limits

  • CPU: AMD Ryzen 7 7700X 8-Core Processor; 16 logical cores; x86_64.
  • OS: Linux-6.17.0-23-generic-x86_64-with-glibc2.39; target: native host.
  • Node: v26.8.1; Perry: perry 0.5.1545; build: release from source.
  • Compile flag: --no-auto-optimize; compiler and both matching runtime archives were rebuilt together.
  • The pinned source revision and compiler/runtime/Node artifact hashes were unchanged throughout the sweep.
  • Load average at measurement start: [3.537109375, 4.07470703125, 5.009765625]; end: [3.58251953125, 3.47265625, 4.1953125].
  • Host contention limits precise constant-factor claims; repeat on a quiet host before asserting an improvement.
  • Timings include timer overhead and checksum calculation. String hashes bound lookup count, not Unicode lookup cost; indexed consumption may also force Node string materialization.

Minimal correctness reductions

This issue is a performance workload; the complete checksum-gated reproducer follows.

Complete standalone benchmark sources

regex-test-literal.ts — sizes [100, 1000, 10000, 100000, 1000000]

Size meanings and fresh-input policy are in the leading metadata. result_on_stderr for this file: False.

// @runtime {"name": "regex-test-literal", "category": "regex", "verification": "checksum", "sources": [{"file": "crates/perry-runtime/src/regex.rs", "function": "js_regexp_test"}], "hypothesis": "Hypothesis: literal-site compilation caching and matcher dispatch determine per-record overhead.", "notes": "Regex baseline, excluded from the main ranking. Literal is intentionally inside the loop; mixed valid/invalid seeded records prevent a constant all-true checksum.", "asynchronous": false, "output_stderr": false, "fresh_input": false}
// Standalone file. Shared helpers/driver are inlined by common.py.

let seed = 0x12345678;
function rnd(): number {
  seed ^= seed << 13; seed ^= seed >>> 17; seed ^= seed << 5;
  return (seed >>> 0) / 4294967296;
}
function numbers(n: number): number[] {
  const a: number[] = [];
  for (let i = 0; i < n; i++) a.push(Math.floor(rnd() * 1000000));
  return a;
}
function hashArray(a: number[]): number {
  let h = a.length;
  for (let i = 0; i < a.length; i++) h = (h * 31 + a[i]) % 1000000007;
  return h;
}
// Bounded checksum work avoids making string slicing/indexing part of every
// string benchmark's asymptotic cost. The workload itself consumes its result.
function hashString(s: string): number {
  let h = s.length;
  const step = Math.max(1, Math.floor(s.length / 32));
  for (let i = 0; i < s.length; i += step) h = (h * 31 + s.charCodeAt(i)) % 1000000007;
  return h;
}

function setup(n: number): string[] {
  const values: string[] = [];
  for (let i = 0; i < n; i++) values.push((rnd() < 0.5 ? 'record_' : '!bad_') + String(i));
  return values;
}
function run(input: string[]): number {
  let h = 0;
  for (let i = 0; i < input.length; i++) h = (h * 31 + (/^[a-z]+_[0-9]+$/.test(input[i]) ? 1 : 0)) % 1000000007;
  return h;
}

// Size is the final argument: both native Perry and Node expose it reliably.
const n = Number(process.argv[process.argv.length - 1]);
if (!(n > 0)) throw new Error("Expected a positive size argument");
function benchmarkMain(): void {
  seed = 0x12345678;
  const preparedInput = setup(n);
  let checksum = 0;
  let seen = false;
  let warmMs = 0;
  let warmRuns = 0;
  while (warmMs < 200 || warmRuns < 5) {
    seed = 0x12345678;
    const input = preparedInput;
    const start = performance.now();
    const value = run(input);
    const elapsed = performance.now() - start;
    if (!(elapsed >= 0)) throw new Error("Invalid monotonic timer");
    warmMs += elapsed;
    warmRuns++;
    if (seen && value !== checksum) throw new Error("CORRECTNESS: unstable checksum during warmup");
    checksum = value;
    seen = true;
  }
  const samples: number[] = [];
  let runs = 0;
  for (let sample = 0; sample < 7; sample++) {
    let elapsed = 0;
    let count = 0;
    // Mutable workloads prepare fresh input BEFORE each timer; immutable
    // workloads reuse setup. Neither preparation nor validation is measured.
    while (elapsed < 20) {
      seed = 0x12345678;
      const input = preparedInput;
      const start = performance.now();
      const value = run(input);
      const duration = performance.now() - start;
      if (!(duration >= 0)) throw new Error("Invalid monotonic timer");
      elapsed += duration;
      count++;
      if (value !== checksum) throw new Error("CORRECTNESS: unstable checksum during sampling");
    }
    samples.push(elapsed / count);
    runs += count;
  }
  // Do not depend on Array.sort to compute the median of a sort benchmark.
  for (let i = 1; i < samples.length; i++) {
    const v = samples[i];
    let j = i - 1;
    while (j >= 0 && samples[j] > v) { samples[j + 1] = samples[j]; j--; }
    samples[j + 1] = v;
  }
  console.log(JSON.stringify({name: "regex-test-literal", category: "regex", n,
    ms_per_run: samples[3], runs, checksum}));
}
benchmarkMain();
regex-test-literal-unicode.ts — sizes [100, 1000, 10000, 100000, 1000000]

Size meanings and fresh-input policy are in the leading metadata. result_on_stderr for this file: False.

// @runtime {"name": "regex-test-literal-unicode", "category": "regex", "verification": "checksum", "sources": [{"file": "crates/perry-runtime/src/regex.rs", "function": "js_regexp_test"}], "hypothesis": "Hypothesis: literal-site compilation caching and matcher dispatch determine per-record overhead.", "notes": "Unicode regex baseline, excluded from the main ranking. Subject and pattern contain umlauts, CJK and/or emoji, using the Unicode flag. n counts records; non-ASCII /g indices are UTF-16 offsets, not UTF-8 byte offsets. Checksums consume test outcomes, actual match contents/indices, or replacement/split outputs. Literal is intentionally inside the loop; mixed valid/invalid seeded records prevent a constant all-true checksum.", "asynchronous": false, "output_stderr": false, "fresh_input": false}
// Standalone file. Shared helpers/driver are inlined by common.py.

let seed = 0x12345678;
function rnd(): number {
  seed ^= seed << 13; seed ^= seed >>> 17; seed ^= seed << 5;
  return (seed >>> 0) / 4294967296;
}
function numbers(n: number): number[] {
  const a: number[] = [];
  for (let i = 0; i < n; i++) a.push(Math.floor(rnd() * 1000000));
  return a;
}
function hashArray(a: number[]): number {
  let h = a.length;
  for (let i = 0; i < a.length; i++) h = (h * 31 + a[i]) % 1000000007;
  return h;
}
// Bounded checksum work avoids making string slicing/indexing part of every
// string benchmark's asymptotic cost. The workload itself consumes its result.
function hashString(s: string): number {
  let h = s.length;
  const step = Math.max(1, Math.floor(s.length / 32));
  for (let i = 0; i < s.length; i += step) h = (h * 31 + s.charCodeAt(i)) % 1000000007;
  return h;
}

function setup(n: number): string[] {
  const values: string[] = [];
  for (let i = 0; i < n; i++) values.push((rnd() < 0.5 ? 'ä中😀_' : '!ä中😀_') + String(i));
  return values;
}
function run(input: string[]): number {
  let h = 0;
  for (let i = 0; i < input.length; i++) h = (h * 31 + (/^[ä中😀]+_[0-9]+$/u.test(input[i]) ? 1 : 0)) % 1000000007;
  return h;
}

// Size is the final argument: both native Perry and Node expose it reliably.
const n = Number(process.argv[process.argv.length - 1]);
if (!(n > 0)) throw new Error("Expected a positive size argument");
function benchmarkMain(): void {
  seed = 0x12345678;
  const preparedInput = setup(n);
  let checksum = 0;
  let seen = false;
  let warmMs = 0;
  let warmRuns = 0;
  while (warmMs < 200 || warmRuns < 5) {
    seed = 0x12345678;
    const input = preparedInput;
    const start = performance.now();
    const value = run(input);
    const elapsed = performance.now() - start;
    if (!(elapsed >= 0)) throw new Error("Invalid monotonic timer");
    warmMs += elapsed;
    warmRuns++;
    if (seen && value !== checksum) throw new Error("CORRECTNESS: unstable checksum during warmup");
    checksum = value;
    seen = true;
  }
  const samples: number[] = [];
  let runs = 0;
  for (let sample = 0; sample < 7; sample++) {
    let elapsed = 0;
    let count = 0;
    // Mutable workloads prepare fresh input BEFORE each timer; immutable
    // workloads reuse setup. Neither preparation nor validation is measured.
    while (elapsed < 20) {
      seed = 0x12345678;
      const input = preparedInput;
      const start = performance.now();
      const value = run(input);
      const duration = performance.now() - start;
      if (!(duration >= 0)) throw new Error("Invalid monotonic timer");
      elapsed += duration;
      count++;
      if (value !== checksum) throw new Error("CORRECTNESS: unstable checksum during sampling");
    }
    samples.push(elapsed / count);
    runs += count;
  }
  // Do not depend on Array.sort to compute the median of a sort benchmark.
  for (let i = 1; i < samples.length; i++) {
    const v = samples[i];
    let j = i - 1;
    while (j >= 0 && samples[j] > v) { samples[j + 1] = samples[j]; j--; }
    samples[j + 1] = v;
  }
  console.log(JSON.stringify({name: "regex-test-literal-unicode", category: "regex", n,
    ms_per_run: samples[3], runs, checksum}));
}
benchmarkMain();

Activity

  1. added
    bugConfirmed defect or regression
    performanceRuntime, compile-time, build-size, or memory performance
    on Sep 13, 2026
  2. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Cost placement: per-call candidates identified, not yet ranked

    Each .test() call currently does all of the following. Reading source lists the work; it cannot rank it — that needs a profile, which is still the first acceptance item here:

    Proposed direction (Perex and Perry together), for per-call loops written in JavaScript: reuse validation across calls. Perex would rebind from a witness of an earlier successful validation, checked by length and header (the same constant-time check with_view already does). Perry would cache that witness on the RegExp and on the string, plus a lastIndex → byte-offset hint for non-ASCII subjects so a JS-level exec loop does not re-seek either. The contract and its GC details need maintainer sign-off before any implementation.

    Status: placement from reading source, by the Perex maintainer and checked line by line against perry main b5a82cfeae and the published perex 0.1.0 crate. It is not yet measured. Ownership and scheduling are awaiting a maintainer decision, so please do not start source changes yet.

    Related: #10164 (seek charge), #10165 (per-operation rebinding).

  3. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Status: not resolved by what has landed. #10183 (train 180, #10195) makes a RegExp's program and subject rebind in constant work across calls. That removed the cross-call quadratic in JS exec/matchAll/test loops on long subjects (see #10165), but it barely moves the short-string per-call cost this issue is about.

    Load-independent instruction counts (perf stat -e instructions:u, spread under 0.1 %): 1,000,000 calls on strings like "record_123" / "!bad_123", minus the program that only builds the strings.

    per call before #10183 after change
    hoisted re.test(v), /^[a-z]+_[0-9]+$/ 24,571 23,634 −3.8 %
    literal in loop 25,967 25,022 −3.6 %
    re.exec(v) + m[2] 41,184 39,835 −3.3 %

    Hoisted and literal stay within about 6 % of each other, so this remains a per-call cost, not construction. About 23.6k instructions per hoisted test means subject binding and program validation were a small share. The first acceptance item still applies: record how the per-call time divides between setup, binding/positioning, matching and result construction. Wall clock on perrymaster is about 1.5–2.8 s per million calls, depending on load; Node takes 15–75 ms.

    https://claude.ai/code/session_01Da12JXeG5XuVBma5yWp5C9

  4. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Same-host rerun of the full reproducers, as the acceptance list asks.

    Setup: perrymaster (Linux x86_64, 16 logical CPUs, shared with other lanes; load average 11–23 during the runs), Node v26.8.1, hostile-benchmarks/runtime/run.py --filter regex, release builds from source.

    Verdict: open (per-call work pending)

    The per-call cost is unchanged by the landed work: it is linear but 39–135× Node, against pre-Perex's 2.7–5.4×. The per-call breakdown (Perry side and Perex side) is with Ralph. Perex 0.1.5 (unpublished) makes an anchored miss return without trying later starts: 2,726 → 1,031 instructions per search on the "!bad_…" case.

    regex-test-literal

    n Node pre-Perex 9495bfc95 main 5d3bf85f9 checksums (main)
    100 0.004 ms 0.01 ms (2.7×) 0.25 ms (65.2×) = Node
    1,000 0.040 ms 0.11 ms (2.8×) 2.90 ms (72.9×) = Node
    10,000 0.411 ms 1.12 ms (2.7×) 27.94 ms (68.0×) = Node
    100,000 2.112 ms 11.25 ms (5.3×) 285.63 ms (135.2×) = Node
    1,000,000 37.791 ms 112.32 ms (3.0×) 1,905.90 ms (50.4×) = Node

    Log-log slope over 1k–100k: Node 0.863, main 0.997 (delta +0.134), pre-Perex 1.003.

    regex-test-literal-unicode

    n Node pre-Perex 9495bfc95 main 5d3bf85f9 checksums (main)
    100 0.004 ms 0.01 ms (2.7×) 0.24 ms (61.1×) = Node
    1,000 0.041 ms 0.11 ms (2.7×) 2.77 ms (66.8×) = Node
    10,000 0.401 ms 1.15 ms (2.9×) 15.57 ms (38.9×) = Node
    100,000 4.022 ms 11.63 ms (2.9×) 157.62 ms (39.2×) = Node
    1,000,000 21.164 ms 114.78 ms (5.4×) 2,013.81 ms (95.2×) = Node

    Log-log slope over 1k–100k: Node 0.993, main 0.878 (delta -0.116), pre-Perex 1.007.

    https://claude.ai/code/session_01Da12JXeG5XuVBma5yWp5C9

  5. proggeramlug commented on Sep 14, 2026

    @proggeramlug
    ContributorAuthor

    Remeasured on current main — 2.9× faster than when this was filed, still 5–6× the pre-Perex ratio

    Same-host, back-to-back, all three arms built from source and run sequentially on perrymaster (AMD Ryzen 7 7700X, 16 logical cores), Node v26.8.1, --no-auto-optimize, TZ=UTC LC_ALL=en_US.UTF-8.

    regex-test-literal

    n Node pre-Perex (× Node) main (× Node) main ÷ pre-Perex checksums
    100 0.0021 ms 0.010 ms (4.8×) 0.054 ms (26.3×) 5.4× = Node
    1,000 0.0208 ms 0.103 ms (4.9×) 0.584 ms (28.0×) 5.7× = Node
    10,000 0.2039 ms 1.094 ms (5.4×) 5.687 ms (27.9×) 5.2× = Node
    100,000 2.0128 ms 11.122 ms (5.5×) 56.639 ms (28.1×) 5.1× = Node
    1,000,000 20.1780 ms 110.424 ms (5.5×) 587.641 ms (29.1×) 5.3× = Node

    Log-log slopes: Node 0.997, pre-Perex 1.013, main 1.006 — all linear.

    regex-test-literal-unicode

    n Node pre-Perex (× Node) main (× Node) main ÷ pre-Perex checksums
    100 0.0021 ms 0.010 ms (4.8×) 0.059 ms (28.5×) 5.9× = Node
    1,000 0.0213 ms 0.110 ms (5.2×) 0.639 ms (30.1×) 5.8× = Node
    10,000 0.2098 ms 1.113 ms (5.3×) 6.558 ms (31.3×) 5.9× = Node
    100,000 2.0292 ms 11.092 ms (5.5×) 65.548 ms (32.3×) 5.9× = Node
    1,000,000 19.8748 ms 112.931 ms (5.7×) 651.449 ms (32.8×) 5.8× = Node

    Log-log slopes: Node 0.994, pre-Perex 1.010, main 1.009.

    Against the numbers in the issue body (72–79× Node), main is now 26–33× Node: 2.4–2.8× faster. Acceptance asks for the pre-Perex ratio (4.8–5.7×); main is 5.1–5.9× that. Not met.

    Hoisted-versus-literal control (the standalone program, 1M calls, one timed pass)

    probe Node pre-Perex main issue body (Perex, 0.5.1545)
    test literal-in-loop 1M 73.2 ms 94.9 ms 548.3 ms 1,603.7 ms
    test hoisted regex 1M 13.6 ms 32.5 ms 471.9 ms 1,509.6 ms

    Hoisted and literal are 16 % apart on main, so this is still per-call cost, not construction — acceptance item 3 holds. Same program, for context on what Perex bought: exec global ascii n=10000 is 50.1 ms on main vs 1,116.9 ms pre-Perex, and replace cb unicode n=10000 is 35.8 ms vs 715.8 ms. The regression is confined to fixed per-call overhead on short subjects; the scanning workloads are 20× better than the old engine.

    Acceptance item 1: where the per-call cost goes

    Load-independent instruction counts, perf stat -e instructions:u, three runs per arm, spread under 0.01 %. Each probe is 1,000,000 calls; the per-call figure subtracts that arm's own control program, which builds the same million strings and does no regex.

    per call Node pre-Perex main main ÷ Node main ÷ pre-Perex vs. 2026-09-13
    hoisted re.test(v) 475 592 8,315 17.5× 14.1× −64.8 % (was 23,634)
    literal in loop 578 2,055 9,631 16.6× 4.7× −61.5 % (was 25,022)
    re.exec(v) + m[2] 824 4,638 11,479 13.9× 2.5× −71.2 % (was 39,835)

    Attribution for the hoisted test, from perf record -e instructions:u on a --debug-symbols build of the same probe (verified to execute within 0.001 % of the stripped binary's instruction count). Percentages are of the whole probe; the instructions/call column scales them by the 9,629 instructions each loop iteration costs including string building.

    phase what is in it % ≈ instr/call
    per-call scratch setup MatchBuffers::new (7.9 %) + the memcpy it and find_near do (20.3 %) 28.2 % 2,710
    matching perex::executor::Vm::{run, initialize, trial, fixed_atom_charge, atom_scan, sought, atom_holds} 23.8 % 2,295
    binding / dispatch / entry find_near self (7.2 %), execute_output, bind_program, perex_dispatch::{execute, test_string}, plus memcmp (6.2 %) reached from dispatch and find_near 18.9 % 1,820
    GC bookkeeping at safepoints from_space_in_use_bytes, influx_driven_nursery_cap_bytes, gc_budgeted_*, ValidPointerSet::contains, layout_note_slot, minus the share the control program shows for the same symbols ~4.5 % ~440
    exception frame setup exception::try_push_with_kind + perry_sjlj_try 3.8 % 370
    property lookup js_object_get_field, object::spill::overflow_get 1.6 % 155
    result construction nothing — test returns a boolean and no capture-materialization symbol appears in the profile 0 % 0
    (string building) the probe's own control work: js_string_concat_value*, main 13.7 % 1,314

    Roughly 95 % of the program is accounted for, and the answer to the question this issue opened with is: only about a quarter of a .test() call is matching. A third is per-call scratch construction, and there is no result-construction cost at all for test.

    The largest single item is concrete. find_near builds a fresh MatchBuffers per call and passes it by value into Search::new; MatchBuffers is 360 bytes (its Slots<usize, INLINE_REGISTERS = 32> inline register array) and Search is 664. The hottest memcpy call site in the profile is find_near+0x361, disassembling to mov $0x150,%edx; call memcpy — a 336-byte struct move on every call, with a second one inside MatchBuffers::new. Nothing about that scratch depends on the subject: for a given program it is the same shape on every call.

    Verdict: still open

    • The gap is no longer dominated by any one thing. Even eliminating the whole scratch-setup bucket leaves ~5,600 instructions per call, against pre-Perex's 592 and Node's 475. Reaching the acceptance ratio needs both a reusable per-call search state and a cheaper engine entry (Vm::initialize alone is 5.8 %, ~560 instructions, for a 10-byte anchored match).
    • Next lever on the Perry side, if the maintainers want it: hold the MatchBuffers (and ideally the whole Search shape) on the bound program and rebind it per call instead of constructing and moving it, so a test loop stops paying ~2,700 instructions of setup. This needs the Perex side's agreement on the borrow/reentrancy contract — a nested regex call inside a replacer callback must not see a scratch another search is using — which is why I am posting it rather than implementing it.
    • Nothing in this comment changed behaviour: all three arms are unmodified source, and every checksum matches Node at every size.
  6. proggeramlug commented on Sep 15, 2026

    @proggeramlug
    ContributorAuthor

    Prototype: lending the thread's scratch removes 42 % of a .test() call

    Branch proggeramlug/perry:perf/10166-lent-scratch (93f49cf876), no PR — it pins perex through [patch.crates-io] at git f0b1b0b (0.1.6). 0.1.6 is now on crates.io, but the 7-day soak window in .cargo/config.toml (global-min-publish-age = "7 days") does not admit it until 2026-09-22, so landing needs a one-time publish-age override, which is Ralph's call.

    perex 0.1.6 adds impl ScratchOwner for &mut O, so a host can lend scratch to a Search instead of giving it up. That answers the 336-byte-per-call move in the attribution above directly: find_near built a MatchBuffers per call whose 32-register inline array was zeroed and then moved by value into Search. One cell per thread now serves every search, borrowed for the duration of one search.

    The cell also keeps whatever frames/undo length an earlier call needed, which turned out to be the other half of the cost — PERRY_REGEX_DIAG counted a scratch growth on nearly every search, so each call was rebuffering as well as constructing. In steady state a loop now grows nothing and constructs nothing.

    Instructions per call

    1M-call probes minus their own string-building control, perrymaster, three runs per arm, spread under 0.01 %. Non-ASCII rows use /^[ä中😀]+_[0-9]+$/u over "ä中😀_<i>".

    per call Node pre-Perex main 92eadb77ab this branch change × Node
    hoisted test, ASCII 475 592 8,315 4,836 −41.8 % 17.5× → 10.2×
    literal in loop, ASCII 578 2,055 9,631 6,151 −36.1 % 16.6× → 10.6×
    exec + m[2], ASCII 824 4,638 11,479 8,016 −30.2 % 13.9× → 9.7×
    hoisted test, non-ASCII 295 604 9,119 5,715 −37.3 % 30.9× → 19.4×
    literal in loop, non-ASCII 520 2,067 10,435 7,031 −32.6 % 20.1× → 13.5×

    Full reproducers

    Median ms per run, same host, back to back, checksums = Node at every size.

    n Node main this branch
    regex-test-literal 1,000,000 20.909 ms 573.805 ms (27.4×) 380.681 ms (18.2×)
    100,000 2.073 ms 58.606 ms (28.3×) 36.866 ms (17.8×)
    regex-test-literal-unicode 1,000,000 20.500 ms 662.350 ms (32.3×) 455.405 ms (22.2×)
    100,000 2.078 ms 66.223 ms (31.9×) 45.353 ms (21.8×)

    The standalone control moves with it: test hoisted regex 1M 471.9 → 280.6 ms, test literal-in-loop 1M 548.3 → 363.8 ms (Node 13.6 / 73.2).

    Semantics and safety

    • Reentrancy is a runtime borrow rather than the compile-time one perex gives within a single frame: a nested regex — a replacer callback that matches, or a poll that re-enters — finds the cell borrowed and takes the owned path, so two searches never share slots.
    • Work charging is unchanged per call. A search that asks for more frames or undo entries than the cell holds grows the cell for the next call and lets this call run the owned path from the budget it entered on, so the work a single call charges is what it charges today.
    • The operation's memory limit still sees the slots. The lent path takes a Charge for the registers, frames and undo it lends, exactly as the owner it replaces did; the thread keeps the memory, the operation only borrows it.
    • No GC pointer is stored in the cell — registers are subject offsets, frames and undo are the engine's opaque scratch, as in the owned buffers.
    • INLINE_REGISTERS drops 32 → 8 for the owned path, which is now only a fallback; the lent cell keeps 32 under LENT_REGISTERS, since it is allocated once per thread rather than moved per call.
    • cargo test -p perry-runtime --lib -- --test-threads=1: 3916 passed, 1 failed — native_stack::tests::stack_top_respects_custom_thread_stack_sizes, which fails on bare main on this host too. cargo fmt --check clean; clippy adds nothing beyond main's 12 pre-existing approx_constant errors.

    Still open

    This is one of the two levers perex 0.1.6 offers. The other, Search::restart_at, applies to the three Perry loops that search one subject repeatedly — perex_split.rs:206, perex_remove.rs:82, and the replace fast path's span collection at perex_replace_direct.rs:213 — and is not in this branch. Even with both, a hoisted test would sit around 10× Node against pre-Perex's 1.25×, so the acceptance bar here still needs a cheaper engine entry: Vm::initialize is ~560 instructions per call for a 10-byte anchored match.

  7. proggeramlug commented on Sep 15, 2026

    @proggeramlug
    ContributorAuthor

    Correction: the −41.8 % above is two changes, not one

    My previous comment attributed the whole drop to lending the scratch. The prototype branch pins perex 0.1.6 through [patch.crates-io], and main resolves perex 0.1.4, so those numbers also carry 0.1.5's start-anchored change (a program anchored at the subject's start tries only its first start). Splitting the two with a middle arm — main with the same perex pin and no lent scratch, same host, deterministic builds:

    per call main (perex 0.1.4) main + perex 0.1.6 + lent scratch perex 0.1.6's share lending's share
    hoisted test, ASCII 8,315 7,356 4,835 −959 (−11.5 %) −2,521 (−34.3 %)
    literal in loop, ASCII 9,631 8,672 6,151 −959 (−10.0 %) −2,521 (−29.1 %)
    exec + m[2], ASCII 11,479 10,524 8,016 −955 (−8.3 %) −2,508 (−23.8 %)
    hoisted test, non-ASCII 9,119 8,251 5,715 −868 (−9.5 %) −2,536 (−30.7 %)
    literal in loop, non-ASCII 10,435 9,567 7,031 −868 (−8.3 %) −2,536 (−26.5 %)

    So the engine bump is worth a flat ~870–960 instructions per call and lending is worth a flat ~2,510–2,540 on top of it. Lending is still the larger lever by roughly three to one, but it is −34.3 % on the hoisted probe rather than −41.8 %; the combined figure is what the earlier table reported.

    The flatness is itself informative: both savings are per-call constants, independent of the subject, which is what a fixed-overhead problem looks like.

  8. proggeramlug commented on Sep 17, 2026

    @proggeramlug
    ContributorAuthor

    The GC safepoint poll is worth 10.5 % of a .test() call — and that is the ceiling

    Recording a measurement so it does not live only in session chat. The per-call path runs one gc_runtime_safepoint_poll before each search (find_near_lent, before Search::new). A profile put it at ~18 % of samples; this prices it by removing it.

    Deliberately unsafe probe, never a PR: the poll()? before Search::new/new_near deleted. Baseline and variant built from the same commit (e6dcb6274d) with the same flags, on the same host.

    Arm 1 — the hoisted .test() loop (1,000,000 calls, string-building control subtracted, instructions:u):

    instructions per call
    baseline 4,792
    poll removed 4,290
    saving 502 (−10.5 %)

    Both arms print the same answer.

    Arm 2 — regex-replace-callback at n=700,000, an allocating workload:

    baseline poll removed
    ms per run 6,084 6,264
    cycle_starts 189 189
    steps 1,474,773 1,474,817
    share_permille 674 680
    mark_barrier_armed_us 33,354,846 33,494,684
    checksum 203210458 203210458

    No win, every counter within half a percent. That workload allocates constantly, so the allocation-site assists already carry the GC stepping and the regex poll is redundant there. The poll earns its cost only in a loop that allocates nothing.

    What this does NOT establish

    Neither arm exercises the hazard. The .test() probe reports cycle_starts=0 on both arms — there is never an open budgeted cycle during it, so nothing could have been left unstepped. In a non-allocating loop the pre-search poll may be the only safepoint on the path, and an open incremental cycle that stops being stepped keeps its mark barrier armed, taxing every write in the program. On the replace workload above, mark_barrier_armed_us is 33 s against a 73 s wall, so that failure mode is expensive when it happens.

    A witness for it needs a third shape that neither arm has: build a large live set, get an old-generation cycle open, then run a long non-allocating regex loop and assert the cycle still completes and the barrier does not stay armed across it.

    What it means for the options

    So the whole question is worth at most 502 instructions per call against the 4,792 a hoisted .test() costs today — set against Node's 475 and the pre-Perex engine's 592.

  9. proggeramlug commented on Sep 17, 2026

    @proggeramlug
    ContributorAuthor

    Open-cycle witness: the pre-search poll does no stepping — and my safety objection does not hold up

    I argued the pre-search poll was load-bearing: in a loop that allocates nothing it might be the only safepoint, so removing it could strand an open budgeted cycle with its mark barrier armed. The witness says otherwise.

    Building it took three designs, and the first two proved nothing

    Worth recording, because each failed its own subject-live assertion rather than producing a misleading green:

    1. Retain a large live set, then loop. cycle_starts=0 — no cycle ever opened. The old-reclaim trigger follows allocation churn, not a static heap. 8M retained strings still gave cycle_starts=0.
    2. Churn first, loop after. 11 starts, 11 completions, and the two arms identical — every cycle finished inside the churn, before the loop began, because the churn's own allocation safepoints stepped them.
    3. Interleave. Each round does a global string replace (known to open cycles) and then a long chunk of .test() calls that allocate nothing. A chunk that begins with a cycle open can only be stepped by the pre-search poll.

    Result

    Baseline versus the no-poll probe, same commit (e6dcb6274d) and flags, 4 rounds of 12,000,000 non-allocating .test() calls each — several seconds of allocation-free work per round, 48,000,000 calls in total:

    cycle_starts completions steps mark_barrier_armed_us
    baseline 6 6 5,286 301,691
    poll removed 6 6 5,286 304,911

    Identical starts, identical completions, identical step counts, and armed time within 1 %. The 8-round × 2M variant reads the same way (15/15, 39,828 steps on both arms).

    The step counts being exactly equal is the finding. If the poll were stepping cycles, removing it would change that number. It does not, so in these programs the poll performs no stepping at all: the #10253 due-check answers "nothing due" and returns, and every actual step comes from an allocation-site assist.

    What this does not establish

    • That a cycle ever spanned a non-allocating window. mark_barrier_armed_us is ~50 ms per cycle against chunks lasting seconds, so the cycles here complete inside the allocation phase. The hazard is therefore unproven, not disproven — I could not construct a program that holds a cycle open across a long allocation-free stretch, and without mid-run observation I cannot tell whether that is because the runtime prevents it or because my setup never achieved it.
    • Cancellation. The poll also returns EngineError::Cancelled. Nothing here tests whether a long regex loop stays interruptible without it. That is a separate witness and it matters for any option that skips polls.

    Where that leaves the options

    The safety objection I raised against removing the poll is not supported by evidence. That makes a stride (option 1b) low-risk on the stepping axis and leaves cancellation as the open question. It also reframes the decision as purely economic: the whole poll is 502 instructions of the 4,792 a hoisted .test() costs, and memoising the due-check can only recover part of that while requiring an invalidation funnel where every input bumps an epoch through one enforced path.

    Probe branch is local and was never pushed; the witness source is open-cycle-witness.ts in my reproducer set, and it carries its assertions in comments so the next person does not repeat designs 1 and 2.

  10. proggeramlug commented on Sep 17, 2026

    @proggeramlug
    ContributorAuthor

    Cancellation: the production poll cannot cancel, so it is not an argument for keeping it

    The remaining safety question was whether a long regex loop stays interruptible without the pre-search poll. It does not need to: nothing in production ever cancels through it.

    host::poll is the only poll the runtime passes into the engine:

    pub(crate) fn poll() -> Result<(), EngineError> {
        crate::gc::gc_runtime_safepoint_poll();
        Ok(())
    }

    It returns Ok(()) unconditionally. And EngineError::Cancelled has no producer anywhere in perry-runtime outside tests — every construction of it is in gc/tests/runtime_roots/perex_{split,glob,strings,execution}.rs, where a test supplies its own cancelling closure. Production carries only the enum variant and its message mapping in perex_api.rs:71.

    That plumbing is not dead weight: perex's contract lets a host cancel, and those tests prove Perry's paths release scratch and preserve consumed work when one does. But since no production path cancels, cancellation cannot be a reason to keep the per-call poll.

    Following gc_runtime_safepoint_poll through confirms the rest of the contract is equally narrow:

    gc_runtime_safepoint_poll -> gc_runtime_safepoint_report
      -> gc_budgeted_step_work_units_inner_with_progress(...)   // report discarded
    

    No user code, no throw, no interrupt delivery. Stepping a budgeted GC is the entire effect.

    The poll's complete account

    can it cancel? No — no production producer of Cancelled
    does it step cycles in practice? No — identical step counts with and without it (5,286 on both arms over 48M allocation-free calls)
    what does it cost? 502 instructions per call, 10.5 % of the 4,792 a hoisted .test() costs
    what does it buy? the option of stepping a due GC from a loop that allocates nothing

    That last row is the whole remaining case for it, and it is a real one — it is the mechanism by which a mutator that is not allocating still participates in an incremental collection. I could not build a program where it mattered (three designs, described in the previous comment), but "I could not construct it" is not "it cannot happen", and the person who owns the pacing should weigh that rather than take my failure to trigger it as proof.

    What has changed is that both of my earlier objections are gone. Removing or striding the poll is not blocked by stepping and not blocked by cancellation; it is a judgement about whether 502 instructions per call is worth the loss of a participation point that nothing currently exercises.

  11. proggeramlug commented on Sep 18, 2026

    @proggeramlug
    ContributorAuthor

    Current number, so the parked decision is not made against a stale one

    Not new analysis — the bar question on this issue still needs a decision rather than more measurement. But the figure in the title has moved by about half since it was written, and the decision reads differently at the new one.

    A hoisted .test() of /^[a-z]+_[0-9]+$/ over 1,000,000 short subjects, instructions per call with a control binary subtracted (same array build, same loop, trivial predicate in place of the regex):

    instructions/call vs Node 26.5.1
    Node 26.5.1 873.6 1.00x
    Perry, current main 4,382.2 5.02x
    Perry, with #10580 4,063.7 4.65x

    The title's "~1.5 µs per call, 44x slower than the old engine" predates perex 0.1.4 -> 0.1.9, the per-thread lent match scratch (#10372), the native replace pieces (#10412), the strided pre-search GC poll (#10494) and the one-shot search entry (#10580, open).

    Instruction counts rather than wall time on purpose: the measuring host carried other sessions' builds throughout, and round-to-round spread reached 68-244% against effects of 0.5-25%, enough to invert the sign of a known-good result. Instruction counts are load-insensitive; the ratio above is the one I would stand behind, and a microsecond figure is not currently measurable here.

    For the bar itself, my read is unchanged and still on this issue: it should be restated as a ratio to Node, because the pre-Perex engine it names was ~20x worse on scanning work, so "1.25x of the old engine" is a moving and mostly meaningless target. At 4.65x Node the gap is real but it is no longer the order-of-magnitude story the title tells.

  12. proggeramlug commented on Sep 27, 2026

    @proggeramlug
    ContributorAuthor

    Package-level measurement from the Phase 3 attribution (#11464, benchmarks/packages/PROFILE.md): origin/main 36420d2, release compiler, auto-optimize, Linux x86-64, perf record -e instructions:u --call-graph dwarf at two N (per-iteration, startup cancels), Node 26.5.1 oracle with outputs checked equal. "% of excess" = share of (Perry − Node) instructions/iter, equal-weight over the 23 packages with Perry/Node ≥ 2×.

    Regex is the second-largest bucket: 17.1% of total excess, ≥5% in 11 packages. Perex frames are on the stack for 18.4% of excess (≥2% in 14 packages). About 73% of the bucket is the matcher itself (Vm::run, trial, atom_scan, begin_class, Cursor::next_unit); ~19% is result building; ~5% is spec-protocol property gets (#10518); ~3% is compilation.

    Per-package share of excess: dotenv 80%, node-cron 74%, uuid 53%, jsonwebtoken 41%, nanoid 35%, date-fns 29%, validator 24%, commander 17%, decimal.js 14%.

    Top sites: jws/lib/verify-stream.js:41 JWS_REGEX.test(string) (60% of jsonwebtoken/decode); uuid validate REGEX.test; validator/lib/isByteLength.js encodeURI(str).split(/%..|./) (11.5% of validator/batch, the #10165 split path); decimal.js str.search(/e/i), where 8.3% of decimal.js/parse_sum is symbol::is_registered_symbol_slow (the @@search lookup, #10518 family); dayjs String.replace with a callback, which sets up an exception::try_push_with_kind per callback call.

  13. proggeramlug commented on Sep 28, 2026

    @proggeramlug
    ContributorAuthor

    Progress: #11548 landed perex 0.1.11 (the PerryTS/perex#3 matcher), under the owner-approved one-time publish-age override. Instructions: jsonwebtoken decode −61%, uuid v4 −48%, v7 −34%, v5_parse −28%, nanoid −30%, uuid/jws regex micro −65%/−84%; commander +0.4%; RSS flat. Together with #11543 this is the planned matcher + host work; leaving it open for the remaining gap to Node.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugConfirmed defect or regressionperformanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions