Skip to content

perf(regexp): constructing a RegExp recompiles the pattern every time (14 KB emoji-regex: 631 µs vs 2 µs in bun, 274×) — string-width makes opencode --help cost 31 s #10179

Description

@proggeramlug

Parent: #10106 (OpenCode startup) / #10107. Found while attributing opencode --help (31 s of CPU vs 0.4 s under bun).

The number

yargs 18 renders help through cliui → string-width@7.2.0, which for every grapheme of every line calls emojiRegex() — i.e. constructs new RegExp(<14,116-char emoji pattern>, "g") — then .test()s it. Per-operation cost, perry (linux-x64, --no-auto-optimize, 8-module probe) vs bun 1.3.14, same machine:

operation perry bun ratio
emojiRegex() construction (14,116-char source) 631 µs 2.3 µs 274×
.test("a") on a reused emoji regex 11.1 µs 0.4 µs 28×
.test("🚀") reused 12.6 µs 0.3 µs 42×
strip-ansi (string.replace(ansiRegex, ""), regex built once) 14.3 µs 0.1 µs 143×
/^\p{Default_Ignorable_Code_Point}$/u.test("a") 1.9 µs 0.2 µs 10×
Intl.Segmenter iterate a 90-char line 68 µs 12 µs 5.6×
eastAsianWidth() ×100 92 µs 9 µs 10×

Net: 9,000 stringWidth() calls on three typical help lines = 340 s under perry vs 72 ms under bun (≈38 ms per call; GC share 1 %). That is the whole --help / run --help cost of the native OpenCode binary. Construction dominates (per grapheme, ~90 graphemes per line); the reused-regex .test and .replace paths are the second tier.

Repro

// inside opencode-v1.18.30/packages/opencode (paths from the bun isolated install)
import emojiRegex from "…/node_modules/.bun/emoji-regex@10.6.0/node_modules/emoji-regex/index.js"
import stripAnsi from "…/strip-ansi@7.1.2/node_modules/strip-ansi/index.js"
function t(name, n, f) { const t0 = performance.now(); for (let i = 0; i < n; i++) f(); console.log(name, ((performance.now() - t0) * 1000 / n).toFixed(1), "us/iter") }
t("construct", 200, () => { emojiRegex() })
const re = emojiRegex(); t("test reused", 2000, () => { re.test("a") })
t("stripAnsi", 2000, () => { stripAnsi("      --log-level   log level   [string] [choices: \"DEBUG\", \"INFO\"]") })

V8/JSC cache the compiled program by (source, flags), so a repeated new RegExp(sameLiteral) costs a hash lookup; perry recompiles the 14 KB pattern every time (see the earlier regex analysis: the engine's per-byte matching is fine, the cost is eager compilation plus per-call wrapper overhead).

Expected

  • new RegExp(source, flags) for a source already compiled in this process returns a RegExp sharing the compiled program (per-object lastIndex stays per object) — construction of the emoji regex drops to the microsecond range.
  • .test/.replace on a reused regex within ~3× of bun for short inputs (the per-call wrapper overhead noted in the regex analysis).
  • opencode --help user CPU falls from 31 s to well under 1 s.

Acceptance

Activity

  1. added
    performanceRuntime, compile-time, build-size, or memory performance
    on Sep 13, 2026
  2. proggeramlug commented on Sep 13, 2026

    @proggeramlug
    ContributorAuthor

    Coordination note: three regex PRs opened today in parallel — #10174 (bind once per split/replace/match), #10176 (no work cap), #10181 (resume searches from the previous position) for #10164/#10165. They are about matching/search calls; this issue is about construction (new RegExp(sameSource) recompiling: 631 µs per emoji-regex construction) plus the per-call .test wrapper. The lane on this issue works on perf/regexp-construction-cache and should rebase onto whichever of those land first to avoid conflicts in crates/perry-runtime/src/regex*.

  3. added a commit that references this issue on Sep 14, 2026
    c76e9c6
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceRuntime, compile-time, build-size, or memory performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions