Repository navigation
perf: OpenCode native binary startup must beat the bun binary — attribute and remove the module-init cost (measured 13× bun CPU on the Aug 21 build) #10106
Description
Activity
- addedenhancementNew capability or improvementNew capability or improvementperformanceRuntime, compile-time, build-size, or memory performanceRuntime, compile-time, build-size, or memory performance
on Sep 12, 2026 Startup attribution, first hard number (2026-09-13, before the symbols profile)
The
--helpfamily is not init cost: it isstring-width. yargs 18 formats help through cliui →string-width@7.2.0, and a 9-module perry probe callingstringWidth()on three typical help lines 3,000 times each (9,000 calls) takesperry (linux-x64, --no-auto-optimize)bun 1.3.14 9,000 stringWidth()calls340,204 ms (≈38 ms per call) 72 ms GC is not the issue there (
share_permille=13, arena 0.4 MB) — it is pure CPU inside the width computation:string-width7 iterates every grapheme withIntl.Segmenterand, per grapheme, testsemojiRegex()(constructs the ~14 KB emoji regex inside the loop, once per character) plus/^\p{Default_Ignorable_Code_Point}$/u. A finer probe separating regex construction, regex.test, the segmenter andstrip-ansiis running; whichever it is, this single path explainsopencode --helpat 31 s /run --helpat 20 s (the time scales with help text) and is the first lever for this ticket. The 1.6 GB live arena seen underPERRY_GC_DIAGon--helpis a separate observation to attribute with the symbols build.Finer attribution of the string-width cost (perry vs bun per operation):
emojiRegex()construction 631 µs vs 2.3 µs (274×), reused.test11 µs vs 0.4 µs (28×),strip-ansireplace 14 µs vs 0.1 µs (143×),Intl.Segmenterline 68 µs vs 12 µs,eastAsianWidth10×. Construction per grapheme dominates. Filed as #10179.perf profiles on the debug-symbols build (linux-x64,
perf record -F 499 -g)--help(27 s user): regex construction and matching plus the GC they cause.- perex compiler (
Parser::escaped_point7.1 %,Parser::sequence7.0 %,class_atom2.9 %,Prepared::emit3.1 %,builtin/leaf/join4.5 %,Program::from_words1.2 %) ≈ 26 % — that isnew RegExp(<14 KB emoji pattern>)compiled once per grapheme bystring-width(perf(regexp): constructing a RegExp recompiles the pattern every time (14 KB emoji-regex: 631 µs vs 2 µs in bun, 274×) — string-width makesopencode --helpcost 31 s #10179, PR perf(regex): cache compiled programs and speed up test/replace #10193). - perex executor (
input::Cursor::next_unit) 10.6 % — matching that regex per grapheme. - GC (
ArenaObjectCursor::next_budgeted8.8 %, trace/classifier/root-marking/layout/sweep steps ≈ 20 %) ≈ 30 % — the allocation volume of those compiles (PERRY_GC_DIAG: 89 cycles, 1.6 GB live arena).
PR perf(regex): cache compiled programs and speed up test/replace #10193 removes the first two buckets (string-width 47.9 ms → 0.46 ms per call in the lane's measurement) and most of the third.
--version(2.0 s user): flat, no single hot symbol. Top entries are generic by-name property access during module initialization —get_field_by_name_object_tail2.3 %,keys_find_slot_by_bytes1.9 %,shape_descriptor_ensure_with_holes1.7 %,js_object_get_field_by_name1.6 %,get_accessor_descriptor1.2 % — pluscore::str::from_utf81.7 %, GC layout/trace bookkeeping ≈ 6 %,js_array_get_f64/js_array_length1.8 %. In other words: 7,897 module initializers (#10180) each doing dictionary-style property reads and writes on freshly built objects, not one pathological site. The levers are the module count (#10180) and cheaper object construction/property access in init code (shape-known stores instead of by-name lookups;from_utf8on every string constant materialization is worth a look).- perex compiler (
Status 2026-09-13 evening:
--helpfix PR #10193 (regex construction cache, stringWidth 41.8 ms → 0.39 ms) is rebased onto 0.5.1557 and waiting for the merge train; the module-pruning ask #10180 (7,897 → toward bun's 4,063 modules, the main lever on the 1.9 s--versioninit cost) has a codex lane on it. Current numbers on perrymaster with the pinned compiler:--version1.9 s user (bun 0.35),--help28 s (bun 0.37),models --help6.7 s.Measured 2026-09-15 on the full v1.18.30 graph (6,926 modules, not the minimal profile)
Same Mac, five warm runs each,
--version. The binary is the lineage-8 build (perry 0.5.1568-era,PERRY_LL_SIZE_OPT=1, 855 MB darwin-arm64) and the oracle is the officialopencode-darwin-arm64release binary.perry official bun binary ratio wall (warm median) 2.32 s 0.32 s 7.3× user CPU 2.10 s 0.41 s 5.1× sys CPU 0.21 s 0.02 s — instructions retired 25.26 G 3.77 G 6.7× cycles elapsed 7.27 G 1.37 G 5.3× max RSS 548 MB 182 MB 3.0× page faults 1,460 70 — The instruction count is the number that matters: this is not I/O, page-in of an 855 MB image, or dyld. It is 21.5 G extra instructions of real work before
--versionprints.Where it comes from
crates/perry-codegen/src/codegen/entry.rs(the entrymainemitter) ends its init prelude withfor (index, prefix) in non_entry_module_prefixes.iter().enumerate() { if cross_module.deferred_module_prefixes.contains(prefix) { continue; } blk.call_void(&format!("{}__init", prefix), &[]); }
so every non-entry module's
__initruns eagerly at startup, and the only exemption isdeferred_module_prefixes. That set is populated incrates/perry/src/commands/compile/run_pipeline.rsfromModuleInitKind::Deferred, which per the #753 doc comment means reachable from the entry only through dynamicimport()edges. A module reached by even one static edge anywhere in the graph is eager, even when the program never calls into it.OpenCode is exactly the adversarial shape for that rule.
packages/opencode/src/index.tsstatically imports all 25 command modules to build the yargs tree, and those command modules then defer their heavy work behindawait import(...)inside the handler (src/cli/cmd/tui.tsdynamically importseffect,../tui/layerand the plugin host). Under bun the handler's imports never run for--version; under perry any of those subgraphs that is also statically reachable from some other module is initialized before argv is parsed.Suggested next step for whoever takes this
Attribution before fixes, per the scope already in this issue: a
PERRY_DEBUG_INITbuild that emits the eager init order, turned into a per-module init-cost table, so we can say how many of the 6,926 modules bun never evaluates for--version. The graph-level lever (point 3 in the scope) now looks like the dominant one rather than a side item, and the general form of it — init-on-first-cross-module-use rather than init-everything-at-entry — is what the target in this issue requires.Tracker: #10107.
Correction to my comment above
A code read from the Proxy/CJS lane work, which I have since confirmed against
init_order.rs, shows I framed the mechanism too strongly. Retracting the part that matters and keeping the part that is measured.What I got wrong. I implied perry eagerly initializes subgraphs bun would never evaluate, and that this is the dominant lever.
classify_eager_modulesalready exempts more than dynamic-import-only modules:.filter(|i| !i.is_dynamic && !i.type_only && !i.runtime_erased && !i.is_deferred_require)
so function-local and conditional
require()edges do not chain into a module's init either. The 2,239 eager modules are therefore mostly genuine static ESM imports — and bun evaluates a static ESM graph at startup too. Its bundle withsplitting: truehoists the static graph exactly the same way; dynamic imports become separate lazily-loaded chunks, which is the analogue of perry's Deferred set. So "perry initializes modules bun skips" is not established by the eager/deferred split alone.What stands, because it is measured rather than inferred. On the full v1.18.30 graph,
--versioncosts 25.26 G instructions against the official binary's 3.77 G, with 548 MB peak RSS against 182 MB, and 2,239 of 6,949 modules initialize before argv is parsed. The gap is real and it is CPU, not I/O.Where the difference is more likely to live, and what I would attribute before changing anything: bun tree-shakes and minifies, so the module count that survives into its bundle is much smaller than 6,949 to begin with; its
__commonJSwrappers make CJS module bodies lazy until first require; and its functions are compiled lazily, where an AOT binary has already paid that cost differently. A per-module init-cost table from a symbolized build is still the right first deliverable — it just needs to be read against bun's surviving module set, not against perry's total.A colleague is taking this issue with a symbolized build (
PERRY_KEEP_SYMBOLS=1is not an object-cache key, so it relinks an existing compile rather than rebuilding it) and will post the attribution. Their sample of a--versionrun so far: ~19% GC (four minors at 40-125 ms each at ~100% survival, plus a 307 ms full incremental cycle that frees 12 MB), ~7% realpath syscalls underjs_register_path_init→canonicalize_module_path, ~6% stack-map index build, the rest module-body execution.Recipe for anyone reproducing the split:
PERRY_COLLECT_ONLY=1writesmodule-graph.jsoninto the cache dir with a per-module eager/deferred field, without running codegen.Attribution of the
--versiongap: module count is not the leverLinux, perf stat -r5,
opencode --version:official Bun Perry ( opencode-v4, 6,949 modules)user instructions 3.53 B 23.52 B wall 0.32 s 2.01 s max RSS 196 MB 605 MB Bun runs the same modules. A
Bun.buildwith the release options (splitting,bun/nodeconditions, Solid plugin) plus a per-module evaluation counter shows Bun evaluates 1,645 modules on--version. Staticimport x from "cjs"lowers to a top-level__toESM(require_x()), so semver, ajv, fastify and js-yaml run at startup under Bun too.Perry initializes 2,238 modules. Some of the difference is intended granularity: Perry compiles
ai/@ai-sdk/*fromsrc/*.ts(225 modules where Bun loads onedistfile). That is ~1.36× the modules for 6.7× the instructions. Function-local and conditional requires are already deferred; #10285 fixes the conditional ones that were still hoisted, a correctness fix with a small startup effect.Where the 23.5 B go (DWARF-unwound perf on a
PERRY_KEEP_SYMBOLS=1relink, which is not an object-cache key: 8.5 min for the full graph):by code by outermost module init by runtime path zod 40% @agentclientprotocol/sdk 21% GC 34% inclusive (incremental full cycle started during startup 14.5%, copying minors 13%) effect 33% schemapackage 13%generic [[Set]](ordinary_set_with_receiver) 24%@modelcontextprotocol/sdk 11% native method dispatch 12% core 7%, ai 6% defineProperty/descriptors/assign ~8% zstd decode of the 34 MB embedded web UI in the pre- mainconstructor 2.8%path-registry realpath~1.7%Startup is dominated by building zod and Effect schemas at module scope. Perry is ~29× slower than Bun at it: 300 zod
z.objectschemas take 23.2 B instructions, against ~0.8 B for Bun. The main pathology is filed as #10287: oneObject.defineProperty(zod's_zod) sends every later store on that object to the slow path, 5–17× the instructions and ~6× the memory. The GC share follows from the same allocation volume. Env knobs (nursery, pacing, incremental) move it at most ~6%.Per-module fixed costs are small for ESM (~3.3 K instructions per trivial module). A CJS wrapper costs ~0.7 M instructions and ~120 KB per module, which is real but ~1–2% here.
Next levers, in order: #10287 (descriptor-bearing objects keep store fast paths and shared shapes), lazy decode of zstd-embedded assets (implementation in progress), then startup GC (no full cycle before init completes; minor root-scan growth).
Parent: #10107. Related: PERRY_STARTUP_PLAN.md (PR #10066 tiny-path work), #6532/#6533/#6534 (cc
--version227 ms), #9441 (idle tail), scavenge/tenuring notes.Measured 2026-09-12 (dev Mac, load avg ~100 — wall inflated; CPU is the fair number)
--versionwallbun run src/index.ts --version(transpiles on the fly)--helpsampleof the stripped binary shows the main thread 100 % in generated code (no wait states), i.e. eager module initialization of the graph (Effect's 225 modules buildSchema/Layer/Contextobjects at import; drizzle tables; yargs command tree). No startup number exists for the full graph (4,063 modules on v1.18.30; 7,136 on the Windows full build #9133), so expect worse before optimization.Target
On the quiet Mac mini (
perry@perry-macos.local), same host, five warm runs each, paired A/B per PERRY_STARTUP_PLAN S1 discipline:opencode --version: user CPU ≤ 0.5 s, wall ≤ 1.0 s, RSS ≤ 185 MB (beat the official v1.18.30 binary fromgh release download v1.18.30 -R anomalyco/opencode -p 'opencode-darwin-arm64.zip').opencode --helpand TUI time-to-first-frame: below bun.Scope
strip=none,PERRY_MAIN_STACK_MBif needed),sample/Instruments over--version, plusPERRY_DEBUG_INITmodule-init trace; produce a per-module init-cost table (top 50) and a per-runtime-helper table (class registration, closure allocation, object-literal construction, string interning, regex compile — regex construction is eager and SipHash-keyed per pattern in perry, see the regex analysis note in secret-tests memory, GC minors during init).@babel/*via the runtime Solid plugin — must not be in the eager init order; fix(compile): support full OpenCode source builds #9133 kept dynamic imports out of eager init, verify on this graph).PERRY_LL_SIZE_OPT=1+ optnone threshold (perf(compile): reduce generated bundle bloat #8418) trade-offs.secret-tests/opencode-1.18.30-inventory/perf/.Acceptance
--versionuser CPU and RSS below the official bun binary on the Mac mini, two independent batches.--helpparity gate stay green.