Skip to content

docs(codegen): record that the per-site concat cache is admitted only for counted-loop induction variables — it never fires on cc (not a defect) #9824

Description

@proggeramlug

Found while counting executions at string-concat sites for the cc-performance
campaign. Filing rather than fixing, because it belongs to whoever owns #9514
and it bears on an attribution outside my lane.

Two independent measurements agree: the cache is not cold, it is not there

Runtime counter. I added a counter at the top of
js_string_concat_site_value, before the slot lookup — so it counts calls, not
hits. One 400-character streamed reply through the offline mock-API rig, on the
compiled claude-code TUI:

[enum-diag] concat calls=4748  site=0  chain=4953      (eager arm)
[enum-diag] concat calls=3395  site=0  chain=4853      (deferred arm)
[enum-diag] concat calls=3719  site=0  chain=4881      (third run, separate binary)

js_string_concat and js_string_concat_chain are called ~8,600 times per
reply between them. js_string_concat_site_value is called zero times, in
three runs across two separately-compiled binaries.

Static check, which explains why. The symbol is not in the linked binary at
all:

$ nm -m /tmp/cc_ksforin | grep -c js_string_concat_site_value
0
$ nm -m /tmp/cc_ksforin | grep js_string_concat
... _js_string_concat
... _js_string_concat_box
... _js_string_concat_chain
... _js_string_concat_value

Every other concat entry point is present; the site-cache one has been
dead-stripped, which it can only be if codegen emitted no call to it for
this module. So this is not a cache that misses — it is a cache with no call
sites in the workload.

(This is the same method that settled #9802: nm on the compiled object said
the outlined IC miss-handler was never referenced, and that turned out to be
the whole story there too.)

Why it is worth someone's time

  1. A cache with no call sites cannot pay back its complexity — the
    CONCAT_SITE_SLOTS table, its GC lifecycle (gc/tests/concat_site.rs) and
    the codegen pass are all carried for a path this workload never takes.
  2. An attribution elsewhere may rest on it. bench_object_property beating
    node was attributed to the per-site concat mechanism. That is a different
    program and the mechanism may well fire there — but "it works" has now been
    shown to be workload-dependent in a way nobody had measured, so the
    benchmark deserves the same nm/counter check before the attribution is
    relied on again.
  3. It may be a regression rather than a design limit. perf(strings): per-site concat cache for "literal" + proven-small value — bench_object_property beats node #9514 landed on the
    strength of a measured admission gate, so either cc's concat shapes never
    matched the specialisation, or something in codegen stopped matching them.
    Those want different fixes and the counter above distinguishes them cheaply.

What I did NOT determine

Whether cc's concat sites should match. concat_site_cache.rs specialises
prefix + <number> (its doc comment gives "field_" + j for j < 20), and I
have not characterised what cc's ~8,600 concatenations per reply actually look
like — they may legitimately all be string+string, in which case the answer is
"expected, and the cache is simply for other programs" and this issue closes as
working-as-intended with a note. The counter to settle it is a split of
js_string_concat* calls by right-hand-side type.

Reproducing

PERRY_ENUM_DIAG=<path> on the branch of #9823 reports the counters above.
Binary: a full compile of cli_2.1.112.js from perf/for-in-deferred-shadow-set
(main c7361c87c + that diff), measured with
secret-tests/cc-permission-harness/stream_scale.py at chunk 100, length 400.

https://claude.ai/code/session_014UZWia6L37DpA93VLtNK9m

Activity

  1. proggeramlug commented on Sep 5, 2026

    @proggeramlug
    ContributorAuthor

    Correction: nothing is broken. I had the premise wrong.

    I filed this asking whether "the workload changed or something broke". Having
    now read the admission gate and checked the benchmark, the answer is neither,
    and the issue should be re-read with that correction in front.

    The gate is narrow by construction, and it is behaving correctly

    concat_site_cache::try_lower_concat_site_cached admits a site only when

    • the left operand is a string literal, and
    • the right operand is proven to lie in 0..=CONCAT_SITE_ADMIT_MAX (255)
      — as a constant, as a LocalGet whose loop-induction interval has
      lo >= 0 && hi <= 255, or as x % C with a small constant C.

    That is a loop-induction range-analysis requirement, not a threshold or an
    outlining interaction. cc has no literal + proven-small-induction-variable
    concat sites, so the lowering correctly declines to fire.

    The bench_object_property attribution stands — I checked it

    benchmarks/suite/bench_object_property.ts is exactly the admitted shape, and
    it is the gate's own doc-comment example:

    const FIELDS = 20;
    for (let j = 0; j < FIELDS; j++) {
      obj["field_" + j] = i * FIELDS + j;    // literal + induction var, hi = 19 <= 255
    }

    So the mechanism fires there, the measured win is real, and no re-checking of
    that attribution is needed. What does not follow is that it transfers. The
    admitted shape — a string literal concatenated with a counted-loop induction
    variable bounded under 256 — is characteristic of microbenchmarks and rare in
    application code. cc performs ~8,600 concatenations per reply and not one of
    them qualifies.

    One thing my nm evidence proves more sharply than I said

    nm -u reports static references, not executions. Because the fill arm's
    js_string_concat_site_value call sits in the emitted instruction stream
    whenever the diamond is lowered, its absence from the object proves the
    lowering was never emitted at all for cc — not that it was emitted and
    never taken. The runtime counter (site=0) is consistent with that but weaker
    on its own. That distinction is the useful part of this issue.

    What is left worth doing

    Not a fix — a fact that is currently unrecorded. Right now "does this
    optimisation apply to real programs?" is answerable only by someone repeating
    this nm check. A per-mechanism record (the shape scripts/gc_rekeyed_key_tables.json
    uses) that names each codegen fast path, states whether it is expected to fire
    on the cc-parity bundle, and justifies the answer, would turn that from an
    unknown into a documented one — and would catch the genuine regression case: a
    gate that used to fire and stops.

    I would drop the "cache does not work" framing from the title if it stays open;
    the accurate framing is "the per-site concat cache is admitted only for
    counted-loop induction variables, so it does not apply to cc — recorded, not
    broken."

    https://claude.ai/code/session_014UZWia6L37DpA93VLtNK9m

  2. changed the title [-]perf(strings): the per-site concat cache (#9514) is entirely absent from the compiled claude-code binary — no call sites emitted[/-] [+]docs(codegen): record that the per-site concat cache is admitted only for counted-loop induction variables — it never fires on cc (not a defect)[/+] on Sep 5, 2026
  3. proggeramlug commented on Sep 6, 2026

    @proggeramlug
    ContributorAuthor

    Fixed by #9826, landed on main via merge train #9866. The train merged as its own branch, so the close-keyword never fired — closing manually.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions