Skip to content

Retire the "upper bound" label in the cause_shares replay crosswalk - #94

Merged
MaxGhenis merged 3 commits into
mainfrom
cause-shares-neutral-replay-labels
Oct 4, 2026
Merged

MaxGhenis merged 3 commits into
mainfrom
cause-shares-neutral-replay-labels

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Aligns analysis/cause_shares.py and analysis/cause_shares.json with the corrected amterr lab reading from #92. That PR withdrew the claim that the 37 replay cases that do not reproduce the issued benefit are an "upper bound on computation-side error". Of the 26 Colorado cases with a layer-2 computational finding, 13 reproduce, 10 do not and 3 were not replayed, and the solver moved nothing in 17 of the 37 (paper/snapshot/labs/amterr/ANALYSIS.md, "Layer 3: engine replay / Results"; claims_audit.json case_level).

Builds on #92 (squash c4333b3). It replaces #93, which GitHub closed when #92's branch was deleted; the commit is the same tree, rebased onto main.

What changes in colorado_replay_reconciliation

Leaf Before After
outcome keys (3 slices, 2 cross-tabs) explained_input_facts, computation_side_upper_bound reproduced, not_reproduced
reference.classification "abs(engine_on_original - RAWBEN) <= 5; a miss is a computation-side upper bound, not a proven computation error" "reproduced if abs(engine_on_original - RAWBEN) <= 5, else not_reproduced"
committed_prose_discrepancy.claim_location paper/snapshot/labs/amterr/ANALYSIS.md:L62-L67 … :L62-L67 at commit 755a7f3179e964dccdf6dc2f7266e7c4758a0141
committed_prose_discrepancy.claim_status (absent) points at the lab's correction (ANALYSIS.md section "The 10 broad-coded cases that do not reproduce"; claims_audit.json case_level.broad_coded_misses)

The labels are single-sourced in REPLAY_OUTCOMES. The generator now raises if any replay row's within5 flag disagrees with abs(engine_on_original - RAWBEN) <= 5, so the classification string is checked against the data rather than asserted. All 88 values under the renamed keys are unchanged.

No other code reads the old keys. A grep of app/, analysis/, paper/, paper-causal/, tests/, scripts/ and docs/ finds them only in these two files.

Hash-chain regeneration

cause_shares.json records the sha256 of its generator, and other artifacts pin the JSON's sha256, some of them transitively. Each artifact below was regenerated with its own generator, in dependency order. A leaf-by-leaf diff confirms that no data value changed. The only changes are input hashes and, where an artifact records it, environment.platform.

Artifact Changed leaves
analysis/cause_shares.json the four leaves above, plus provenance.generator.sha256
analysis/coding_consistency.json provenance.reference_artifact_sha256
app/public/engine_scenario_data.json provenance.cause_shares_sha256
app/public/app.js ADOPT_DATA_SHA256 pin; ASSET_V 20260820b → 20261003a (with index.html), per the cache-buster rule, so browsers do not pair a cached payload with the new pin
analysis/adoption_national.json provenance.engine_scenario_data_sha256
analysis/persistence_results.json input_hashes.coding_consistency; environment.platform
analysis/persistence_backtest_results.json input_hashes.{coding_consistency,persistence_results}; environment.platform
analysis/interventions_results.json input_hashes.{coding_consistency,persistence_results}
analysis/uhip_decomposition_results.json input_hashes.{cause_shares,coding_consistency}; environment.platform
analysis/fixed_donor_decomposition_results.json input_hashes.{cause_shares,coding_consistency,joint_fit_results}; environment.platform

Environment strings. I regenerated on Python 3.14.4 with numpy 2.5.1, pandas 2.3.3 and pyreadstat 1.3.5 (the recorded versions), so python and package strings are unchanged. environment.platform moves from macOS-26.5.1-… to macOS-26.6.2-… because the Mac's OS was updated; this is environment noise, not data. A side effect: the two local raw-regeneration tests that failed on clean main only because of that string (test_fixed_donor_decomposition, test_uhip_decomposition) now pass on this machine. They still skip in CI.

The memos the generators also write (PERSISTENCE.md, PERSISTENCE_BACKTEST.md, INTERVENTIONS.md, UHIP_DECOMPOSITION.md, FIXED_DONOR.md) came out byte-identical. build_interventions_app_data.py reproduces app/public/interventions_data.json byte for byte.

Left as historical records (deliberately not changed). Both still contain the old cause_shares.json hash c5b16617…:

Coordination with the manuscript task

The manuscript prose (paper/index.qmd, paper/FACTS.md D4) still calls the 37 a "computation-side upper bound". That text belongs to the separate manuscript task ("Correct FSBEN and reconstruction claims in the snap-qc-sim paper"), so this PR does not touch paper/. Neither file names the renamed keys, so nothing there is forced. FACTS K2 cites analysis/cause_shares.json colorado_replay_reconciliation by section name, which still resolves.

Invariants and tests (tests/test_cause_shares.py)

Invariants, each executed:

  • Classification is the documented test. For every committed replay row, within5 == (abs(engine_on_original - rawben) <= 5). The generator enforces this too.
  • Labels. Every slice and cross-tab carries exactly {reproduced, not_reproduced}, and the section contains no "upper bound" or "computation_side" text.
  • Partitions. In every slice the two outcomes sum to the total in n, weighted n and dollars, and their shares sum to 1. The official (97) and subthreshold (186) slices add up to all 283, outcome by outcome. The fractional cause cross-tab and the any-agency cross-tab each sum to the official slice. The broad residual is no larger than the not-reproduced count.
  • Differential (two implementations). cause_shares.py reads the FY2024 SAV; the lab's audit_claims.py reads the May 2026 CSV posting. They agree on 283/246/37 cases, on dollars to the cent ($99,128,122.62 / $71,388,303.72 / $27,739,818.90), on the 76/21 above-threshold split, on 283/283 solver-engine concordance, and on the broad residual (10 cases, $4,469,987.48).
  • Pointer. claim_location names a full 40-hex commit. The cited ANALYSIS.md section and figures exist in the revised file, and lines 62–67 at that commit hold the "$3.28M/yr = 3.3%" claim. The git-history check skips in shallow CI checkouts.
  • Generator guard. One-row synthetic replays: a flag that contradicts the test raises (a $100 gap marked reproduced, a $3 gap marked not, a missing engine benefit marked reproduced). A matching flag gets past the guard, including a gap of exactly $5.
  • Property (Hypothesis). _replay_metric is additive over any two-way partition of a slice, for arbitrary weights and amounts.

Against the pre-change JSON, four of the new tests fail; against the new JSON every test in the file passes.

Verification

  • Full local suite with the QC data present (Python 3.14.4, on the tree of 7e2ea21): 393 passed, 3 skipped, 0 failed in 50 minutes. Two of the skips are the matplotlib-gated figure tests in tests/test_migrations_figures.py. Every raw-regeneration test passes, including test_fixed_donor_decomposition and test_uhip_decomposition, which failed locally on main only because of the recorded macOS version.
  • tests/test_cause_shares.py at 58eb5ab: 20 passed. With the generator's within5 guard removed, the three rejection cases fail; with its <= changed to <, the exact-$5 acceptance case fails.
  • ruff check (the CI lint step) passes.
  • Pin closure: after regeneration, every live hash pin in the repo matches its file. The only remaining copies of a pre-change hash are the two historical records described above.

Review

  • Independent review, GPT-6.1 Sol via Subfleet (hard tier), on 7e2ea21: APPROVE. It checked that every changed JSON leaf is a renamed key, one of the three strings, a hash, or the platform string; that numbers and library versions are unchanged; that classification and each part of claim_status agree with ANALYSIS.md and claims_audit.json; that the historical commit holds the cited claim; that no consumer of the old keys remains; that the chain pins, ADOPT_DATA_SHA256 and the cache-busters match; and that leaving NATURE_REPORT.md and the read-only snapshot untouched is correct. Two minor findings:
    1. No test exercised the generator's new mismatch guard. Fixed in a3ab7ee: three synthetic one-row replays must raise.
    2. docs/v2-error-model.md still frames misses as needing a computation-error channel and says the replay is not in the repo. This predates the PR and is outside its scope, so it is filed as a separate task.
  • A second Subfleet reviewer (Opus 5.5) was lost in a network outage and then sat queued; I cancelled it. With the Subfleet daemon down, an in-session Opus 5.5 reviewer checked the follow-up commits, each reviewed only on its own delta:
    • a3ab7ee: APPROVE. Each rejection case reaches this guard rather than an earlier check, fails without it, and needs no data; only the test file changed, and the generator's hash still matches its pin. Its one optional suggestion: no acceptance case at the $5 boundary, so an le → lt slip would go unnoticed.
    • 58eb5ab adds that case and shares the setup. APPROVE: the sentinel is the first call after the guard, and the le → lt mutation fails only the exact-$5 case.

axiom: n/a: analysis artifact labels, regenerated provenance hashes and tests; no policy encoding.

🤖 Generated with Claude Code

The lab's 2026-10-03 revision (PR #92) withdrew the reading that the 37
non-reproduced replay cases bound computation-side error. The
colorado_replay_reconciliation section of analysis/cause_shares.json still
carried it, so:

- rename the replay outcomes explained_input_facts -> reproduced and
  computation_side_upper_bound -> not_reproduced, single-sourced in
  REPLAY_OUTCOMES;
- state only the test in reference.classification, and raise if a replay
  row's within5 flag disagrees with abs(engine_on_original - RAWBEN) <= 5;
- qualify committed_prose_discrepancy.claim_location with commit 755a7f3
  and point claim_status at the lab's correction.

Regenerate every artifact that pins the changed hashes, transitively, with
its own generator: coding_consistency, engine_scenario_data (and the
app.js pin, with an ASSET_V bump), persistence, persistence_backtest,
interventions, uhip_decomposition, fixed_donor_decomposition and
adoption_national. Data leaves are unchanged; only input hashes and the
recorded macOS version move.

Tests lock the labels and classification, check the committed replay
against the documented test, assert the partition identities, compare the
reconciliation with the lab's independently computed claims_audit.json,
and property-test _replay_metric's additivity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 2 commits October 3, 2026 22:03
Addresses the independent review: no test exercised the new guard in
colorado_replay_reconciliation, so removing it left the suite green. Three
synthetic one-row replays (a $100 gap flagged reproduced, a $3 gap flagged
not reproduced, a missing engine benefit flagged reproduced) now must raise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The rejection cases alone would not notice an le -> lt slip, and the
committed replay has no gap of exactly $5. Three consistent one-row replays
(gap of exactly $5 flagged reproduced, $6 gap and missing engine benefit
flagged not reproduced) must now get past the guard to the next call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merge audit (head 58eb5ab):

  • CI: gh pr checks exits 0 (run 37170294389, test pass). GitHub reports MERGEABLE and CLEAN; not a draft; no CHANGES_REQUESTED review.
  • Independent review: GPT-6.1 Sol via Subfleet (hard tier) on 7e2ea21 — APPROVE, with two minors. The untested guard is fixed in a3ab7ee. The stale prose in docs/v2-error-model.md predates this PR and is filed as a separate task.
  • In-session Opus 5.5 review of each follow-up delta: a3ab7ee — APPROVE, with an optional suggestion for an exact-$5 acceptance case; 58eb5ab adds that case — APPROVE. The Subfleet daemon was down at the time.
  • Local full suite with QC data (tree of 7e2ea21): 393 passed, 3 skipped, 0 failed. tests/test_cause_shares.py at 58eb5ab: 20 passed.
  • Data values unchanged; only key labels, three strings, provenance hashes, the platform string, the app.js pin and ASSET_V move.

@MaxGhenis
MaxGhenis merged commit 4b7ecab into main Oct 4, 2026
1 check passed
@MaxGhenis
MaxGhenis deleted the cause-shares-neutral-replay-labels branch October 4, 2026 02:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant