Skip to content

Compare replayed populations by content seal instead of by retained object - #950

Closed
MaxGhenis wants to merge 76 commits into
native-scale-transportfrom
native-retention-seal
Closed

MaxGhenis wants to merge 76 commits into
native-scale-transportfrom
native-retention-seal

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Implements Max's 2026-09-17 decision on the transport lane's report §10
question 2 — option (b), replace the object comparison with a content
seal
— so the base US financial run stops retaining nineteen detached
populations between the executor's observation and its own replay comparison.

Design authority: docs/us-native-retention-seal.md (this PR).
It answers the four questions the brief asked, in order.

0. The report's proposed mechanism does not hold, and that is the first finding

Report §10 question 2(b) said _population_stamp "already folds everything
same_replayed_population compares except the type assertions". That is
wrong in both directions, and both halves are proved by scripts committed
under experiments/native-retention-seal/:

  • _population_stamp is too strict. On a US_SCHEMA population whose
    frame is round-tripped through ContentStore.put_frame/load_frame,
    same_replayed_population accepts (the store zeroes values beneath a
    null mask and NONCANONICAL_NULL_BACKING permits exactly that) while the two
    stamps differ. A stamp-equality seal would turn every resume="require"
    replay red. _population_stamp's own docstring says so.
  • _frame_identity is too weak. It misses float64 NaN payload bits,
    quiet-versus-signalling NaN (NATIVE_BITS) and
    DataFrame.flags.allows_duplicate_labels (TABLE_TYPE_OR_FLAGS), because
    _cell spells every NaN null and nothing folds the flags.

So the seal is purpose-built and lives in survey_population_replay.py,
beside the comparison it replaces. The rule it follows: every comparison of
byte strings becomes a comparison of their sha256; every other predicate keeps
the small object it applies to — a dtype, its class, an index class, an axis
name, a WeightKind, a flags value — and applies the identical is/==;
every predicate about one operand is asserted when the seal is built, with the
same refusal code. The record is O(columns), not O(rows).

What is no longer proved

Nothing about the content. Four things change in character, all stated in
§4 of the design note: byte equality becomes sha256 equality; one-sided
assertions fire on arrival rather than at comparison; a one-sided defect now
takes precedence over a two-sided one when a population carries both; and on an
object-dtype axis a difference refuses under AXIS rather than
OBJECT_VALUE, because Index.equals there is an element-wise != over
arbitrary Python objects that no digest reproduces. Every defect is still
refused, so this did not need to go back to the owner.

The battery

test_us_survey_population_replay.py now drives every mutation through the
object comparison and the seal and requires the same verdict with the same
code. That driver found four gaps in the seal's first draft; a separate
adversarial pass over the finished seal found five more, including one in the
dangerous direction (an object axis sealed through pd.Series(index.array) lost
its null-sentinel identity, because pandas 3 re-infers that Series to str).
All nine are fixed and pinned.

An adversarial pass found three real divergences in the seal, and a gate

Six independent review dimensions over the finished head, every finding put to
two independent verifiers whose default answer is REFUTED: 31 findings, 25
survived
, five answered in code. Receipt with every claim, both verdicts and
each disposition:
experiments/native-retention-seal/adversarial-verification-receipt.json.

divergence, on an object-dtype or masked axis direction now
a value carrying its own __eq__ comparison refuses, seal ACCEPTED refuses on both, at seal construction
None against float("nan") comparison accepts, seal REFUSED accepted by both
a masked-integer axis above 2**53 both refuse; the fold was lossy and the code moved exact, AXIS on both

The first is the dangerous direction — a run accepting a replay the comparison
refuses — and it was reproduced before being believed. The fix is not a patch
per case: Index.equals over an object axis is an element-wise ==, so the
fold now reproduces the equivalence classes that operator actually has,
measured on this pandas pin rather than read off its source ({None, NaN} is
one class, pd.NA and pd.NaT their own, bool/int/float one numeric class by
exact value, str and bytes their own), and refuses AXIS at seal construction
for the one case a digest cannot represent. A side effect: True against 1
and -0.0 against 0.0 now agree on OBJECT_VALUE, where the seal used to
say AXIS. 124 new battery cases pin all three, and each half of the fix
was reverted to confirm its own cases go red.

This change also broke a gate that passes at the base.
test_us_spine_blindness.py::test_runtime_population_operators_are_source_spine_blind
is a static analyser over a roster survey_population_replay.py has been on
since before this branch — the branch does not touch the test — and it refuses
a statically unresolvable subscript on a name it infers to be a column
container. Every seal record is a positional tuple, so it reported 48 sites
in one module. Verified both ways before changing anything: 1 passed in 223.35s at a64f7b733, red here. The records are now unpacked into named
fields
instead of indexed, and the one dynamic getattr over
_comparables reads through a literal reader per name — which is also strictly
more closed, since an unknown comparable now refuses instead of folding None.
No record changed shape, order or repr, so no seal identity moves; the module
is in no stage roster and in none of the inventory's 124 contracts. 1 passed in 217.42s at this head.

Two tests were vacuous and both now fail against the defect they pin: an
executor test whose default-mode arm compared one run's objects against a
different run's, and the dtype census, which asserted the same predicate
twice and never called a seal function.

Four documented mechanisms were not in the code. The design note described
a dtype token of strings and booleans, a _comparables digest and a
repr() fallback, and spent its whole residual-risk paragraph on the fallback.
The code does the opposite of tokenising — it retains the objects — and the
note's own §4(1) had argued a token would be wrong. Both sections now describe
what ships, with the three residual risks that are real.

Fail-closed

Every refusal code that exists today still fires on the same defect — which is
now checked rather than asserted, in both directions, by the battery's
agreement driver and by the 121-case sweep of the measured equality classes.
FINANCIAL_NODE_POPULATION_CHANGED means: for a node whose Population the
run still retains, its in-process stamp differs from the one recorded at
issuance; for every other node, the retained seal record no longer digests to
what it did. Neither arm is vacuous.

Six refusal codes are added, and five now have tests — they had none.
ATOMIC_OBSERVER_RETENTION is still unpinned, and both documents say so,
including that its second conjunct re-derives seal_identity from the record
whose identity it compares against and so cannot fire as written.

Declared consumers. result.financial_population is observed[final_node]
is an identity check, so the base run retains financial.ATTACH_NODE, the
property graph's attach node, and the tax gate when the rebase is enabled —
detached with the executor's own snapshot function — and seals and drops the
rest. The completion host turned out to be a consumer of the base run's whole
roster (graph_survey_completion_host.py:814-816 hands those populations to
_states, which reads their frames), so the recursive base call sets a private
retention flag and that path keeps today's behaviour exactly. Found by test,
not by reading; said plainly in the design note.

An inherited defect, repaired here

survey_population_preparation._spill_roster gained a Path.read_bytes on this
base branch at b6081efcb, which added a resource_accesses entry and left
graph_implementation_inventory.json's declared resource_accesses_sha256
stale. implementation_manifest("authenticated_survey_population_v1") therefore
refused, and SurveyPopulationCreateKernel.implementation_hash calls it —
so no 19-node graph run was possible at the base branch's tip. Bisected:
clean at 5ff889814 and 5307249b3, refusing from baaf4270c onward. The base
branch's report §4 "No committed pin moves" was recomputed before b6081efcb.

The repository had no test over the whole inventory — the one caller built three
of the ten stages — so nothing was red.
test_us_implementation_inventory_contracts.py builds all ten and checks all
124 contracts; reverting the re-pin turns both arms red, verified before
committing.

What moves

experiments/native-retention-seal/implementation-identity-receipt.json:

comparison roster module digests that move
base branch tip → this PR's US-only commits none
base branch tip → this PR's head microcosm.graph/executor.py, in all ten stages

survey_population_replay.py, graph_atomic_survey_financial.py and
survey_atomic_geography.py are in no stage roster; executor.py is in every
one, and the inventory re-pin moves inventory_sha256 in every one. Node keys
and store addresses therefore move, for the reason the base branch's own report
gave: a US stage's implementation hash is over its whole module roster.

packages/microcosm-graph hunks — main-only

Commit 9bef866c5, on its own, is the main-only change. It adds
run_graph(_population_observer_detach=False): an observer that only seals
receives the live admitted population and no snapshot is allocated. Opt-in,
default unchanged, no US import, and the keyword enters no key, no receipt and
no cache record. Four tests in test_graph_executor.py, modelled on
test_absent_observer_allocates_no_snapshot. It is deliberately not called
population_retention: that keyword existed on #893's branch (6c0f24c77,
layer 08) and made the manifest's attached population views lazy, which is a
different target.

file main-only
packages/microcosm-graph/src/microcosm/graph/executor.py yes
packages/microcosm-graph/tests/test_graph_executor.py yes
everything else no

The main-only change is now two commits, not one, and the split has to
carry both: 9bef866c5, which adds the mode, and the graph-shard hunks of
6e3b3091c, which answer two of the adversarial pass's findings — the
docstring that withdraws both guarantees by name, and the fix to the
executor test whose default-mode arm was vacuous. git log a64f7b733..HEAD -- packages/microcosm-graph/ returns exactly those two, and the whole graph-shard
diff is executor.py +30/−1 and test_graph_executor.py +152. The brief asked
for the mode in its own commit and it is; the two follow-up hunks arrived after
the review that found them.

Runs

Both required runs have run. Records, launchers and measurement JSON are in
experiments/native-retention-seal/; the full write-up is that directory's
out.md. Both measured trees carry packages/*/src byte-identical to this
head
— the only packages/ difference is two test files added afterwards,
and no source file references them.

The 1/1000 cold run and its required replay — COMPLETED

baseline 5ff889814 transport after 5307249b3 this PR vs after
runner call CPU s 2,026.78 1,907.58 1,714.92 −192.65
financial node loop wall s (19) 277.36 238.33 196.48 −41.85
prefix node loop wall s (9) 55.93 45.49 39.55 −5.94
outside both node loops, wall s 1,712.62 1,692.17 1,473.68 −218.49
required replay CPU s — 1,664.45 1,566.82 −97.63
whole-process CPU s 2,028.69 3,574.86 3,284.70 −290.16

Regenerated from the three runs' own measurement JSONs by
before_after_table.py, so no cell is transcribed by hand. CPU is
process_time and load-independent; the wall rows are not. −192.65 CPU-s on
the runner call, 1.112×
, and the comparison understates it, because this head
also carries the base branch's b6081efcb, which adds work.

The replay proof. Manifest key identical
(a08d5536bcc93e0e065e5756a8453c78c526673b998274523906fd77546c015f), node
roster identical, content_addressed projection identical with
nodes_differing: [], every node a store hit, store bytes identical — 0
paths added, 0 changed — and the cold and warm projection sha256 equal. 5,869
store objects, the same count as the base branch's after-run.

The key is not the base branch's bd511d92…, and the receipt names why:
survey_population_preparation.py (the base branch's own b6081efcb),
inventory_sha256 (this PR's re-pin, forced by that), and
microcosm.graph/executor.py (this PR's observer mode).

Which half moves a key, precisely. The seal and its battery move nothing —
the receipt's base → seal_and_battery row is empty, and
survey_population_replay.py, graph_atomic_survey_financial.py and
survey_atomic_geography.py are in no stage roster. But the US half is not
key-neutral: it carries the inherited-contract re-pin, and inventory_sha256
is in every stage manifest. Exactly one commit moves each of the two things
that move — 9bef866c5 the executor, 43fb39270 the inventory — and
git log a64f7b733..HEAD -- <path> returns a single commit for both.

Peak RSS went up and this PR is not going to spin it. 17.86 GB
whole-process against the base branch's 12.42 GB; by phase, cold
13.16 / 12.42 / 13.49 GB and replay — / 11.71 / 17.83 GB. The retention
this change removes is about 0.31 GiB at 1/1000 — an order of magnitude
below the difference — so the comparison is not evidence about retention in
either direction, and no mechanism is offered for it. Report question 5 asks
whether to chase it with a repeat on a quiet machine.

The 1/15 run — refused, as predicted, before any node

STOPPED_…PREPARATION_ISSUANCE_REFUSED at 605.65 CPU-s and 8.28 GB,
with 0 nodes run. A row-count ceiling did not fire; a byte-budget ceiling
did.
A diagnostic re-run with a sys.monitoring RAISE block — no byte of the
measured tree changed — names it: ACSCoverageAuthenticationError: CANONICAL_SIZE from issue_acs_native_coverage, the 1 MiB canonical-JSON cap
on the selected-ACS-SERIALNO list at
acs_native_coverage_binding.py:563-565. Computed from measured inputs, it
refuses above about 4.28% of source — below the 6.10% the base branch
lifted
, and it was in no ceiling census. 1/24 fits, 1/20 and everything above
refuses. So this run measures the ceiling, not the seal; the seal's measurement
is the 1/1000 run above.

Refusal codes this PR adds

Six, and until now none had a test. Five now do, and each of the three seal
tests was verified to go red against its own reverted guard.
ATOMIC_OBSERVER_RETENTION is still unpinned, and both the design note and the
report say so — including that its second conjunct re-derives seal_identity
from the record whose identity it compares against, so as written it cannot
fire. Report question 7 asks whether to test the half that can, or cut the half
that cannot.

Scope

Nothing here is a build, a certification or a release artifact; every
measurement carries "release_eligible": false, no gated data was touched and
none left the machine.

Draft, and it stays draft.

🤖 Generated with Claude Code

Reconciled with main (18 September 2026)

This branch was merged level with main (through #946 and #938's split executor) rather than rebased. Main's executor split (run_graph → _execute_graph) left _population_observer_detach referenced inside _execute_graph without being passed in, a NameError on the first detached run; commit 98b698593 threads the flag through the split. Batteries on the merged head, each rc=0 with no failure markers: packages/microcosm-graph/tests and test_us_survey_population_replay.py.

MaxGhenis and others added 30 commits September 17, 2026 17:33
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_population_stamp refuses a store round trip that same_replayed_population
accepts; _frame_identity misses NaN payload bits and DataFrame.flags. The
transport report's proposed mechanism for option (b) does not hold as written.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…l is it

Answer (4) is 'nothing about the content'; the two things that change in
character (sha256 for bytes, unary assertions firing on arrival) are stated
rather than hidden, and neither narrows what a run proves about the data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every byte-string comparison becomes a sha256 comparison; every other
predicate keeps the object it applies to and the identical is/== operator;
every one-sided predicate is asserted when the seal is built, with the same
code. The directional NONCANONICAL_NULL_BACKING disjunction stays per series.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The agreement driver caught four real gaps in the seal's first draft: an axis
needs a value-equality fold as well as a byte fold, so AXIS and the series
codes each keep their own defect; and the exact index class, the columns axis
dtype and the index _comparables were folded too coarsely. Three mutations the
battery asked for are not constructible through Population's own validation
and are written past it the way test_graph_executor writes its own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
run_graph(_population_observer_detach=False) passes the live admitted
population instead of allocating a detached snapshot per reached node. Opt-in,
default unchanged, no US import, and the keyword enters no key, receipt or
cache record. Four tests: no snapshot is allocated, the run exports byte-
identical store objects and an identical manifest key, the observer really
does receive the executor's own object, and the keyword is a bool that does
nothing without an observer.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rd it

survey_population_preparation._spill_roster gained a Path.read_bytes on
native-scale-transport at b6081ef, which added a resource_accesses entry and
left graph_implementation_inventory.json's declared
resource_accesses_sha256 stale. implementation_manifest therefore REFUSED
authenticated_survey_population_v1 -- the stage whose kernels' own
implementation_hash calls it -- at every commit from b6081ef to the branch
tip a64f7b7, so no 19-node graph run was possible there. The transport
report's section 4 'No committed pin moves' was recomputed before that commit
and is stale at its own tip.

The repository had no test over the whole inventory: the one caller built
three of the ten stages. test_us_implementation_inventory_contracts.py builds
all ten and checks all 124 declared contracts; reverting the re-pin turns both
arms red, which was verified before committing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The base financial runner now names its declared consumers before the run --
financial.ATTACH_NODE, the property graph's attach node, and the tax gate when
the rebase is enabled -- detaches those three with the executor's own snapshot
function so object identity is exactly what it is today, and seals every other
observation on arrival and drops it. Both same_replayed_population call sites
become seal comparisons with the same codes.

_node_population_stamp takes either arm: a retained Population is stamped as
before, a sealed node re-derives seal_identity from its retained record, so
FINANCIAL_NODE_POPULATION_CHANGED stays non-vacuous on both. The completion
host keeps today's behaviour untouched: it runs its own observer, retains
everything, and passes no seal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A probe fleet over the comparison surface found eleven discriminations the
first battery missed, all of the same shape: the pair is byte-identical and the
comparison refuses anyway, because it is asserting something about one operand.
A seal cannot carry those as content, so it must assert them when it is built,
and the agreement driver proves it does -- a MassChangeRecord or MassRecord
subclass, a Population subclass, a pyarrow or sparse masked carrier, a uint8
mask, an ndarray-subclass backing, a numpy str_ cell, metadata key order,
object NaN sign, and non-finite metadata, which the store codec spells in hex
rather than refusing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The 1/1000 launcher runs the transport lane's committed replay harness
unedited, byte-identical, at NEED_GB 40. The 1/15 harness is that lane's
committed tenth harness with exactly three changes -- the fraction, the RSS
ceiling the retention brief names, and the label -- verified by diffing the
bodies.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The dangerous one first: an object axis sealed through pd.Series(index.array)
lost its null-sentinel identity, because pandas 3 re-infers that Series to str
and maps both None and pd.NA to nan -- so the seal ACCEPTED a pair
Index.identical refuses. The axis now folds np.asarray(index.array).

Four code divergences, each pinned: array_equivalent is byte-tolerant for
complex and bool as well as float, so the value fold collapses all three and
the byte difference keeps reaching NATIVE_BITS; Index.identical compares
dtypes with == only, so the dtype CLASS check belongs to the series code and
not to AXIS; _array_bytes_equal does not raise, so a structured dtype with an
object field must refuse under its caller's code; and the masked canonical
flag is np.all(data[mask] == 0), which an empty string fails and truthiness
does not. Also: _comparables are compared element-wise with ==, as identical
does, rather than as a tuple that short-circuits on identity.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Found by test rather than by reading: graph_survey_completion_host:814-816
uses the base run's whole per-node population roster as its own expected
populations and hands them to _states, which reads each frame's tables. The
recursive base call now sets a private retention flag, so that path retains
exactly what it retains today, and the design note says where the saving does
not apply.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A declared consumer is sealed after the snapshot detaches it -- which is the
object the comparison ran against before this change -- and every other node is
sealed live, with no copy taken at all. One seal per node instead of two, and
the property that makes the two equivalent is now a test rather than a runtime
double-seal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…agnostic

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The root journal is cumulative: a lane appends a section above the rule and
leaves everything below it. This lane's first commit replaced the whole file,
deleting 1,623 lines of four earlier lanes' history. Restored verbatim from
a64f7b7, with this lane's section above it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
108 -> 112 tests. The battery now records each comparison's verdict on both
paths when MICROCOSM_BATTERY_RECEIPT is set, and the committed receipt shows
the tally: every comparison agrees. Two codes are unreachable and the receipt
says why -- FRAME_TYPE because Population validates its frame, STRING_POLICY
because StringDtype equality already compares storage and na_value. The strata
name is only reachable after construction, because Frame normalises it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The 1/15 harness takes the same per-file F401 ignore the native-scale lane's
three harnesses take, for the same reason: it is committed byte-identical to
what ran and its imports feed the stray-module assertion. The two probe scripts
were reformatted and RE-RUN, so the committed copy is the one that produced the
output the report quotes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Counted off the module rather than off the table. The battery reaches twenty of
them plus eleven accepted pairs; the two it cannot reach are named with the
reason each is unreachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Twenty-one files import a module this lane touched. One pytest process per
file, bounded parallelism: the native-scale lane's section 9a is the write-up
of what happens when they share one process, and this avoids it by
construction rather than by explaining it afterwards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…iagnostic

Round 1 named a second catch-all. Round 2 -- patching the other catch-alls --
refused in 3 CPU-s because acs_native_coverage_binding pins the sha256 of four
ACS module files, which is itself worth recording: those modules cannot be
instrumented. Round 3 observes the raise with sys.monitoring instead, changing
no byte of the measured tree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first attempt at that measurement was wrong and the report says so: the
workspace's editable install resolved microcosm.* to this lane's own sources,
so the run measured the head it was meant to compare against. Caught by
printing module.__file__.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MaxGhenis and others added 17 commits September 17, 2026 21:32
An adversarial pass found that sections 6 and 7c concluded "the US half of
this change moves no key at all" from a revision, 1b0b915, that is only the
lane's first five commits -- journal, probes, note, seal, battery -- and
precedes both the executor commit and the inherited-contract re-pin. The
conclusion is false of the half: the re-pin moves inventory_sha256, which is in
every stage manifest, so the US half moves every node key.

What is true, and is what the sections now say, is the two-commit attribution:
9bef866 moves microcosm.graph/executor.py in all ten stage rosters, 43fb392
moves inventory_sha256 in all ten stage manifests, and git log over each path
returns exactly one commit. The seal and the retention change themselves move
nothing, because none of their three modules is in any stage roster. The
revision is relabelled seal_and_battery so it cannot be read as the US half
again, and both wrong drafts are recorded rather than quietly replaced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…it broke

An adversarial verification pass over the finished head returned 31 findings,
25 of which survived two independent verifiers. Four are answered here in code;
the rest are documentation, answered below and in the same pass.

**The seal accepted a pair the comparison refuses.** On an object-dtype axis,
`Index.equals` is an element-wise `==`, and the value fold was a digest of the
store codec's bytes. A value carrying its own `__eq__` has equal bytes to the
plain int of the same value and is not `==` to it, so `same_replayed_population`
refused AXIS and the seal accepted -- the one direction that matters. Both
classes were reproduced first: a subclass whose `__eq__` always refuses, and
one that is equal to itself but not to the plain int.

**The seal refused a pair the comparison accepts.** pandas holds `None` and
every NaN interchangeable on an object axis, and the codec spelled them apart,
so a green run would have turned red. Measured the operator's real equivalence
classes rather than reading its source: {None, nan} is one, pd.NA and pd.NaT are
their own, bool/int/float share one numeric class by exact value, str and bytes
are their own. `_object_axis_equivalence` reproduces exactly those, refuses AXIS
at seal construction for the one case a digest cannot represent, and leaves the
byte-exact discriminations to the arm `_axis` itself falls through to -- which
also makes `True` against `1` and `-0.0` against `0.0` agree on OBJECT_VALUE
where the seal used to say AXIS.

**A masked-integer axis folded through float64.** `np.asarray` on a masked
integer array returns float64 with NA as NaN, so an Int64 axis's fold was a
float fold, lossy above 2**53, and the refusal moved to PRESENT_BITS. It now
takes the exact object view, measured for Int64, UInt64, boolean and string.

**And the change broke a gate that passes at the base.**
`test_us_spine_blindness.py::test_runtime_population_operators_are_source_spine_blind`
is a static analyser over a reviewed roster this module has been on since
before this branch, and it refuses a statically unresolvable subscript on a
name it infers to be a column container. Every seal record is a positional
tuple, so it reported 48 sites. Verified both ways before touching anything:
the test passes at a64f7b7 (1 passed in 223.35s) and failed here. The
records are now unpacked into named fields instead of indexed, and the one
dynamic getattr over `_comparables` reads through a literal reader per name --
which is also strictly more closed, since an unknown comparable now refuses
rather than folding None. No record changed shape, order or repr, so no seal
identity moves; the module is in no stage roster and in none of the inventory's
124 contracts. `1 passed in 217.42s` at this head.

124 new battery cases pin all three divergences, including the 121-case sweep
of the measured equality classes, and each half of the fix was reverted to
confirm its own cases go red. The design note's section 3(a) and section 6
described a dtype token and a repr() fallback that are not in the code -- its
own section 4 argued the token would be wrong -- and both now describe what
ships, with the three real residual risks in place of the invented one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first after-run measured a head that predates this session's seal fixes,
and the brief asks for the cold run and its required replay at the implemented
head. It is also the repeat on a quiet machine that question 5 asked for: the
run holds the machine alone, and the 21-file battery waits for it to exit.

The record states, before the outcome, what the run must reproduce -- including
the first after-run's own manifest key, because the only source file that
differs between the two heads is in no stage roster. If the key moves, that
reasoning was wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…awal, two gaps

`test_a_seal_only_observer_receives_the_live_population`'s default arm compared
one run's snapshots against a *different* run's frames, and no object from one
run_graph call can be `is`-identical to one from another, so it passed whatever
the default did. Both arms now compare within their own run, and reducing
`_observer_snapshot` to the identity turns it red -- verified.

`test_the_dtype_token_is_faithful_to_the_predicate_it_replaces` asserted
`type(a) is type(b) and a == b` against `(type(a), a) == (type(b), b)`: the
same predicate written twice, calling no seal function. Renamed to
`..._dtype_head_...` and rewritten to drive every ordered pair of the census
through `_series` and through the seal, requiring the same verdict with the
same code. Replacing the retained head with `str(dtype)` turns it red, which
is the divergence the note's own section 4(1) predicts.

`run_graph`'s new docstring withdrew one guarantee and the paragraph above it
promises two: "changes to the snapshot cannot alter execution or persistence".
The persistence half is withdrawn too, and its damage outlives the run -- a
mutating observer can leave the store holding bytes that are not the content
the node key names, with the payload digest rewritten to match, so later runs
serve them as cache hits under an unchanged key. Both halves are now withdrawn
by name.

Two battery gaps: a non-finite axis name refuses AXIS on the comparison and
UNSUPPORTED_AXIS_NAME on the seal -- the one unary assertion whose code really
does move, now pinned instead of latent; and nothing distinguished the
per-column canonical-null disjunction from a frame-wide flag, so a mixed pair
now does, and making the flag pessimistic turns it red.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first attempt measured f4b4db9, and four more findings were still open --
one of them changing executor.py. Stopped by exact pid about eight minutes in,
before any result, and relaunched at 6e3b309, whose packages/*/src is
identical to the lane head's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
68 agents over six review dimensions, each finding put to two independent
verifiers whose default answer is refuted: 31 findings, 25 survived. Five were
answered in code, four in tests, sixteen in the documents, and none was left
open. The receipt carries every claim, both verdicts and what was done, and
the six refuted findings are in it too, because a review that reports only
hits cannot be calibrated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
agreement_fuzz.py walks a seeded sweep over seven axis kinds, a 21-value
object pool and nine column mutations, driving every pair through both paths.
3,852 pairs, nine refusal codes, 466 mutual acceptances, 0 disagreements. It
is committed rather than described, because a disagreement it reports is a
finding -- this is how two of the three divergences were characterised.

Section 1's answer (4) had four items and now has three. The fourth claimed
the object-axis fold is never weaker than equals and stricter on exactly two
named pairs; both halves were false, and the section now says so and says what
replaced them: the fold reproduces the equality classes the operator actually
has, so the object-axis code no longer moves at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Question 3 asked which refusal code an object-dtype axis should keep, on a
premise the adversarial pass disproved: the fold was both weaker and stricter
than the predicate it replaced, and it has been rewritten, so the code no
longer moves and there is nothing to decide.

What replaced the question is the one worth a ruling. Two of the three
divergences lived behind a battery of 116 agreeing comparisons and a report
claiming every defect is still refused. The battery agreed because it tested
the cases its author thought of; a 40-line seeded sweep that runs in seconds
would have caught two of them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Four more things a reader should not take from it: that every measured run is
at one head (the 1/1000 was re-run after the fixes; the 1/15 was not, and why
it did not need to be), that all 21 dependent files have a result at this head
(two do not, because I invalidated them), that the agreement sweep proves
absence rather than failing to find one, and that the report's final state is
its only state -- seven claims it got wrong are left visible beside what
replaced them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
After the fix pass changed both survey_population_replay.py and executor.py,
the receipt says the same thing where it matters: base -> seal_and_battery is
empty, base -> head is microcosm.graph/executor.py alone, and ten of ten
contracts are accepted at the working tree. The executor's own digest moved
again, because the docstring is part of the file the manifest folds; the seal's
rewrite moved nothing, which re-verifies this section's claim at the head that
ships rather than at the head it was written against.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The fix pass's 124 new cases mostly drive the agreement driver, so the battery
is now 248 collected items driving 564 comparisons over the same 21 verdicts,
564 agreements and 0 disagreements -- against 116 before. The per-code table,
the acceptance count and the driving/silent decomposition are all recomputed
from the receipt, and both earlier versions of the decomposition sentence are
named as wrong rather than quietly replaced.

Also records why AXIS and OBJECT_VALUE now dominate: the equality-class sweep
crosses the members of two classes per parametrisation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…under-stated

The mode's admission named one of the two guarantees the paragraph above it
promises. Both are withdrawn, and the persistence half is the worse one,
because its damage outlives the run. Section 4 now says that, and says the
split has to carry 9bef866 plus the graph-shard hunks of 6e3b309 -- which
git log over that path returns exactly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The note's section 4 claimed the object-axis fold is never weaker than equals
and stricter on exactly two named pairs, and concluded that none of its four
items narrows what a run proves. A second adversarial pass ran both paths and
found the fold weaker in one case and stricter in another it had not named, so
for the period between the first version of this note and the fix, the change
did narrow what a run proves.

Section 4 now has three items, section 4a sets out the withdrawn fourth with
both reproductions and what replaced it, and the masked-integer axis is listed
among the exact kinds because it now is one. The item is withdrawn in place
rather than deleted, because a design authority that only shows its final
state hides which of its claims were ever checked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The battery tests the cases its author thought of, and two of the three
divergences lived behind 116 agreeing comparisons. The sweep would have found
two of them, so the note asks for one beside any hand-built battery whenever a
change replaces a predicate with a fold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
MaxGhenis and others added 2 commits September 18, 2026 10:38
…7c938)

Brings main through 8c44daa up the stack. The only conflict was the
shared journal, which keeps both sides.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Main splits run_graph into run_graph and _execute_graph, and this branch's
keyword was read inside what used to be one function, so after the merge
every run_graph call raised NameError. Same repair as 7a74751 on the
main-based PR #951 branch: pass the keyword through and accept it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Sep 23, 2026
…on line

Conflicts (2 files), resolved so both retention answers survive, separated
by profile:

- pyproject.toml: union of the per-file-ignores blocks (row-ceilings and
  byte-transports from this line, the retention-seal harnesses from #950).
- graph_atomic_survey_financial.py (12 hunks):
  - _population_retention="compact" (ee70493) is unchanged: witness every
    node, retain the compact roster, keep the executor's detachment.
  - _population_retention="all" (the default) adopts #950: seal every node on
    arrival, retain only the declared consumers (every node when
    _retain_every_node_population), detach them itself and pass
    _population_observer_detach=False; replay comparisons use the seals and
    node_populations carries the seal record for each dropped node.
  - Declared consumers come from _compact_retained_roster (financial attach,
    development attach when present, property attach, tax gate): a superset
    of #950's three that adds the development attach node this line makes
    final_node.
  - _retain_every_node_population is accepted by _run_survey_financial and
    both public wrappers; the recursive completion base call sets it exactly
    when the profile is "all"; the flag check also refuses the compact
    profile.
  - docs/us-native-retention-seal.md gains an integration note saying so.

Targeted tests on the resolved tree: test_graph_executor,
test_us_survey_population_replay, test_us_implementation_inventory_contracts,
test_us_graph_pre_geography_survey_financial and
test_us_graph_survey_completion pass. test_us_graph_atomic_survey_financial
has one failure, test_final_owner_return_cannot_mutate_materialized_geography
("DID NOT RAISE"): its profiler waits for a caller that is
run_atomic_survey_financial, which became a thin wrapper in 48ce525, so the
mutation never fires. That predates this merge and is fixed in a follow-up
commit. Completion-host and compact-retention files were still running.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Sep 25, 2026
…ock the interface

Main recorded the live-population observer opt-in (#950, #951) as graph amendment 25
while #918 was open, so the two amendments #918 records as 25 (a same-kind
WeightUpdate is declarable) and 26 (the context carries the version's metadata,
mass log and column order) become 26 and 27. Every code comment, docstring, test
docstring, changelog fragment, acceptance heading and receipt moves in lockstep;
the interface lock is re-recorded because the comments live in decl.py and
kernel.py. The _project_context docstring now points at _execute_graph, where
#938 moved the boundary selection, and amendment 27 states that the executor's
boundary mass logs are live references under amendment 25's opt-in. The three
root-level review artefacts of the shared-contract lane are dropped; the
experiments/ receipts stay, with a note on the renumbering and the pre-cherry-pick
commit hashes they cite.

Verified: packages/microcosm-graph/tests 690 passed, 1 skipped (junit);
tools/graph_acceptance_burndown.py --verify ok; ruff clean on the edited files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Sep 28, 2026
…ock the interface

Main recorded the live-population observer opt-in (#950, #951) as graph amendment 25
while #918 was open, so the two amendments #918 records as 25 (a same-kind
WeightUpdate is declarable) and 26 (the context carries the version's metadata,
mass log and column order) become 26 and 27. Every code comment, docstring, test
docstring, changelog fragment, acceptance heading and receipt moves in lockstep;
the interface lock is re-recorded because the comments live in decl.py and
kernel.py. The _project_context docstring now points at _execute_graph, where
#938 moved the boundary selection, and amendment 27 states that the executor's
boundary mass logs are live references under amendment 25's opt-in. The three
root-level review artefacts of the shared-contract lane are dropped; the
experiments/ receipts stay, with a note on the renumbering and the pre-cherry-pick
commit hashes they cite.

Verified: packages/microcosm-graph/tests 690 passed, 1 skipped (junit);
tools/graph_acceptance_burndown.py --verify ok; ruff clean on the edited files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Sep 28, 2026
…ock the interface

Main recorded the live-population observer opt-in (#950, #951) as graph amendment 25
while #918 was open, so the two amendments #918 records as 25 (a same-kind
WeightUpdate is declarable) and 26 (the context carries the version's metadata,
mass log and column order) become 26 and 27. Every code comment, docstring, test
docstring, changelog fragment, acceptance heading and receipt moves in lockstep;
the interface lock is re-recorded because the comments live in decl.py and
kernel.py. The _project_context docstring now points at _execute_graph, where
#938 moved the boundary selection, and amendment 27 states that the executor's
boundary mass logs are live references under amendment 25's opt-in. The three
root-level review artefacts of the shared-contract lane are dropped; the
experiments/ receipts stay, with a note on the renumbering and the pre-cherry-pick
commit hashes they cite.

Verified: packages/microcosm-graph/tests 690 passed, 1 skipped (junit);
tools/graph_acceptance_burndown.py --verify ok; ruff clean on the edited files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Sep 28, 2026
…ock the interface

Main recorded the live-population observer opt-in (#950, #951) as graph amendment 25
while #918 was open, so the two amendments #918 records as 25 (a same-kind
WeightUpdate is declarable) and 26 (the context carries the version's metadata,
mass log and column order) become 26 and 27. Every code comment, docstring, test
docstring, changelog fragment, acceptance heading and receipt moves in lockstep;
the interface lock is re-recorded because the comments live in decl.py and
kernel.py. The _project_context docstring now points at _execute_graph, where
#938 moved the boundary selection, and amendment 27 states that the executor's
boundary mass logs are live references under amendment 25's opt-in. The three
root-level review artefacts of the shared-contract lane are dropped; the
experiments/ receipts stay, with a note on the renumbering and the pre-cherry-pick
commit hashes they cite.

Verified: packages/microcosm-graph/tests 690 passed, 1 skipped (junit);
tools/graph_acceptance_burndown.py --verify ok; ruff clean on the edited files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Closing as superseded: this head (98b6985) is an ancestor of #1011's head (3de7348, branch native-integration-20260923), so nothing here is lost. The branch is kept. Whether the graph-native US line itself continues is decision US-2 in the 2026-10-09 consolidation review.

@MaxGhenis MaxGhenis closed this Oct 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant