diff --git a/docs/BACKLOG.md b/docs/BACKLOG.md index 55cc19edd..b830d42d0 100644 --- a/docs/BACKLOG.md +++ b/docs/BACKLOG.md @@ -16171,7 +16171,15 @@ blind window is about a day rather than open-ended. ## 1328. no cell-to-item map exists, so an ASVS re-score cannot be handed to the seat that must flip the banner -> 🚧 **RUN FOR THE FIRST TIME 2026-09-03. The date comparison this row calls "the open work" has now produced an answer, and reading the 28 flagged entries is what remains.** The screen used to exit 3 on the only clone that exists here; it does not any more. Measured against the vault record at engine `2b8bccb4`: 62 pairs over 48 items, **51 pairs decided exactly and 11 held back as undecidable**, **35 hits over 28 distinct items**. Of those 28, **11 carry an open banner** (the missing-flip direction this row was filed for) and **17 carry a closed one** (the reverting direction its own amendment names, where a closed banner may rest on a withdrawn premise). **A hit is still a prompt to read the entry, not a finding** -- the first run's two apparent hits both dissolved on reading, and #1026 is flagged again here exactly as that worked negative case predicts. **Filed 2026-08-22.** A research-verdict item closes by TWO acts in different seats: a vault scorecard re-score, then a ledger banner flip. Nothing maps a cell to the item it governs, so the first act cannot be handed to the seat that performs the second. +> 🚧 **READ 2026-09-04. The 28 entries this row called "what remains" are read, and the answer is THREE items, not 28.** Re-run against the vault record at engine `a2eef0f37` over a full-history clone: 62 pairs over 48 items, **zero undecidable pairs and zero floored dates**, so every pair is decided against a true last-touch date -- a strictly better run than 2026-09-03's, which held 11 back. The same 35 hits over 28 distinct items. **22 of those 28 items carried NO covering grade change at all after the banner moved**, so four in five hits were a `last_verified` bump from a re-check that CONFIRMED the grade already there. **Surviving the read: `BACKLOG #1172`, `BACKLOG #1127` and `BACKLOG #1184`** -- all three in the missing-flip direction, all three carrying `closing-act: scorecard-rescore`. `BACKLOG #1172` is the cleanest instance this row has ever had: its own banner says the row stays open only because the re-score "is explicitly not run", and that re-score landed 2026-08-26, the day after. `BACKLOG #1184` is flagged in the same shape but its prose still describes the defect as live and high-value, so the direction to fix is the Tracker's call rather than an automatic flip. **ZERO re-opens.** Only two downward grade transitions exist in the entire scorecard history, and neither leaves a closed banner resting on a withdrawn premise: `BACKLOG #65`'s predates its banner's last touch, and `BACKLOG #1049`'s reverted and recovered inside a single day, ending at the grade its closure rested on. **The three flips are a ledger act this seat does not perform, so this row stays open until they land.** +> +> **A DATE MOVE IS NOT A RE-SCORE, AND THAT IS THE QUESTION THIS ROW'S OWN DISCRIMINATOR ASKS.** `last_verified` bumps on every re-verification, including one that changes nothing, so the date comparison fires on confirmations exactly as loudly as on real re-scores -- an instrument answering a question adjacent to the one asked (SDS-3.8). The tool now reads the scorecard's own git history and marks each hit GRADE MOVED, GRADE MOVED DOWN or GRADE UNCHANGED, which cuts the reading burden from 28 items to 6. It prints a direction and never a grade value; cell ids and grades stay vaulted. +> +> **TWO DISAGREEMENTS THE DATE RULE STRUCTURALLY CANNOT SEE, found by reading and not by the screen.** `BACKLOG #1148`'s covering grade moved to a pass on the SAME DAY its banner last moved, and same-day is the in-sync convention every comparison here uses, so it renders as a confirmation. `BACKLOG #1143`'s covering grade has passed since 2026-08-01, before that item's banner ever moved, while the item still says every named condition holds at HEAD. **The screen detects ORDERING, not DISAGREEMENT** -- a record that contradicted the ledger from the start is invisible to it at any depth of history. That is a different screen and it is unfiled; naming the subject rather than a number, it is the grade-versus-banner-state comparison. +> +> ***AND THE ENVIRONMENT MOVED UNDER THE 2026-09-03 RUN, WHICH IS WHY THE REFUSAL CAME BACK ONE DAY LATER.*** The shared clone re-shallowed: the graft boundary went 2026-08-04 to **2026-08-31**, visible history fell to 110 commits and the two ledger walks to 76 and 2 revisions, and the newest re-score in the record is 2026-08-29. Every pair therefore sits at or before the boundary and the tool refused, correctly, deciding nothing. **The remedy is not the owner-gated one this row assumed.** `git fetch --unshallow` is the owner's call because it writes to an object store shared by every worktree, 3.7 GB here. A THROWAWAY clone shares no object store, so that objection does not reach it, and `--root` already takes a separate history source: measured 2026-09-04 at **35 MB in 17 seconds**. Do NOT pass `--filter=blob:none` -- a blobless clone fetches one ledger revision per network round trip and did not finish in ten minutes. The refusal now prints this route instead of only naming the shared-store one. **Also measured: five items are referenced by the record but absent from every ledger revision walked (246, 274, 296, 301, 314), up from the two this row recorded, because a full history sees more of them.** +> +> **RUN FOR THE FIRST TIME 2026-09-03. The date comparison this row calls "the open work" has now produced an answer, and reading the 28 flagged entries is what remains.** The screen used to exit 3 on the only clone that exists here; it does not any more. Measured against the vault record at engine `2b8bccb4`: 62 pairs over 48 items, **51 pairs decided exactly and 11 held back as undecidable**, **35 hits over 28 distinct items**. Of those 28, **11 carry an open banner** (the missing-flip direction this row was filed for) and **17 carry a closed one** (the reverting direction its own amendment names, where a closed banner may rest on a withdrawn premise). **A hit is still a prompt to read the entry, not a finding** -- the first run's two apparent hits both dissolved on reading, and #1026 is flagged again here exactly as that worked negative case predicts. **Filed 2026-08-22.** A research-verdict item closes by TWO acts in different seats: a vault scorecard re-score, then a ledger banner flip. Nothing maps a cell to the item it governs, so the first act cannot be handed to the seat that performs the second. > > **Scored 2026-09-03 -> P1.** Value **6/10** · Difficulty **2/10** · _quick win_. The check itself shipped -- scripts/asvs/rescore_handoff_check.py implements the date comparison at :260-275, and tests/test_asvs_rescore_handoff.py passes 21 of 21 under the project venv -- so that limb is met, but the item's own stated remainder is not. Driving the tool against the vault scorecard from both this worktree and the primary checkout exited 3 with "REFUSING: the ledger history is TRUNCATED at a shallow graft boundary", because the guard at rescore_handoff_check.py:214 fired when a ledger walk began at a graft and ca8a7488, the oldest visible revision of docs/BACKLOG.md, is listed in .git/shallow; no non-shallow engine checkout exists beside this one. **CLEARED 2026-09-03 without deepening any clone, and the fix was NOT the one this paragraph proposed.** The graft-predates-the-ledger case it names was already handled by construction: `git log` over a path lists only revisions where that path CHANGED, so a graft older than the ledger's creation is never the oldest line, and the oldest line is a graft exactly when the path already existed at the boundary -- which is the real state here, so that distinction alone would have gone on refusing. **The truncation test was necessary but not sufficient, and the missing half is a date.** A floored last-touch is always later than or equal to the truth, so the floor can only suppress a hit dated AT OR BEFORE the graft boundary; a re-score strictly after it is decided identically on the floored date and on the true one. The verdict is therefore per PAIR, and the whole run is refused only when NOTHING is decidable. Value 6 because a ledger-integrity screen that refuses in the only environment it ships into leaves the step-one-done step-two-missing disagreement undetected exactly as filed; difficulty 2 for a small additive change on an existing tested seam plus the record edit. > Verdict: build diff --git a/scripts/asvs/rescore_handoff_check.py b/scripts/asvs/rescore_handoff_check.py index 5a0953bf4..f623eb3e9 100644 --- a/scripts/asvs/rescore_handoff_check.py +++ b/scripts/asvs/rescore_handoff_check.py @@ -62,7 +62,21 @@ exactly when the path already existed at the boundary. That is why the test below is on the walk's own first revision and not on the repository's shallowness. -**IT NAMES NO CELLS.** Output is item numbers and dates. Cell identifiers stay vaulted. +***AND A DATE MOVE IS NOT A RE-SCORE, WHICH IS THE QUESTION THIS FILE'S OWN TITLE ASKS.*** +``last_verified`` bumps every time a cell is looked at again -- when the grade changes, and +identically when a re-check CONFIRMS the grade already there. So the date comparison above answers +*"was this cell revisited after the banner moved"*, which is adjacent to the question asked and not +the same sentence (``CLAUDE.md`` section 11, SDS-3.8). + +**Measured 2026-09-04 against the vault record at engine ``a2eef0f37``: 35 hits over 28 items, of +which 22 items had NO covering grade change at all after the banner moved.** Four in five hits were +confirmations, and each one still cost a reader the full two-record read. The screen was not wrong -- +over-firing is its documented safe direction -- but a prompt list that is four-fifths noise is one +nobody finishes, which is the failure the item was filed about. ``grade_history`` and ``classify`` +below split the list, and the reading order is MOVED first. + +**IT NAMES NO CELLS.** Output is item numbers, dates, and a direction. Cell identifiers and grade +VALUES stay vaulted; "the grade moved down" is the actionable fact and discloses no value. Usage:: @@ -77,6 +91,7 @@ import subprocess import sys from dataclasses import dataclass +from itertools import pairwise from pathlib import Path sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "docs")) @@ -85,6 +100,10 @@ parse_items, ) +# The AUTHORITATIVE verdict vocabulary, imported rather than retyped. ``scorecard.py`` is this +# script's own sibling, so it is already importable wherever this file runs. +from scorecard import VERDICTS # noqa: E402 + #: The ONLY accepted spelling. See the module docstring: a bare ``#N`` is not reliably a backlog number. ITEM_REF = re.compile(r"BACKLOG #(\d+)") #: Counted so the ambiguous population is reported rather than silently dropped -- an unreported @@ -93,14 +112,59 @@ #: ``[[cell]]`` tables are flat, so a cell's own span runs to the next top-level table header. CELL_START = re.compile(r"^\[\[cell\]\]", re.M) LAST_VERIFIED = re.compile(r'^\s*last_verified\s*=\s*"([^"]*)"', re.M) +#: Read to JOIN a cell to its own grade history. Never printed -- cell ids stay vaulted. +CELL_ID = re.compile(r'^\s*id\s*=\s*"([^"]*)"', re.M) +VERDICT = re.compile(r'^\s*verdict\s*=\s*"([^"]*)"', re.M) +#: Strength order, used ONLY to say which DIRECTION a grade moved. The ORDER is this file's own +#: judgement and is not derivable from :data:`scorecard.Verdict`, which declares the vocabulary +#: without ranking it -- so the ranking is written here and the VOCABULARY is checked against the +#: type below rather than retyped a third time. +GRADE_RANK = { # nosec B105 - "pass" here is an ASVS grade name, not a credential + "unverified": 0, + "fail": 1, + "needs-review": 2, + "partial": 3, + "pass": 4, +} +#: Outside the ordering on purpose: ``na`` means the requirement does not apply, so a move into or out +#: of it is a scope change, and ranking it would report that as a strengthening or a weakening. +UNRANKED = frozenset({"na"}) + +# A SEVENTH VERDICT MUST NOT LAND HERE SILENTLY, AND THIS EXACT DRIFT HAS ALREADY HAPPENED ONCE. +# ``scorecard.py`` derives VERDICT_ORDER from the ``Verdict`` type precisely because a hand-written +# second list enumerated five states against a stated six (BACKLOG #1012). GRADE_RANK is a third such +# list, and an unranked verdict does not raise in ``classify`` -- it makes the comparison skip, so a +# genuine downgrade renders as GRADE MOVED instead of GRADE MOVED DOWN, which is the one direction the +# classifier exists for. Failing at import is the whole point: a checker whose vocabulary has silently +# gone stale gives a confident wrong answer, which is what every refusal in this file exists to stop. +if set(GRADE_RANK) | UNRANKED != VERDICTS: + raise RuntimeError( + "GRADE_RANK has drifted from scorecard.Verdict. Ranked " + f"{sorted(set(GRADE_RANK) | UNRANKED)}, but the record defines {sorted(VERDICTS)}. " + "Rank the new verdict or add it to UNRANKED -- leaving it out makes a downgrade into or out " + "of it render as a neutral move." + ) @dataclass(frozen=True) class Pair: - """One (item, re-score date) linkage, read from a single cell's own text.""" + """One (item, re-score date) linkage, read from a single cell's own text. + + ``cell`` is carried so the grade history below can be looked up per cell, and it is NEVER + printed. Cell identifiers stay vaulted (``CLAUDE.md`` section 12); this field exists only to join + two in-memory tables. + """ item: int last_verified: str + cell: str = "" + + +#: What ``classify`` returns, in the order a reader should care about them. +GRADE_UNCHANGED = "GRADE UNCHANGED since before the banner moved" +GRADE_MOVED = "GRADE MOVED after the banner" +GRADE_MOVED_DOWN = "GRADE MOVED DOWN after the banner -- a closed banner here may rest on a premise" +GRADE_UNKNOWN = "grade history unavailable" @dataclass(frozen=True) @@ -108,6 +172,13 @@ class Flag: item: int last_verified: str banner_touched: str + cell: str = "" + + +def _by_item(entry: tuple[Flag, str]) -> int: + """Sort key for the printed hit list. A named function rather than a lambda, so mypy --strict + checks the tuple shape at the call site instead of inferring it.""" + return entry[0].item def read_pairs(scorecard_text: str) -> tuple[list[Pair], int]: @@ -131,12 +202,145 @@ def read_pairs(scorecard_text: str) -> tuple[list[Pair], int]: verified = LAST_VERIFIED.search(block) if verified is None: continue + cell = CELL_ID.search(block) ambiguous += len(set(BARE_REF.findall(block))) for num in sorted({int(n) for n in ITEM_REF.findall(block)}): - pairs.append(Pair(item=num, last_verified=verified.group(1))) + pairs.append( + Pair( + item=num, + last_verified=verified.group(1), + cell=cell.group(1) if cell else "", + ) + ) return pairs, ambiguous +def grade_history(scorecard: Path) -> dict[str, list[tuple[str, str]]] | None: + """Per cell, the dates its GRADE actually changed -- or None when that cannot be read. + + ***THIS IS THE HALF THAT SEPARATES A RE-SCORE FROM A RE-VERIFICATION, AND WITHOUT IT THIS TOOL + ANSWERS A QUESTION ADJACENT TO THE ONE ITS OWN TITLE ASKS.*** ``last_verified`` moves every time a + cell is looked at again. It moves when the grade changes, and it moves identically when a re-check + CONFIRMS the grade already there. The date comparison cannot tell those apart, so it fires on both. + + **Measured 2026-09-04 against the vault record at engine ``a2eef0f37``: of 28 flagged items, 22 + had no covering grade change at all after the banner moved.** So roughly four in five hits were + confirmations, and every one of them cost a reader the full two-record read the item demands. The + screen was not wrong -- over-firing is its documented safe direction -- but a prompt list that is + four-fifths noise is one nobody finishes, which is the failure mode the item was filed about. + + The grade is read from the scorecard's OWN git history, oldest first, recording only the + revisions where a cell's verdict differs from the previous one. A cell absent from a revision + simply contributes nothing there. + + **RETURNS None RATHER THAN AN EMPTY DICT WHEN THE HISTORY CANNOT BE READ**, because an empty dict + would classify every hit as ``GRADE UNCHANGED`` -- the most reassuring possible answer, produced + by having failed to look. That is the same shape as the refusals this file already carries twice. + + **IT STORES GRADES BUT PRINTS NONE.** The caller turns this into MOVED / UNCHANGED / MOVED DOWN. + Cell ids and grade values stay vaulted; a direction is the actionable fact and discloses no value. + """ + repo = subprocess.run( # nosec B603 B607 - fixed argv, no shell; read-only + ["git", "-C", str(scorecard.parent), "rev-parse", "--show-toplevel"], + capture_output=True, + text=True, + # EXPLICIT, like the two calls below it. This one sets the root the others use, and it was + # the one call in this function that first went without -- which is exactly how the cp1252 + # trap recorded at length further down got in the first time. + encoding="utf-8", + ) + if repo.returncode != 0: + return None + root = Path(repo.stdout.strip()) + try: + relative = scorecard.resolve().relative_to(root.resolve()).as_posix() + except ValueError: + return None + walk = subprocess.run( # nosec B603 B607 - fixed argv, no shell; read-only git log + ["git", "-C", str(root), "log", "--format=%H %cs", "--reverse", "--", relative], + capture_output=True, + text=True, + encoding="utf-8", + ) + if walk.returncode != 0 or not walk.stdout.strip(): + return None + history: dict[str, list[tuple[str, str]]] = {} + for line in walk.stdout.splitlines(): + sha, _, date = line.partition(" ") + blob = subprocess.run( # nosec B603 B607 - fixed argv, no shell; one blob read + # EXPLICIT ENCODING, for the reason the ledger walk below records at length: locale + # decoding on this box is cp1252 and silently mangles the file. + ["git", "-C", str(root), "show", f"{sha}:{relative}"], + capture_output=True, + text=True, + encoding="utf-8", + errors="replace", + ) + if blob.returncode != 0: + continue + starts = [m.start() for m in CELL_START.finditer(blob.stdout)] + if not starts: + continue + for lo, hi in zip(starts, starts[1:] + [len(blob.stdout)], strict=True): + block = blob.stdout[lo:hi] + cell, verdict = CELL_ID.search(block), VERDICT.search(block) + if cell is None or verdict is None: + continue + seen = history.setdefault(cell.group(1), []) + if not seen or seen[-1][1] != verdict.group(1): + seen.append((date.strip(), verdict.group(1))) + return history or None + + +def classify( + flag: Flag, history: dict[str, list[tuple[str, str]]] | None, boundary: str = "" +) -> str: + """Did this cell's GRADE move after the banner did, or was it merely looked at again? + + A change strictly AFTER ``banner_touched`` is the real handoff signal. A cell whose grade has not + moved since before the banner moved was re-verified and confirmed, which is the in-sync state. + + ``na`` sits outside the ordering rather than at one end of it: it means the requirement does not + apply, so a move into or out of it is a scope change and not a strengthening or a weakening. + + ***THE FLOORED-DATE REASONING RUNS THE OPPOSITE WAY HERE, AND REUSING IT UNEXAMINED INVERTED THE + ANSWER TO THE REASSURING SIDE.*** ``evaluate`` compares ``re-score > banner_touched``, where a + floored banner date is safe: it is later than or equal to the truth, so it can only SUPPRESS a + hit. This function asks whether any grade change lands after that same date -- and pushing the + date later DELETES change points from the window, turning a real MOVED into a confident + UNCHANGED. Same date, same inequality, opposite direction, because one asks whether a single + point clears the date and the other asks what survives above it. + + So a floored banner date is answered per case rather than as a whole: + + * a change ABOVE the floor is still a true change, since ``change > floored >= true``, and it is + reported as MOVED exactly as it would be on the real date; + * NO change above the floor decides nothing, because a change could sit between the true touch + and the floor where this walk cannot see it. That is UNKNOWN, never UNCHANGED. + + Without ``boundary`` no date is treated as floored, which is right for an untruncated run. + """ + # No history at all and no history FOR THIS CELL are the same answer: nothing was read, so + # nothing is claimed. + timeline = history.get(flag.cell) if history else None + if not timeline: + return GRADE_UNKNOWN + # A touch date sitting exactly ON the boundary is a floor, not a measurement -- the same test the + # printer uses when it refuses to render that date bare. + floored = bool(boundary) and flag.banner_touched == boundary + if not any(point[0] > flag.banner_touched for point in timeline): + # UNCHANGED is a claim about everything above the reference date, so it is only available + # when that date was measured. See the docstring: the floor hides the window it would need. + return GRADE_UNKNOWN if floored else GRADE_UNCHANGED + for earlier, later in pairwise(timeline): + if later[0] <= flag.banner_touched: + continue + before_rank, after_rank = GRADE_RANK.get(earlier[1]), GRADE_RANK.get(later[1]) + if before_rank is not None and after_rank is not None and after_rank < before_rank: + return GRADE_MOVED_DOWN + return GRADE_MOVED + + @dataclass(frozen=True) class LedgerWalk: """What ONE ledger path's walk actually saw, including where it began and why it stopped there. @@ -365,11 +569,13 @@ def evaluate(pairs: list[Pair], touched: dict[int, str]) -> tuple[list[Flag], li unknown.append(pair.item) continue if pair.last_verified and pair.last_verified > when: - flags.append(Flag(pair.item, pair.last_verified, when)) + flags.append(Flag(pair.item, pair.last_verified, when, pair.cell)) return flags, sorted(set(unknown)) -def split_by_boundary(pairs: list[Pair], boundary: str) -> tuple[list[Pair], list[Pair]]: +def split_by_boundary( + pairs: list[Pair], boundary: str, touched: dict[int, str] +) -> tuple[list[Pair], list[Pair]]: """Split pairs into the ones a graft boundary cannot affect and the ones it can. A graft floors an affected item's last-touch date AT the boundary, and the floor is always later @@ -379,12 +585,33 @@ def split_by_boundary(pairs: list[Pair], boundary: str) -> tuple[list[Pair], lis of the visible history can say which -- undecidable, and the reason it is named rather than silently answered. - An empty ``boundary`` means no walk began at a graft, so everything is decidable. + ***THE SAME INEQUALITY DECIDES A SECOND CLASS, AND READING IT ONLY ONE WAY CALLED MEASURED DATES + UNDECIDABLE.*** Hidden revisions all sit at or before the boundary, so an item's true last touch + is ``max(visible, something <= boundary)``. When the VISIBLE touch is already strictly after the + boundary, that max is the visible date itself -- the walk measured it, the graft cannot raise it, + and the pair is decidable whatever its own re-score date is. Only an item whose visible touch is + still at or before the boundary is genuinely floored. + + **The measured gain is small and is stated rather than implied**: against the vault record at + engine ``a2eef0f37``, this widening decides one further pair at the 2026-09-03 graft boundary and + none at the 2026-08-31 one. It is here because calling a measured date undecidable is a wrong + answer, not because it moved a count. + + ``touched`` is REQUIRED, not defaulted. There is one production caller and nothing outside this + file imports the function, so a compatibility default protects nobody (section 0: with zero + deployments the cost of a breaking signature is zero). It would instead leave a silently weaker + path where "the caller passed no map" and "these items have no measured touch" collapse into the + same expression -- and that path errs toward UNDECIDABLE, which suppresses hits and makes the + whole-run refusal more likely. An empty ``boundary`` means no walk began at a graft, so + everything is decidable. """ if not boundary: return list(pairs), [] - decidable = [pair for pair in pairs if pair.last_verified > boundary] - undecidable = [pair for pair in pairs if pair.last_verified <= boundary] + decidable: list[Pair] = [] + undecidable: list[Pair] = [] + for pair in pairs: + clears = pair.last_verified > boundary or touched.get(pair.item, "") > boundary + (decidable if clears else undecidable).append(pair) return decidable, undecidable @@ -446,15 +673,31 @@ def main(argv: list[str] | None = None) -> int: return 3 touched = walk.touched boundary = walk.boundary - decidable, undecidable = split_by_boundary(pairs, boundary) + decidable, undecidable = split_by_boundary(pairs, boundary, touched) if boundary and not decidable: + oldest = min((pair.last_verified for pair in pairs if pair.last_verified), default="") sys.stderr.write( "REFUSING: the ledger history is TRUNCATED at a shallow graft boundary dated " f"{boundary} for {[walked.path for walked in walk.truncated]}, and EVERY pair read is " "dated at or before it, so this run can decide nothing. A floored date reads as a LATER " "touch, which SUPPRESSES real hits rather than inventing them -- the one direction this " - "check exists to rule out. Deepen the clone (git fetch --unshallow, which writes to an " - "object store shared by every worktree) and re-run.\n" + "check exists to rule out.\n" + "\n" + "THE CHEAP REMEDY IS A SEPARATE CLONE, NOT A DEEPER SHARED ONE, and the difference is " + "who has to approve it. 'git fetch --unshallow' writes to an object store shared by " + "every worktree, which is why deepening this checkout is the owner's call. A throwaway " + "clone shares no object store, so that objection does not reach it, and this tool " + "already takes the history source as an argument:\n" + "\n" + " git clone --single-branch --branch main --no-checkout \n" + f" {Path(sys.argv[0]).name} --scorecard --root \n" + "\n" + "Measured 2026-09-04: 35 MB in 17 seconds, against 3.7 GB for the shared store. " + f"The oldest re-score this run must cover is {oldest}, so a clone reaching that date is " + "enough; --unshallow is more than the question needs.\n" + "\n" + "DO NOT pass --filter=blob:none for this. A blobless clone fetches each ledger revision " + "on demand, one network round trip per revision, and did not finish in ten minutes.\n" ) return 3 # THE GUARD THIRTY LINES ABOVE, FOR THE OTHER INPUT, AND THE REASONING TRANSFERS VERBATIM. With @@ -544,8 +787,37 @@ def main(argv: list[str] | None = None) -> int: else: print("no item was re-scored after its banner was last touched") return 0 - print(f"RE-SCORED AFTER THE BANNER WAS LAST TOUCHED: {len(flags)}") - for flag in sorted(flags, key=lambda f: f.item): + history = grade_history(args.scorecard) + # ONE ORDERED LIST, COUNTED AND PRINTED FROM THE SAME STRUCTURE. Keying this on the Flag itself + # made the counts silently disagree with the lines below: two flags equal in every field collapse + # to one dict entry, so the summary would describe fewer hits than it went on to print. + ranked = sorted(((flag, classify(flag, history, boundary)) for flag in flags), key=_by_item) + if history is None: + print( + "GRADE HISTORY UNAVAILABLE -- the scorecard is not in a readable git checkout, so every " + "hit below is a bare date move and cannot be told apart from a re-verification that " + "confirmed the grade already there." + ) + else: + # THREE BUCKETS, NOT TWO, AND COUNTING THEM AS TWO MADE THIS LINE SAY THE OPPOSITE OF THE + # TRUTH. The first cut counted everything that was not UNCHANGED as "moved", so a hit whose + # grade history could not be read was reported as a grade that MOVED -- failure-to-look + # rendering as the strongest possible signal, in the one summary sentence a reader acts on. + # Each bucket is now counted by naming it, so a fourth verdict cannot silently join another. + confirmed = sum(1 for _, verdict in ranked if verdict == GRADE_UNCHANGED) + unreadable = sum(1 for _, verdict in ranked if verdict == GRADE_UNKNOWN) + moved = sum(1 for _, verdict in ranked if verdict in (GRADE_MOVED, GRADE_MOVED_DOWN)) + print( + f"of these, {moved} sit on a cell whose GRADE moved after the banner, {confirmed} on a " + f"cell that was only re-verified, and {unreadable} could not be decided. A " + "re-verification bumps last_verified without changing anything, so it is the IN-SYNC " + "state and not a handoff." + ) + # A COUNT THAT DOES NOT ADD UP TO THE LINES PRINTED BELOW IS THE BUG THIS BLOCK ALREADY HAD + # ONCE, so the arithmetic is asserted rather than trusted. + assert confirmed + unreadable + moved == len(ranked), "a verdict escaped every bucket" + print(f"RE-SCORED AFTER THE BANNER WAS LAST TOUCHED: {len(ranked)}") + for flag, verdict in ranked: # A TOUCH DATE SITTING EXACTLY ON THE BOUNDARY IS A FLOOR, NOT A MEASUREMENT, AND PRINTING # IT BARE WOULD HAND A READER A FLIP DATE THAT NEVER HAPPENED. The FLAG is still sound -- # the re-score cleared the floor and the floor is later than or equal to the truth -- but @@ -557,8 +829,18 @@ def main(argv: list[str] | None = None) -> int: if boundary and flag.banner_touched == boundary else flag.banner_touched ) - print(f" BACKLOG #{flag.item}: re-scored {flag.last_verified}, banner last touched {when}") + print( + f" BACKLOG #{flag.item}: re-scored {flag.last_verified}, banner last touched {when}" + f" [{verdict}]" + ) print() + # ONLY WHEN THERE IS SUCH A LINE TO READ. Printed unconditionally, this told a reader to start + # with a category the output above did not contain. + if any(verdict in (GRADE_MOVED, GRADE_MOVED_DOWN) for _, verdict in ranked): + print( + "READ THE 'GRADE MOVED' LINES FIRST. A 'GRADE UNCHANGED' line is a cell that was looked" + ) + print("at again and confirmed, which is what an in-sync pair looks like.") print( "A hit is a PROMPT TO READ THE ENTRY, not a finding. The item's own first run produced two" ) diff --git a/tests/test_asvs_rescore_handoff.py b/tests/test_asvs_rescore_handoff.py index e01d3b244..eed843c0a 100644 --- a/tests/test_asvs_rescore_handoff.py +++ b/tests/test_asvs_rescore_handoff.py @@ -27,10 +27,16 @@ sys.path.insert(0, str(ROOT / "scripts" / "asvs")) from rescore_handoff_check import ( # noqa: E402 + GRADE_MOVED, + GRADE_MOVED_DOWN, + GRADE_UNCHANGED, + GRADE_UNKNOWN, Flag, Pair, banner_last_touched, + classify, evaluate, + grade_history, main, read_pairs, split_by_boundary, @@ -721,7 +727,7 @@ def test_a_pair_after_the_boundary_is_decidable_and_one_at_or_before_it_is_not() is unavailable, and the answer is unknown rather than clean. """ pairs = [Pair(1, "2026-05-06"), Pair(2, "2026-05-05"), Pair(3, "2026-05-04")] - decidable, undecidable = split_by_boundary(pairs, "2026-05-05") + decidable, undecidable = split_by_boundary(pairs, "2026-05-05", {}) assert [p.item for p in decidable] == [1] assert [p.item for p in undecidable] == [2, 3] @@ -730,7 +736,7 @@ def test_with_no_boundary_every_pair_is_decidable() -> None: """The unbounded branch. An empty boundary means no walk began at a graft, so nothing is floored and holding anything back would be a refusal with no cause.""" pairs = [Pair(1, "2026-05-06"), Pair(2, "2026-01-01")] - decidable, undecidable = split_by_boundary(pairs, "") + decidable, undecidable = split_by_boundary(pairs, "", {}) assert decidable == pairs assert undecidable == [] @@ -754,3 +760,221 @@ def test_the_run_reports_what_it_actually_scanned(ledger: Path, tmp_path: Path, assert f"walked {LIVE}: 3 of 3 revisions" in out assert f"walked {ARCHIVE}: 2 of 2 revisions" in out assert "truncation branch: NONE" in out + + +# ------------------------------------------- a measured touch is decidable whatever the graft says + + +def test_an_item_whose_touch_was_MEASURED_is_decidable_below_the_boundary() -> None: + """The half of the inequality the first fix read only one way. + + Hidden revisions all sit at or before the boundary, so a true last touch is + ``max(visible, something <= boundary)``. When the visible touch already clears the boundary, that + max IS the visible date -- the graft cannot raise it. Item 2's own re-score date sits below the + boundary, and it is still decided exactly, because its BANNER date was measured rather than + floored. + """ + pairs = [Pair(1, "2026-05-04"), Pair(2, "2026-05-04")] + touched = {1: "2026-05-05", 2: "2026-05-09"} + decidable, undecidable = split_by_boundary(pairs, "2026-05-05", touched) + assert [p.item for p in decidable] == [2] + assert [p.item for p in undecidable] == [1] + + +def test_the_per_item_rule_never_narrows_the_old_one() -> None: + """A widening must not turn a previously decided pair undecidable. Pinned because the two rules + are OR-ed, and an OR written as an AND would silently shrink the answer instead of growing it.""" + pairs = [Pair(1, "2026-05-06"), Pair(2, "2026-05-04")] + old, _ = split_by_boundary(pairs, "2026-05-05", {}) + new, _ = split_by_boundary(pairs, "2026-05-05", {2: "2026-05-01"}) + assert {p.item for p in old} <= {p.item for p in new} + assert [p.item for p in new] == [1] + + +def test_a_FLOORED_banner_date_cannot_produce_a_clean_GRADE_UNCHANGED(tmp_path: Path) -> None: + """THE INEQUALITY RUNS THE OPPOSITE WAY HERE, AND REUSING IT UNEXAMINED INVERTED THE ANSWER. + + ``evaluate`` is safe against a floored banner date because a later date can only suppress a hit. + ``classify`` asks what survives ABOVE that date, so pushing it later DELETES change points and + turns a real move into a confident confirmation. Same history, same cell: on the true date the + grade moved; on the floored one the walk can see nothing above it and must say so. + """ + repo = scorecard_repo(tmp_path, [("2026-01-01", "partial"), ("2026-05-01", "pass")]) + history = grade_history(repo / "card.toml") + true_date = Flag(10, "2026-05-01", "2026-02-01", "C1") + assert classify(true_date, history) == GRADE_MOVED + floored = Flag(10, "2026-05-01", "2026-08-31", "C1") + assert classify(floored, history, "2026-08-31") == GRADE_UNKNOWN + # A change ABOVE the floor is still a true change, so it is answered rather than withheld. + low_floor = Flag(10, "2026-05-01", "2026-02-01", "C1") + assert classify(low_floor, history, "2026-02-01") == GRADE_MOVED + + +def test_an_UNKNOWN_verdict_is_not_counted_as_a_grade_that_MOVED( + ledger: Path, tmp_path: Path, capsys +) -> None: + """FAILURE-TO-LOOK MUST NOT RENDER AS THE STRONGEST SIGNAL IN THE ONE LINE A READER ACTS ON. + + The first cut counted every verdict that was not UNCHANGED as "moved", so a hit whose grade + history could not be read was summarised as a grade that moved. Here the scorecard's committed + history carries a different cell id than the working tree, which is what an uncommitted re-score + looks like, so the join misses and the verdict is UNKNOWN. + """ + repo = scorecard_repo(tmp_path, [("2026-01-01", "partial")]) + card = repo / "card.toml" + card.write_text( + '[[cell]]\nid = "UNCOMMITTED"\nverdict = "partial"\n' + 'last_verified = "2026-05-01"\nresidual = "BACKLOG #10"\n', + encoding="utf-8", + ) + assert main(["--scorecard", str(card), "--root", str(ledger)]) == 0 + out = capsys.readouterr().out + assert "0 sit on a cell whose GRADE moved after the banner" in out + assert "1 could not be decided" in out + assert GRADE_UNKNOWN in out + # The steer to read the moved lines first must not appear when there are none. + assert "READ THE 'GRADE MOVED' LINES FIRST" not in out + + +def test_the_ranked_verdicts_cannot_drift_from_the_lines_printed( + ledger: Path, tmp_path: Path, capsys +) -> None: + """The summary counts and the detail lines come from one structure, so they cannot disagree. + + Keying the verdicts on the Flag dataclass collapsed two flags equal in every field into one + entry, which would have described fewer hits than it went on to print. + """ + repo = scorecard_repo(tmp_path, [("2026-01-01", "partial"), ("2026-05-01", "pass")]) + assert main(["--scorecard", str(repo / "card.toml"), "--root", str(ledger)]) == 0 + out = capsys.readouterr().out + printed = [line for line in out.splitlines() if line.strip().startswith("BACKLOG #")] + assert "RE-SCORED AFTER THE BANNER WAS LAST TOUCHED: 1" in out + assert len(printed) == 1 + assert "1 sit on a cell whose GRADE moved after the banner" in out + + +# --------------------------------------------- a date move is not a re-score: the grade classifier + + +def scorecard_repo(tmp_path: Path, revisions: list[tuple[str, str]]) -> Path: + """A scorecard repo whose single cell takes each ``(date, verdict)`` in turn.""" + repo = tmp_path / "vault" + repo.mkdir() + git("init", "-b", "main", str(repo), cwd=tmp_path) + git("config", "user.email", "t@example.com", cwd=repo) + git("config", "user.name", "t", cwd=repo) + card = repo / "card.toml" + for when, verdict in revisions: + card.write_text( + f'[[cell]]\nid = "C1"\nverdict = "{verdict}"\n' + f'last_verified = "{when}"\nresidual = "BACKLOG #10"\n', + encoding="utf-8", + ) + commit_at(repo, f"{verdict} at {when}", when) + return repo + + +def test_a_cell_only_RE_VERIFIED_after_the_banner_is_not_a_re_score(tmp_path: Path) -> None: + """THE DEFECT THIS CLASSIFIER EXISTS FOR. ``last_verified`` bumps on a confirming re-check just as + it does on a real re-score, so the date comparison alone fires on both. + + Measured 2026-09-04 against the vault record: 22 of 28 flagged items were this case. + """ + repo = scorecard_repo(tmp_path, [("2026-01-01", "partial"), ("2026-05-01", "partial")]) + history = grade_history(repo / "card.toml") + assert classify(Flag(10, "2026-05-01", "2026-03-01", "C1"), history) == GRADE_UNCHANGED + + +def test_a_cell_whose_grade_moved_after_the_banner_is_the_real_signal(tmp_path: Path) -> None: + repo = scorecard_repo(tmp_path, [("2026-01-01", "partial"), ("2026-05-01", "pass")]) + history = grade_history(repo / "card.toml") + assert classify(Flag(10, "2026-05-01", "2026-03-01", "C1"), history) == GRADE_MOVED + + +def test_a_grade_that_moved_DOWN_after_a_banner_is_called_out_separately(tmp_path: Path) -> None: + """The reverting direction, which is the one #1328's amendment warns about: a CLOSED banner + resting on a pass that a later re-score withdrew is a compensating control on a false premise + (SDS-3.7). It must not render the same as a strengthening.""" + repo = scorecard_repo(tmp_path, [("2026-01-01", "pass"), ("2026-05-01", "partial")]) + history = grade_history(repo / "card.toml") + assert classify(Flag(10, "2026-05-01", "2026-03-01", "C1"), history) == GRADE_MOVED_DOWN + + +def test_a_move_into_na_is_a_scope_change_not_a_weakening(tmp_path: Path) -> None: + """``na`` means the requirement does not apply. Ranking it at either end of the scale would + report every scope change as a strengthening or a weakening, which is a grade claim the record + never made.""" + repo = scorecard_repo(tmp_path, [("2026-01-01", "pass"), ("2026-05-01", "na")]) + history = grade_history(repo / "card.toml") + assert classify(Flag(10, "2026-05-01", "2026-03-01", "C1"), history) == GRADE_MOVED + + +def test_an_unreadable_grade_history_is_UNKNOWN_and_never_silently_UNCHANGED( + tmp_path: Path, +) -> None: + """THE WHOLE POINT OF RETURNING None. An empty history would classify every hit as UNCHANGED -- + the most reassuring possible answer, produced by having failed to look. That is the same shape as + the two refusals this tool already carries.""" + loose = tmp_path / "loose" + loose.mkdir() + card = loose / "card.toml" + card.write_text('[[cell]]\nid = "C1"\nlast_verified = "2026-05-01"\n', encoding="utf-8") + assert grade_history(card) is None + assert classify(Flag(10, "2026-05-01", "2026-03-01", "C1"), None) == GRADE_UNKNOWN + + +def test_a_cell_absent_from_the_grade_history_is_UNKNOWN_rather_than_confirmed( + tmp_path: Path, +) -> None: + """A cell the history never saw is not a cell whose grade held steady.""" + repo = scorecard_repo(tmp_path, [("2026-01-01", "partial")]) + history = grade_history(repo / "card.toml") + assert classify(Flag(10, "2026-05-01", "2026-03-01", "MISSING"), history) == GRADE_UNKNOWN + + +def test_read_pairs_carries_the_cell_so_the_grade_can_be_joined() -> None: + """The join key. It is READ but never PRINTED -- cell ids stay vaulted.""" + pairs, _ = read_pairs( + '[[cell]]\nid = "C1"\nlast_verified = "2026-01-01"\nresidual = "BACKLOG #10"\n' + ) + assert [(p.item, p.cell) for p in pairs] == [(10, "C1")] + + +def test_the_run_splits_confirmations_from_real_grade_moves( + ledger: Path, tmp_path: Path, capsys +) -> None: + """End to end: the hit list states how much of itself is noise, so a reader knows where to start. + + #10's banner is flipped 2026-02-01 and the cell is re-verified 2026-05-01 WITHOUT moving, which + is the in-sync state and the dominant real-world case -- 22 of 28 items when this ran against the + vault record. + """ + repo = scorecard_repo(tmp_path, [("2026-01-01", "partial"), ("2026-05-01", "partial")]) + assert main(["--scorecard", str(repo / "card.toml"), "--root", str(ledger)]) == 0 + out = capsys.readouterr().out + assert "RE-SCORED AFTER THE BANNER WAS LAST TOUCHED: 1" in out + assert "0 sit on a cell whose GRADE moved after the banner" in out + assert "1 on a cell that was only re-verified" in out + assert "0 could not be decided" in out + assert GRADE_UNCHANGED in out + + +def test_the_refusal_names_the_route_that_needs_no_owner_decision( + shallow_ledger: Path, tmp_path: Path, capsys +) -> None: + """A refusal whose only remedy is 'ask the owner' is why this tool sat unrun for twelve days. + + A throwaway clone shares no object store, so the shared-store objection that makes --unshallow + the owner's call does not reach it, and --root already takes a separate history source. + """ + card = tmp_path / "card.toml" + card.write_text( + '[[cell]]\nid = "C1"\nlast_verified = "2020-01-01"\nresidual = "BACKLOG #10"\n', + encoding="utf-8", + ) + assert main(["--scorecard", str(card), "--root", str(shallow_ledger)]) == 3 + err = capsys.readouterr().err + assert "git clone --single-branch" in err + assert "shares no object store" in err + assert "2020-01-01" in err + assert "--filter=blob:none" in err