Skip to content

cargo test -p perry-runtime cannot attribute a regression: the failing SET differs between runs in both directions (0 / 11 / 13 failures on comparable trees) #10944

Description

@proggeramlug

This makes the "full suite green" condition on every piece of work in the One Path campaign currently unsatisfiable. Filing it separately because it blocks verification, not any one change.

The observation

While verifying #10943 I nearly reported "my change causes 14 regressions". It does not — I had no baseline. Taking one changed the conclusion:

arm failures
pristine upstream/main (train 253), my own worktree, changes stashed 13
the same worktree, changes applied 14

14 is not 13 + 1. The sets differ in both directions — 4 tests fail only in the baseline, 5 only with the change:

baseline only: json::stringify_tojson_probe::…::object_proto_signature_fast_match_agrees_with_the_validated_builder · json_tape::cached_read::…::materialized_descriptor_invokes_getter_through_rooted_fallback · json_tape::cached_read::…::materialized_mutations_override_sparse_cache_and_original_length · object::class_meta_registry::dense_parent_tests::re_registering_the_same_edge_flushes_nothing

with the change only: json::stringify_flat::…::flat_output_reuses_only_a_matching_object_prototype_signature · json::stringify_record_output::…::cached_empty_object_reuses_only_an_unchanged_receiver · json::stringify_tojson_probe::…::object_proto_signature_fast_match_rejects_a_replaced_prototype_address · object::proto_validity::…::a_structural_mutation_of_an_unmarked_object_does_not_bump_validity · string::tests::concat_memo_governor_stays_on_when_hits_are_frequent

Corroboration from the same night: three lanes ran the same command on comparable trees and got 0, 11 and 13 failures. One reported that all 181 tests in the affected modules pass under --test-threads=1.

Why: process-global memo counters, shared across tests in one binary

Every one of the unstable tests asserts a reuse count, and fails by exactly one:

assertion `left == right` failed: an unchanged live signature must reuse the prior negative verdict
  left: 2, right: 1

assertion `left == right` failed: an unchanged generation triple must reuse the recorded verdict
  left: 2, right: 1

These are memos the runtime keeps per process — JSON prototype-signature and toJSON verdict memos, proto_validity generations, the concat-memo governor, the class-meta dense-parent edge cache. The tests assert "this probe ran once". Any other test running concurrently in the same binary that touches the same memo makes it run twice. Nothing is wrong with the runtime; the tests are asserting a global count from a multi-threaded harness.

A second group looks environment- rather than memo-shaped and should be triaged separately: child_process::reactor::lifecycle_tests (3), pty::platform_impl::tests (2), stdlib_pump::tests (2) — plausibly timing/load-sensitive on a shared 16-core box running several lanes' builds.

What this costs right now

  • No lane can attribute a regression. A green-vs-red comparison is meaningless when the failing set moves on its own; I could not tell whether my own change broke anything, which is why it was not pushed.
  • It silently inverts the campaign's own standard. "Sabotage it and watch it redden" assumes reddening means something.

Recommendations

  1. Interim rule: cargo test --release -p perry-runtime -- --test-threads=1 on BOTH arms of any comparison. The 181-tests-pass-single-threaded datapoint says this is sufficient today. Cheap, and it makes verdicts trustworthy again.
  2. Make the memo assertions robust rather than the harness serial, as the real fix: have each test reset the memo it asserts on (several modules already expose test_reset_* helpers), or assert a delta it captured itself rather than an absolute count.
  3. Triage the pre-existing 13 on their own. They are not one bug: at least two families (memo-count, and process/pty/pump timing), and until they are green the suite has no clean baseline to compare against at all.

Reproduced on upstream/main @ train 253 (0fa391529), Linux x86-64, release profile.

Activity

  1. proggeramlug commented on Sep 22, 2026

    @proggeramlug
    ContributorAuthor

    Measured: every failure is interference, and the fix is a lint, not a patch

    The number that settles it

    Same binary, same tree (upstream/main train 253), one flag apart:

    --test-threads=1   4215 passed;  0 failed
    parallel (x6)      4198-4205 passed; 10-17 failed, a DIFFERENT set each time
    

    Zero genuine failures. Every failure this suite has produced for anyone tonight — my 13, the 11 and the 0 from other lanes — is interference. Nothing in the runtime is broken; the suite cannot attribute.

    Mechanism, with the smoking gun

    The flaky tests assert on process-global counters. json_tape::cached_read carried the premise in a comment:

    // The runtime suite is serial. This witness holds no managed values.
    static ROOTED_READS: AtomicU32 = AtomicU32::new(0);

    It is not serial. libtest runs tests in one process on many threads, so a sibling taking the same safepoint bumps the witness and the assertion fails by exactly one — left: 2, right: 1, the shape every one of these has.

    Partial fix pushed: fix/10944-per-test-memo-isolation (6e9c6852)

    Three sets of counters converted to per_test_global! — the mechanism the codebase already has for exactly this, which gives each test thread its own instance in a test build and expands to the plain static byte for byte outside one:

    • intl::segments_view — OPENS + four DECLINE_*
    • object::proto_validity — PROTO_VALIDITY, ANY_PROTOTYPE_MARKED
    • json_tape::cached_read — ROOTED_READS (and the false comment corrected)

    Over six parallel runs afterwards, none of those three modules appears again.

    Why that does not close this, and what does

    Six parallel runs after the fix still show 17 distinct failing tests, including names earlier runs never produced — builtins::fn_metadata, object::class_registry::dispatch::obj_dispatch_ic_tests, json::stringify_flat, four in json::stringify_tojson_probe, object::class_meta_registry. Converting statics one at a time is whack-a-mole across modules owned by several lanes, and the population is not bounded by anything I can enumerate.

    The systemic fix is to extend scripts/global_sink_isolation.py. It already runs in lint and already fails on a bare static that is neither thread-local nor declared via per_test_global! — but only for storage behind a GC clear helper. Extending its rule to any process-global counter a test asserts on makes this class structurally impossible, which is precisely what per_test_global!'s module docs argue for:

    A new sink cannot be added quietly, and a new reader never has to remember anything.

    That module's own history says this is recurring, not new: it cites #7665, #7671, #7672 and #7975 — "three flakes in two days", each diagnosed from a wrong VALUE rather than from the timing. This is the fifth.

    Interim rule stands

    --test-threads=1 on both arms of any comparison, until the lint covers it. It costs ~32 s against ~12 s, and it is the difference between a verdict and a coin flip.

    One more thing worth fixing separately

    The timing-shaped family — child_process::reactor (3), pty (2), stdlib_pump (2) — is unaffected by any of this and fails under load on a shared box. It needs its own triage and should not be counted against the memo class.

  2. added a commit that references this issue on Sep 22, 2026
  3. proggeramlug commented on Sep 22, 2026

    @proggeramlug
    ContributorAuthor

    Fixed as a ratchet — PR #10947

    Two commits: three conversions, and the gate that stops new instances arriving.

    6e9c6852 converts intl::segments_view's five counters, object::proto_validity's two and json_tape::cached_read's witness to per_test_global!. None of those three modules reappears across six parallel runs. The comment that stated the false premise — "The runtime suite is serial" — is corrected in place.

    5279c16d extends this script with a second rule: a bare static of a shared-mutable type, outside thread_local! / per_test_global! / perry_thread_local!, whose name appears inside an assert*!(…) in test code. Baseline scripts/global_sink_asserted_baseline.txt, 62 entries, may only shrink, with --asserted-no-raise-vs wired into the PR job beside the raw-handle-debt rule.

    Why a ratchet: six runs after the conversions still produced 17 distinct failures including names no earlier run showed, so the population is not enumerable and a sweep would touch several lanes' modules at once. Recorded entries get converted by whoever owns each file; nothing new arrives quietly.

    Why only assertions count: a first draft flagged any mention and produced 481 entries of mostly noise. A test that arms a feature flag or reads a census counter it never checks cannot be broken by a sibling. Tightened to assertions: 62, matching the actual failure shape.

    Proven able to fail — four self-test fixtures (hazard reported, three near-misses not), plus end-to-end: the gate passes on the baseline, a planted CANARY_10944 fails it by name, and --update-asserted refuses to absorb the canary.

    The three conversions are absent from the baseline rather than listed in it.

    Leaving this issue open until the 62 are worked down; the gate stops it growing meanwhile. The timing family (child_process::reactor, pty, stdlib_pump) is untouched and still wants its own triage.

  4. added 5 commits that reference this issue on Sep 22, 2026
  5. added 2 commits that reference this issue on Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugConfirmed defect or regression

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions