Skip to content

Kept external stream replay on the live run's drain schedule. - #39

Open
moedash wants to merge 2 commits into
moe/AI-198-if-pyext-2-core-repairs-pinfrom
moe/AI-198-if-pyext-3-replay-schedule
Open

moedash wants to merge 2 commits into
moe/AI-198-if-pyext-2-core-repairs-pinfrom
moe/AI-198-if-pyext-3-replay-schedule

Conversation

@moedash

@moedash moedash commented Oct 3, 2026

Copy link
Copy Markdown
Owner

This PR makes external stream replay take the same drains the live run took.

What changed?

  • An activation with no stream activity still records its segment once a subscription exists, so input and output share one replay schedule. A prerelease marker that left those segments out is rejected with a NondeterminismError rather than guessed at.
  • A replay marker is applied right before the first drain that can publish, and it no longer earns a drain of its own.
  • One handoff case reads the marker right before the step it checks, because every closing task now writes its own marker.

Part of AI-198 (epic AI-37).

Why?

Replays of runs that read input, publish and wait on an activity diverged. Without the empty segments, wait_condition predicates ran a different number of times on replay. A marker installed after the Signal jobs had already published wiped those records and failed the manifest check, and its extra drain put the activation one drain ahead of the task the marker came from. These are two fixes, but they share the activation path and test_output_runtime.py, so they ship as one PR with a commit each.

How did you test it?

Link to a test plan if any -

  • Unit Tests
  • Staging
  • End to End Tests

poe lint is clean. The new schedule and drain cases and the whole external stream suite pass on the dev server. With only the first fix, the handoff finalization case read the wrong marker. That's why its read moved in the same commit.

An empty activation still ran one drain, so the marker has to record it or
replay fires wait_condition predicates a different number of times. The
marker also goes in with the first set that drains, rather than after the
signal and update jobs have already published.
Behind a patch-only job set the install added a second drain, so every
wait_condition predicate fired once more than the recorded task fired it.
The dispatch moved to its own method so the schedule can be asserted.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant