Conversation
An empty activation still ran one drain, so the marker has to record it or replay fires wait_condition predicates a different number of times. The marker also goes in with the first set that drains, rather than after the signal and update jobs have already published.
Behind a patch-only job set the install added a second drain, so every wait_condition predicate fired once more than the recorded task fired it. The dispatch moved to its own method so the schedule can be asserted.
This was referenced Oct 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR makes external stream replay take the same drains the live run took.
What changed?
NondeterminismErrorrather than guessed at.Part of AI-198 (epic AI-37).
Why?
Replays of runs that read input, publish and wait on an activity diverged. Without the empty segments,
wait_conditionpredicates ran a different number of times on replay. A marker installed after the Signal jobs had already published wiped those records and failed the manifest check, and its extra drain put the activation one drain ahead of the task the marker came from. These are two fixes, but they share the activation path andtest_output_runtime.py, so they ship as one PR with a commit each.How did you test it?
Link to a test plan if any -
poe lintis clean. The new schedule and drain cases and the whole external stream suite pass on the dev server. With only the first fix, the handoff finalization case read the wrong marker. That's why its read moved in the same commit.