I am implementing this as a focused follow-up to #882 and #883.
Problem
The offline replay tool evaluates one configured base threshold. That makes it hard to see how changing the threshold affects fallback frequency, task quality, and cost on the same recorded cases. The routing benchmark in #762 shows why this matters: a policy can match a fixed target on completed tasks while costing much more.
Proposed scope
Extend offline replay to evaluate an explicit set of thresholds against the same recorded probability distributions and per-target outcomes. For each threshold, report fallback count, selected-target counts, total quality and cost, and regret relative to the best fixed target. Preserve the existing single-threshold report and deterministic output. Missing costs must stay unknown rather than becoming zero.
Document that thresholds should be chosen on development cases and checked on separate held-out cases. The checked-in fixture remains synthetic; it does not establish production routing quality.
Out of scope
- New provider calls or credential handling.
- Runtime routing or automatic threshold changes.
- A new fixture schema or duplicate replay implementation.
- Claims about production accuracy from synthetic data.
Acceptance criteria
- A user can supply multiple thresholds and get a stable comparison in one offline run.
- Boundary behavior matches the runtime rule: confidence below the threshold falls back.
- Tests cover ties, repeated or invalid thresholds, missing costs, fallback selection, and unchanged default output.
- Focused and repository-wide checks pass.
I am implementing this as a focused follow-up to #882 and #883.
Problem
The offline replay tool evaluates one configured base threshold. That makes it hard to see how changing the threshold affects fallback frequency, task quality, and cost on the same recorded cases. The routing benchmark in #762 shows why this matters: a policy can match a fixed target on completed tasks while costing much more.
Proposed scope
Extend offline replay to evaluate an explicit set of thresholds against the same recorded probability distributions and per-target outcomes. For each threshold, report fallback count, selected-target counts, total quality and cost, and regret relative to the best fixed target. Preserve the existing single-threshold report and deterministic output. Missing costs must stay unknown rather than becoming zero.
Document that thresholds should be chosen on development cases and checked on separate held-out cases. The checked-in fixture remains synthetic; it does not establish production routing quality.
Out of scope
Acceptance criteria