You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Does the learned-v3 result survive when the confirmatory queries and support annotations are authored by a contributor who did not tune the development methodology or inspect selector outcomes?
Why now
Phase O confirmed that q_0011–q_0030 were held out from model-side inspection, but the same person authored them and tuned the development process. This leaves an author-time dependence concern even though the benchmark freeze itself was respected.
Hypothesis
Independent authoring will reduce the risk that query selection, gold annotations, or distractor composition reflect implicit development-time expectations. The resulting effect may be positive, null, or negative.
rules for structural validation and author access.
The author must not see learned-v3 outcomes, model scores, selector comparisons, or outcome-driven tuning notes.
Experimental design
Recruit a separate contributor to author and annotate approximately 20–30 confirmatory queries. Provide only the corpus, schema, authoring guide, distractor guide, and a written target mix. Require:
query, gold answer, and gold_support_ids annotation;
explicit difficulty, topic, question-family, and multi-hop labels;
multi-hop questions with at least two gold-support chunks;
typed distractor rationale where pools are constructed;
structural validation only during authoring.
After the authoring set is complete, run schema/support validation, create a hash manifest, freeze the artifact, and only then generate outcomes. Decide up front whether this set extends PG-Context-Select-v1 or is the confirmatory portion of Context-Select-v2 from #28. Prefer the latter if it avoids duplicate corpus and annotation work; document the decision before authoring.
Outputs
Independently authored query artifact with distinct version/identity.
Gold-support and distractor annotation record.
Author independence/access log sufficient to establish no outcome access.
Structural validation output and frozen hash manifest.
Frozen model outcomes and paired analysis report, if this issue includes evaluation after the freeze.
Acceptance criteria
The issue is complete when the authoring independence protocol is documented, the target set is validated and frozen before outcome inspection, provenance and hashes are recorded, and the resulting learned-vs-static comparison is reported regardless of direction or significance.
Research-integrity constraints
Do not let the author inspect model outputs or alter queries after seeing outcomes. Do not reject unfavorable queries without a pre-specified structural reason. Do not author replacement queries after outcome inspection. Do not tune learned_v3 or prompt thresholds for this set.
Question
Does the learned-v3 result survive when the confirmatory queries and support annotations are authored by a contributor who did not tune the development methodology or inspect selector outcomes?
Why now
Phase O confirmed that q_0011–q_0030 were held out from model-side inspection, but the same person authored them and tuned the development process. This leaves an author-time dependence concern even though the benchmark freeze itself was respected.
Hypothesis
Independent authoring will reduce the risk that query selection, gold annotations, or distractor composition reflect implicit development-time expectations. The resulting effect may be positive, null, or negative.
Frozen variables
Before authoring begins, freeze and publish:
learned_v3 - topk_pool_order;gold_plus_distractors;The author must not see learned-v3 outcomes, model scores, selector comparisons, or outcome-driven tuning notes.
Experimental design
Recruit a separate contributor to author and annotate approximately 20–30 confirmatory queries. Provide only the corpus, schema, authoring guide, distractor guide, and a written target mix. Require:
gold_support_idsannotation;After the authoring set is complete, run schema/support validation, create a hash manifest, freeze the artifact, and only then generate outcomes. Decide up front whether this set extends PG-Context-Select-v1 or is the confirmatory portion of Context-Select-v2 from #28. Prefer the latter if it avoids duplicate corpus and annotation work; document the decision before authoring.
Outputs
Acceptance criteria
The issue is complete when the authoring independence protocol is documented, the target set is validated and frozen before outcome inspection, provenance and hashes are recorded, and the resulting learned-vs-static comparison is reported regardless of direction or significance.
Research-integrity constraints
Do not let the author inspect model outputs or alter queries after seeing outcomes. Do not reject unfavorable queries without a pre-specified structural reason. Do not author replacement queries after outcome inspection. Do not tune learned_v3 or prompt thresholds for this set.
Dependencies
Provenance
.planning/CONFIRMATORY_INTEGRITY_AUDIT.md, sections 10 and 15.data/annotation/query_authoring_guidelines.md.data/annotation/distractor_guidelines.md..planning/CONFIRMATORY_BENCHMARK_FREEZE.md.