Skip to content

[Phase P] Create an independently authored confirmatory query set #29

Description

@pro-utkarshM

Question

Does the learned-v3 result survive when the confirmatory queries and support annotations are authored by a contributor who did not tune the development methodology or inspect selector outcomes?

Why now

Phase O confirmed that q_0011–q_0030 were held out from model-side inspection, but the same person authored them and tuned the development process. This leaves an author-time dependence concern even though the benchmark freeze itself was respected.

Hypothesis

Independent authoring will reduce the risk that query selection, gold annotations, or distractor composition reflect implicit development-time expectations. The resulting effect may be positive, null, or negative.

Frozen variables

Before authoring begins, freeze and publish:

  • the corpus snapshot and chunk IDs;
  • query-authoring and distractor guidelines;
  • required schema and split metadata;
  • primary comparison learned_v3 - topk_pool_order;
  • secondary status of gold_plus_distractors;
  • selector, estimator, prompt regime, scoring weights, evaluator version, and token budget;
  • inference policy from [Phase P] Define inference policy for multi-model and multi-domain experiments #25;
  • rules for structural validation and author access.

The author must not see learned-v3 outcomes, model scores, selector comparisons, or outcome-driven tuning notes.

Experimental design

Recruit a separate contributor to author and annotate approximately 20–30 confirmatory queries. Provide only the corpus, schema, authoring guide, distractor guide, and a written target mix. Require:

  • query, gold answer, and gold_support_ids annotation;
  • explicit difficulty, topic, question-family, and multi-hop labels;
  • multi-hop questions with at least two gold-support chunks;
  • typed distractor rationale where pools are constructed;
  • structural validation only during authoring.

After the authoring set is complete, run schema/support validation, create a hash manifest, freeze the artifact, and only then generate outcomes. Decide up front whether this set extends PG-Context-Select-v1 or is the confirmatory portion of Context-Select-v2 from #28. Prefer the latter if it avoids duplicate corpus and annotation work; document the decision before authoring.

Outputs

  • Independently authored query artifact with distinct version/identity.
  • Gold-support and distractor annotation record.
  • Author independence/access log sufficient to establish no outcome access.
  • Structural validation output and frozen hash manifest.
  • Frozen model outcomes and paired analysis report, if this issue includes evaluation after the freeze.

Acceptance criteria

The issue is complete when the authoring independence protocol is documented, the target set is validated and frozen before outcome inspection, provenance and hashes are recorded, and the resulting learned-vs-static comparison is reported regardless of direction or significance.

Research-integrity constraints

Do not let the author inspect model outputs or alter queries after seeing outcomes. Do not reject unfavorable queries without a pre-specified structural reason. Do not author replacement queries after outcome inspection. Do not tune learned_v3 or prompt thresholds for this set.

Dependencies

Provenance

  • Phase O author-dependence limitation: .planning/CONFIRMATORY_INTEGRITY_AUDIT.md, sections 10 and 15.
  • Existing authoring workflow: data/annotation/query_authoring_guidelines.md.
  • Existing distractor rules: data/annotation/distractor_guidelines.md.
  • Frozen benchmark example: .planning/CONFIRMATORY_BENCHMARK_FREEZE.md.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    phase-pPhase P external-validity researchresearchResearch planning or analysis

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions