Skip to content

Epic: earn the complete permission-review journey, from a real change to voluntary reuse #778

Description

@pengfei-threemoonslab

PM backlog review — 2026-10-01

Application-agent PR proof (#868) remains the primary product track. Host correctness is bounded maintenance. This review organizes the 100 open issues; it does not establish delivery, user adoption or release qualification.

Workstream Current focus Exit evidence
Application comparisons #908 measurement, then #909 + #874, #910, #913 + #914; #911/#912 next #868's 10 useful, correct, complete cases; development and holdout counts reported separately on the released build
Host maintenance P1 #918, #920, #921; P1 #698 diagnosis Correct expansion/scope wording on each pinned reproduction; measured large-checkout latency with timeouts retained
Control correctness P0 #610 Empty planning completion cannot confer verifier publication, merge or completion authority; compatible runtime/schema/consumer behavior
Workflow and adoption #915 application delivery; remaining #780/#855 hosted evidence; #839 → #840 for the host decision/correction journey Actual hosted results and reviewer interpretation; voluntary use stays a separate observed outcome
Qualified claims #572 and its evidence dependencies Deferred, explicit qualification obligations; no substitution of advisory or engineering evidence
Shared maintenance Existing build, reader and diagnostics backlog Bounded acceptance and a reproduced current need before pickup

Evidence, not a delivery forecast. The latest #868/#908 ledger (2026-09-30, main 6ced6f70) reports 1 Q2 case / 49 development PRs, 9 with rows and zero false-complete answers in that sample. These are reported development measurements, not a fresh replay, adoption counts or representative market rates. The Oct 14 / Nov 13 / Dec 13 checkpoints remain evidence reviews.

Execution order. Resolve #908's corpus preservation and holdout-freeze prerequisite before merging the reader changes it gates. #909 owns list-expression evaluation; #874 owns parameter/constructor/instance flow. #910 owns object/built-in reader work, with #866 retaining its specific regression acceptance. #913 supplies richer observed effects; #914 supplies readable findings. #915 can prepare integration and ship with existing truthful text rather than waiting for #914. #916 is a bounded framework-selection study, not authorization to add a reader.

Avoid duplicate work. #584 is covered by #909 but stays open until its acceptance is demonstrated. #655/#867 need an acceptance-to-evidence reconciliation against the delivered application route before new implementation is selected. #795/#812 have delivered slices; #780 has ten recorded hosted runs; #855 still owes hosted digest proof. Do not repeat completed work or close residual obligations solely because a PR merged.

Queue policy. Every issue has one workstream:* label and one P0/P1/P2 label. Priority describes importance within its workstream; it is not a commitment to start every P1. status:deferred preserves an explicit prior deferral or conditional decision. #563 retains its historical P0 qualification significance with that status; it is not promoted into the advisory application sprint. Existing assignees are preserved; name an accountable implementer when work is picked up rather than assigning the entire backlog to one person.

Next product decisions. #908 must record a feasible freeze sequence for its >=30 holdout PRs created after Oct 1; no backdating or silent threshold change. #772 remains owner-deferred on privacy/noise. #502 owns contribution policy; #917 is a proposed example requiring version-pinned evidence, not a verified product defect. The recorded #830 and per-note outreach requirements remain in their owning issues; this triage sends no outreach and changes none of those requirements.

The dated selection above supersedes older execution order where they differ. Historical observations and acceptance below remain intact.


Selected execution — 2026-09-24

Days 1–5 execution is selected and recorded in PR #869. Application-agent proof (#868, milestone 13) is primary; host correctness is bounded maintenance. Accountable product owner: @pengfei-threemoonslab; technical executor: Codex current task. #580/#655 design precedes #867 and the reproduced #864/#865/#866 readers. #610 is reproduced; #787 is the bounded recipe repair. Milestones 9–11 were reconciled, retaining Oct 14 / Nov 13 / Dec 13 as evidence checkpoints. #830 and external outreach holds remain unchanged. Older recruitment-first and hosted-deferred text is superseded by this dated selection; no external value or ten-case result is claimed.


Earlier sequencing below is retained as history and superseded where it conflicts with this dated execution selection.

Product outcome

A reviewer can turn a real agent-permission change into an explicit, evidence-grounded decision, carry out the next step, verify the selected correction and choose to use the workflow on the next eligible change.

Real change → understand → decide → verify correction → voluntary next use.

This is the sole product program epic. #811 owns the selected experience; its implementation children #839 and #840 close the fact-to-decision and correction-verification gaps. #791 remains discovery coordination. No separate dashboard, control plane, research cohort or approval engine is introduced.

Current sequence — 2026-09-18

The owner's 2026-09-16 decision controls execution: build the selected work in milestones 10–12, run #830 on a released build, and begin external recruitment/outreach only after that gate passes and an explicit owner decision permits it. Internal preparation, fixture replay and product design may proceed now. This update performs no outreach and does not change #830's thresholds or per-note approval.

The original 2026-09-14 program clock and Oct 14 / Nov 13 / Dec 13 checkpoints remain. A missing participant is a dated shortfall, not value. Each participant's existing four-week observation window starts at first value; preserve any prior observations and never reset a window for this plan. #811's known-change decision-support hypothesis and #830's novel automatic author-actionable finding hypothesis retain separate measures and denominators.

Why this is the next product work

The evidence chain has useful foundations, but a generic "Does the team intend this declared capability change?" leaves the reviewer to infer the available choices and the condition a rerun would verify. More findings or severity labels do not supply business intent.

The selected product question is: Does the result let a non-author reviewer make and verify a concrete decision without a maintainer translating it? A justified decision to retain a known expansion is eligible value. Correctly identifying missing evidence is useful recovery; an incomplete comparison remains outside #571's first-valid-result definition.

Current facts — observed 2026-09-18

One decision-ready result

Reviewer question Product answer Owner
What changed? Exact before/after, source, original refs #795 / #819
What does this establish? Supported static semantics and direction; no invented runtime or business outcome #816 delivered; #820 / #827 and bounded readers
What is still unknown? Compared scope, partial/unread input, named recovery where established #812 / #821 / #808 / #822
What must I decide? One case-specific missing-intent question and relevant conditional options #839
Who acts, and what happens next? Existing operational control or advisory owner-to-be-assigned; exact verification target #839, existing fix/control contracts
Did the correction address that change? Original B→C and B→F, selected correction plus remaining/new deltas and current coverage #840
Will this help on the next change? Actual eligible opportunity and voluntary review consumption #571 / #796, after #830

Start with the supported shell-permission route. A refund demonstration from an application tool surface is a separate scenario and cannot be inferred from a host allow rule. Agent/framework breadth follows demonstrated value and the existing support boundaries.

Roadmap and phase exits

Phase / checkpoint Selected result Engineering proof Product evidence
Phase 1, 2026-10-14 Truthful, concrete facts a reviewer can use Close remaining release/presentation proof around #795/#812/#792/#781; distinguish no-change from unknown; first #798 case Internal cold-reader readiness and explicit shortfalls only; #653 recruitment remains gated
Phase 2, 2026-11-13 One supported change supports a decision and a verified correction Existing fields/direction/coverage work, then #839 + #840 inside #811; B/C/F/R and negative controls; source/package/hosted evidence separately Record whether independent internal readers identify fact, limits, decision, action and correction correctly; no external-adoption claim
Phase 3, 2026-12-13 Establish automatic value, then observed reviewer value and reuse when permitted #796 exercised route; #830 on released Phase 2 build, retaining V1 unchanged Only after #830 and owner approval: #653/#571 existing cohort, concrete decisions, real corrections and next eligible use; otherwise dated shortfall and bounded next decision

Dates are evidence-review checkpoints, not promised releases, recruitment dates or automatic permission to start outreach. Phase 3 engineering/gate readiness precedes outreach; its external-value outcome is necessarily observed afterwards. Do not require an observation produced only by outreach before allowing the prerequisite gate to run. Do not close the user outcome as achieved merely because the readiness work or V1 is complete.

Delivery order and capacity

Preserve the owner's Sept 16 engine order. Phase 1's merged fixes and #837/#838 are delivered inputs; remaining installed proof and owner-reviewed publication remain. Proposed 1.1.0/1.2.0 in the prior plan are release targets, not already published versions.

Phase 2: #819 → #820 → #821 → #808/#822 → #809 → #827 → #824 → #823 → #825 → #826 → #807 → #702 → #811, including #839 then #840 → #780/#570 owner-approved hosted evidence → #828 decision → conditional #829 → owner-reviewed release. A supported shell prototype may be designed earlier without expanding coverage claims or silently changing this delivery order.

Keep the capacity assumption of two engineers plus product/research ownership and about 20% engineering reserve. The new implementation slices are provisionally 3–5 and 5–8 engineering days after prerequisites; they are part of #811, not an additional platform project. Validate those estimates at pickup. If the date is missed, name delivered scope and the remaining dependency; do not drop an identity/coverage test or manufacture user evidence.

Measurement: three distinct questions

  1. Engineering readiness: are facts, direction, conditional guidance and correction explanations correct and reproducible? Require zero false correction confirmations across the selected adversarial matrix. Internal task timing measures readability, not adoption.
  2. Automatic discovery: Value gate before any outreach: measure automatic author-actionable value on a released build #830's M1–M6, strata, novelty, yield, precision/noise, independent adjudication and released-build requirements stay unchanged. Manual examples and known intentional changes cannot raise M4.
  3. Reviewer decision value: Run the published advisory host-diff pilot: first reviewer value and the next eligible change #571's existing first-value definition, full denominators and four-week window stay unchanged. Add decision/rationale, assistance, missing evidence, next owner/action, requested static correction and whether it was correctly verified. Count a justified retention, a correction request and useful recovery separately. A dispute never suppresses a finding.

Use the same cohort/ledger. Preserve invited, attempted, first_valid_result, first_value, second_change_eligible, second_change_observed. Compare the ordinary workflow and Agents Shipgate with balanced ordering where practical; record total reading/decision/correction time, errors and assistance. Scanner latency and positive comments are not saved review time.

Evaluate #571's existing ladder in order, on the same route: Continue at ≥2 unaided first values and ≥1 observed second eligible change; Stop at zero first value with ≥1 first valid result; Narrow otherwise. Report all denominators. No second eligible change means repeat use remains unobserved. Voluntary use is reported separately from automated CI execution and from the existing second-change threshold.

Strategic caution: a known-change result can improve a real decision without finding a novel problem. #830 and #811 therefore test different propositions. This roadmap preserves the owner-selected outreach gate; changing it later requires an explicit decision, not a metric substitution.

Authority, privacy and conditional expansion

Program acceptance

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

P1Next after P0; blocks other work or ships a misleading resultarea:releaseRelease pipeline, packaging, and safety qualificationepicTracking issue coordinating a group of related issuesworkstream:adoptionReviewer journey, distribution, workflow delivery and observed reuse.

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions