Skip to content

Handoff: what to do on 1 October 2026 and in which order #95

Description

@martinfrancois

Handoff: prove PR #94, release 1.2.2 to the registry

State on 2026-09-21: the registry serves 1.2.0 (quality 100, 2.22x). The v1.2.1 GitHub release exists but its Tessl publish failed at the gate in scripts/validate_publish_ready.sh (tessl review run --threshold 100), because the released SKILL.md scores 91 under the current Tessl reviewer (#88, run 35561430400). PR #94 rewrites SKILL.md to score 100 (review run 01a0c22c-f9e8-7376-933c-240b76d3cf1d) without changing any rule; it has no hosted evidence yet. The maintainer decided: keep 1.2.1 GitHub-only, publish 1.2.2 once #94 is proven. Do this after java-functional-style-skill PR #18 and its 1.0.0 release, and after merging the optionals release PR, because this is the most expensive proof.

Steps

  1. git fetch origin && git checkout fix/skill-review-100, rebase onto main if behind, force-push with lease. Run the local validators from docs/agents/workflow.md.
  2. Quality review must still be 100 on the rebased head:
    tessl review run --workspace martinfrancois --threshold 100 --force skills/java-streams/SKILL.md
  3. scripts/pre_submit_gate.sh --plan-only prints the targeted probe order for this runtime change. Run the probes it names, one or two at a time, with scripts/run_eval_suite.sh <suite> <scenario> -- --wait -y. Every with-context result must be 100 before widening.
  4. Full main suite, both variants (80 credits for 4 scenarios): scripts/run_eval_suite.sh main -- --wait -y. The ratio must stay at or above 2.22x with all four at 100 with context.
  5. Reference, both variants per docs/agents/evals.md (120 for 6 scenarios), or with context only (60) if the budget is tight; state which in the PR: scripts/run_eval_suite.sh reference -- --wait -y.
  6. Regression with context only (190 for 19 scenarios): scripts/run_eval_suite.sh regression -- --wait -y. All at 100.
  7. Composition with the companion, from a checkout of java-functional-style-skill next to this one: scripts/run_composed_eval.sh ../java-streams-skill main -- --wait -y (40). All at 100.
  8. Record run IDs in the numbering notes and criteria-meta.json, fill the "Hosted Evidence" table in the PR body, push.
  9. Squash-merge fix(skill): tighten the stream workflow for the current Tessl reviewer #94 (title fix(skill): ...). Release Please opens "chore(main): release 1.2.2"; merge it. This time the publish gate passes (review 100) and the registry evaluates 1.2.2.
  10. Confirm on https://tessl.io/registry/martinfrancois/java-streams: version 1.2.2, quality 100, score at or above 2.22x. If the registry run drops below 2.22x, read tessl eval view <id> --json for the scenario that moved before changing anything; variance on one scenario is normal, a with-context miss is not.

Budget: about 400 credits with both-variant reference, about 340 with context-only reference. Together with the functional-style proof (about 300) and two registry evals this fits one monthly window only if nothing is repeated, so do the functional-style repository first and check tessl org usage before each stage.

Also open, lower priority: #44 (identity-mapper eval, now solved by the default solver; comment on the issue explains the re-scope), #93 (share the commitlint install block).

Ground rules for whoever picks this up

  • Do this on or after 1 October 2026, when the Tessl free plan's credit window resets (1,000 credits per month, 300 eval credits per day). Check first: tessl org usage.
  • Sign in with tessl login --device (the loopback flow prints nothing on a headless host). Any machine works; clone all three repositories side by side, because the composition runner in java-functional-style-skill reads the sibling checkouts by relative path (../java-streams-skill, ../java-optionals-skill).
  • Costs, so the budget can be planned: one scenario with both variants costs 20 credits, with context only 10, a quality review 10, a registry eval at publish time 30 to 60. Cells drop at random on this plan; tessl eval retry <run-id> creates a new run that reuses the completed cells. Never rerun a run because it looks slow; poll it with tessl eval view <run-id> --json.
  • Every hosted result must be recorded with its run ID in the numbering notes (evals/NUMBERING.md, evals-reference/NUMBERING.md, evals-regression/NUMBERING.md) and the per-scenario criteria-meta.json, never in docs/agents/* (the validator rejects run IDs there).
  • Any commit that touches the runtime bundle (skills/, and rules/ where it exists) invalidates every with-context result; freeze the runtime text before spending credits.
  • Commitlint rejects merge commits in these repositories, and the streams and optionals rulesets require branches to be up to date, so update a branch by rebase and force-push with --force-with-lease --force-if-includes (the maintainer approved that for these branches), never with a merge commit.
  • Model selection (--model) is a paid-plan feature and returns an entitlement error on the free plan; use the default solver.

AI Assistance (if used)

  • AI-assisted issue
  • I confirm I understand and reviewed this request

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions