Skip to content

Add a ci-status skill that reports a PR's real CI state - #33

Open
Babissimo wants to merge 1 commit into
mainfrom
feat/ci-status
Open

Babissimo wants to merge 1 commit into
mainfrom
feat/ci-status

Conversation

@Babissimo

@Babissimo Babissimo commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

ClickUp: 2.1 Add a ci-status command that reports a PR's real CI state

Summary

gh pr checks and gh run watch have each reported green for a PR that was untested, conflicting, failed or unreviewed. On retina-server that happened in at least eight distinct ways, each kept as a separate memory note that worktree sessions never load. This adds a ci-status skill to core whose script reads the PR's runs directly and exits non-zero unless every gate passed:

python3 "${CLAUDE_PLUGIN_ROOT}/skills/ci-status/scripts/ci_status.py" <pr> -R <owner>/<repo> [--watch]

Exit status: 0 green, 1 failed or will not run until someone acts, 2 the question was wrong, 3 not settled yet, 4 GitHub could not be read.

What it handles

  • Runs for the PR's recorded head only, checked against its branch on GitHub and the local branch. While the head lags its branch nothing about the old head is judged. Unpushed local commits fail the reading.
  • A conflicting PR, which gets no merge ref and so no CI, is reported as never going to run.
  • Runs whose jobs all skipped (the run a title or body edit fires) are set aside. A newer run supersedes an older one job by job, so an edit's run cannot hide tests that failed, or that a timeout cut short, before it.
  • Cancelled duplicates from a stack push are set aside, and lend no job that never queued. The latest attempt decides, and failed earlier attempts are still listed, because a rerun overwrites the run's conclusion.
  • Job conclusions decide, not the run's, since a run can read completed while its jobs are still going.
  • Other apps' checks and commit statuses count as gates.
  • The Claude review counts only when a bot comment from the latest attempt of this head's review run links back to it and carries a fully ticked checklist. It reports when the PR edits the review workflow, or when the head carries an older copy of it than the default branch; either way the action skips itself without a word.
  • --watch polls until the answer settles. It rides out outages, and after a few minutes it calls a wait that has not cleared final: no runs, no review run, or a head that has not moved.

The script's docstring records what it cannot see.

Testing

  • 126 unit tests against a fake GitHub (tests/ci-status/), a few of them against real git for the local-branch comparison, run by a new ci-status tests workflow with SHA-pinned actions and a read-only token.
  • Every branch was checked by deliberately breaking it: 85 such changes, and each one turns at least one test red.
  • Run against live PRs:

Not done

The last three review passes each found one misreport, all now fixed:

  1. A test run cancelled by a timeout, followed by a body edit, read green.
  2. The fix for that let a stack push's cancelled duplicate lend jobs that never queued, which turned a green PR red (retina-server #622).
  3. The first fix for #622 also withheld tests that had queued, so a run cancelled by hand before its tests started, followed by a body edit, read green.

These points from the passes are left as they are:

  • --commit on a main merge whose deploy job was cancelled in a burst, and so superseded by the next merge's run, reads failed. It judges the commit's runs, not whether the change shipped.
  • Under --watch, a continue-on-error job that fails while its run is still going ends the watch with exit 1, even if the run then succeeds. The docstring says so.
  • After GitHub's "Update branch" rebases the PR, a local copy of the branch from before it reads as diverged, and fails the reading, where main changed lines beside the PR's.
  • A run's own allowed failure still gates while a job lent by an older, still-running run is unfinished. It needs two overlapping runs for one head plus continue-on-error.
  • A job waiting on an environment approval reads as not settled (exit 3), not as needing action. No PR workflow in the org uses environments.
  • A job list kept once its run reads completed is not read again should the same attempt later read in progress. The verdict still follows the run's own status.

Review notes

  • The core plugin goes from 0.6.0 to 0.7.0 so installs pick the skill up.
  • This repo's own review workflow does not set track_progress, so its runs post nothing to link. ci-status says so rather than calling such a PR reviewed.

🤖 Generated with Claude Code

gh pr checks and gh run watch have each reported green for a PR that was untested, conflicting, failed or unreviewed. On retina-server that happened at least eight ways, each kept as a separate memory note that worktree sessions never load, so every session rediscovered the trap it hit.

ci_status.py judges the runs for the PR's recorded head instead of the checks table. It sets aside runs whose jobs all skipped (a title or body edit), cancelled duplicates (a stack push) and non-gating events; lets a newer run supersede an older one job by job, so an edit's run cannot hide tests that failed before it; takes each run's latest attempt while listing failed earlier ones (a rerun overwrites the conclusion); decides by job conclusions rather than the run's; reports a conflicting PR as never going to run; checks the PR's head against its branch on GitHub and the local branch; counts other apps' checks and commit statuses; and counts the Claude review only when a bot comment from the latest attempt of this head's review run links back to it with its checklist ticked. It also says when the head's copy of the review workflow differs from the default branch's, since the action then skips itself without a word until the branch is rebased.

What may yet arrive (a run, the review run, the head following its branch) is 'not settled' in one reading, because GitHub records no push time to judge the wait against; --watch calls it final once it has seen it last a few minutes. The exit status separates failed (1) from a wrong question (2), not settled (3) and GitHub being unreadable (4), so neither a 502 nor a typo reads as a red build.

It lives here rather than in one repo because every repo scaffolded by setup-repo carries the same review workflow and the same gh habits. Its tests run in a workflow of their own so a failure is labelled as theirs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant