Skip to content

Latest commit

 

History

History
455 lines (391 loc) · 21.8 KB

File metadata and controls

455 lines (391 loc) · 21.8 KB

Integration Recipes

GitHub Actions

Host-only advisory PR review (no manifest)

A repository whose pull requests change coding-agent host configuration needs no shipgate.yaml for PR review. Use examples/github-actions/14-host-only-advisory-pr.yml: the Action compares each PR with its base branch in Git history and posts one comment that later pushes update. What it finds never fails the job, and missing base history shows in the comment as Host capability comparison unavailable. The job does fail on a setup or execution error — the install, a non-zero CLI exit such as an invalid shipgate.yaml added by the PR, or a comment API error other than a missing permission. The examples README covers those, permissions, forks, pinning and what each comment means. The recipes below apply to a repository with a shipgate.yaml.

With a manifest

The public Marketplace action installs from its tagged source by default; set shipgate_version when you want the action to install a pinned PyPI package version.

name: Agents Shipgate

on:
  pull_request:

permissions:
  contents: read

jobs:
  agents-shipgate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
        with:
          fetch-depth: 0
      - id: agents-shipgate
        uses: ThreeMoonsLab/agents-shipgate@v1.2.0
        with:
          config: shipgate.yaml
          ci_mode: advisory
          diff_base: target
          shipgate_version: '1.2.0'

To post PR comments, set:

permissions:
  contents: read
  pull-requests: write

with:
  pr_comment: "true"

To apply organization policy packs from CI, pass a comma- or newline-separated list:

with:
  policy_packs: policies/org-release.yaml,policies/security.yaml

To make the verifier merge verdict load-bearing in CI, configure fail_on_merge_verdicts. The recommended agent-PR policy is either to block only blocked, or to require can_merge_without_human == true in a separate workflow step:

with:
  fail_on_merge_verdicts: blocked

This is opt-in. When configured, the action fails closed if the installed agents-shipgate package does not emit verifier.json, so pinned older versions should be upgraded before enabling the input.

Action outputs:

Output Meaning
decision Release decision (blocked, review_required, insufficient_evidence, or passed). v0.8+; insufficient_evidence added v0.14. Use this as the CI gating signal. Switch on the value with a review_required fallback for unknown future values.
merge_verdict PR/control projection of decision (mergeable, human_review_required, insufficient_evidence, blocked, or unknown). Used by fail_on_merge_verdicts when configured; this is an explanatory projection, not a second release gate.
can_merge_without_human true only for a verified passed result or a completed deterministic not_applicable skip.
agent_control_state Authoritative operational state from verifier.json.control.state: complete, agent_action_required, review_publishable, or human_review_required. review_publishable authorizes commit/push/PR updates and denies merge and completion.
agent_control_reason Deterministic reason from verifier.json.control.reason.
agent_controller_must_stop One-cycle compatibility mirror of verifier.json.control.must_stop.
agent_controller_stop_reason One-cycle compatibility mirror of verifier.json.control.stop_reason.
agent_controller_completion_allowed One-cycle compatibility mirror of verifier.json.control.completion_allowed.
blocker_count Number of blockers in release_decision.blockers. v0.8+.
review_item_count Number of review items in release_decision.review_items. v0.8+.
ci_would_fail true/false — whether the active fail policy would fail CI. v0.8+.
status Legacy report summary status, such as release_blockers_detected. Baseline-blind; preserved for v0.7 compat.
critical_count Unsuppressed critical finding count.
high_count Unsuppressed high finding count.
medium_count Unsuppressed medium finding count.
baseline_new_count New finding count when baseline is set.
baseline_matched_count Baseline-matched finding count when baseline is set.
baseline_resolved_count Resolved baseline finding count when baseline is set.
adk_agent_count Statically detected Google ADK agent count.
adk_dynamic_toolset_count Google ADK dynamic or unresolved toolset count.
report_json Path to report.json.
report_markdown Path to report.md.
report_sarif Path to report.sarif.
verifier_json Path to verifier.json.
verify_run_json Path to verify-run.json, which validates against verify-run-schema.v4.json.
run_id Stable verify-run input identity from verify-run.json.run_id.
pr_comment_markdown Path to pr-comment.md.
exit_code Agents Shipgate CLI exit code. Matches release_decision.fail_policy.exit_code.

The action runs agents-shipgate verify, which writes Markdown, JSON, SARIF, packet JSON, verifier JSON, verify-run JSON, and PR-comment Markdown artifacts. It intentionally emits packet.json only for the packet; pr-comment.md is the human PR surface. Run agents-shipgate agent control --workspace . first and require it to succeed: it validates current-control.json against every artifact it binds and against the live repository, and exits 4 (workspace_changed) once HEAD or the working tree has moved since the decision. Reading the pointer file directly does not check that — it returns the old state until verification is re-run. Then read current-control.json, agent-handoff.json for the compact agent handoff, verifier.json for detailed control context, verify-run.json for reproducibility metadata, and report.json.release_decision.decision for the gate. Capability diffs and capability_review.top_changes are supporting/provisional review context. Verify never fetches; use fetch-depth: 0 on checkout or fetch the base ref before the action when diff_base: target is set. An explicit head_ref is scanned from an isolated archive; without it, the checked-out workspace is scanned. Upload report.sarif to GitHub code scanning from your workflow if you want SARIF annotations. Policy-pack findings use stable policy rule IDs as SARIF ruleId values when present, so the first upgrade run can close/reopen existing GitHub code-scanning alerts whose identity was previously the built-in Shipgate check ID.

After adoption, choose an explicit merge policy in the workflow rather than leaving advisory mode load-bearing. 07-block-on-blocked-verdict.yml blocks only when merge_verdict == 'blocked'; 08-require-mergeable.yml requires can_merge_without_human == true; 11-fail-on-insufficient-evidence.yml fails only on insufficient_evidence. Strict, baseline, SARIF, Check Run and multi-config recipes are in examples/github-actions/; the full input and output catalog is action.yml.

CI is advisory by default. Strict mode exits 20 only on unsuppressed critical findings, so on an existing project it fails on the backlog the first time it runs. Record that backlog as a baseline, then gate on what is new:

# 1. see what strict would do today — expect exit 20 if there is any backlog
agents-shipgate scan --config shipgate.yaml --ci-mode strict
# 2. accept the current findings as the baseline
agents-shipgate baseline save --config shipgate.yaml --out .agents-shipgate/baseline.json
# 3. strict from here: fails only on findings the baseline does not carry
agents-shipgate scan --config shipgate.yaml --baseline .agents-shipgate/baseline.json --ci-mode strict

Severity and failure thresholds are configurable in the manifest (checks.severity_overrides, ci.fail_on) — see baseline.md.

For source-only testing in this repository:

- uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd
  with:
    fetch-depth: 0
- uses: actions/setup-python@a309ff8b426b58ec0e2a45f0f869d46889d02405
  with:
    python-version: "3.12"
- run: python -P -m pip install -e ".[dev]"
- run: agents-shipgate verify --workspace . --config shipgate.yaml --base origin/main --head HEAD --ci-mode advisory --format json

Local Diagnostics

agents-shipgate init --workspace . --write
agents-shipgate doctor --config shipgate.yaml
AGENTS_SHIPGATE_LOG_FORMAT=json agents-shipgate scan --config shipgate.yaml --verbose
agents-shipgate scenario suggest \
  --from agents-shipgate-reports/report.json \
  --out agents-shipgate-reports/suggested-scenarios.yaml

The scenario YAML is derived from report.json.suggested_scenarios[] and fans static findings out into concrete sandbox/adversarial validation steps. Baseline-matched findings remain in this export because they are accepted debt, not resolved risk.

Claude Code hooks (local advisory)

After agents-shipgate verify and CI are working, install project-scoped Claude Code hooks for faster local feedback:

agents-shipgate install-hooks --target claude-code --write

The installer writes .claude/settings.json and .claude/hooks/agents-shipgate.py. The PostToolUse hook runs a cheap agents-shipgate trigger check after Edit|Write|MultiEdit so Claude Code gets immediate context when an edit touches an agent-related surface. It evaluates the edited paths without the manifest-present force-run rule, so irrelevant docs edits do not produce a nudge just because the repo is opted in. The Stop hook runs full agents-shipgate verify only when the working tree or current branch has a relevant change that has not already been checked, then routes on the authoritative verifier.control.state: complete ends the turn silently, agent_action_required soft-blocks the Stop once and names the one exact remaining command, and human_review_required lets the turn end with a hand-off notice — a Stop-hook block forces the agent to keep working, which is the opposite of what must_stop means.

Without a configured manifest, when every changed file is host configuration — .claude/settings.json, .mcp.json, .cursor/mcp.json, .codex/config.toml, .vscode/mcp.json or a GitHub workflow — the Stop hook runs agents-shipgate diff instead of advising you to initialize a manifest. It stays quiet when no row widens what the agent can do and none is of unknown direction. An edit the engine cannot order, such as a hook command or MCP argument edit, or a setting no documented rule ranks, is not called a widening but is named under its own heading, never silently dropped (#820). A workflow step moved to a different action reference, such as a pinned SHA to @main, is a non-widening row, so the hook stays quiet about it; diff and the PR comment still show it. The same holds for an agent launch in a workflow whose settings change without gaining a documented widening rule, for a checkout's ref, and for an edit to an argument input the audit does not read, which is compared by a digest, named in the row and named as a limit in audit --host. A run: that mentions an agent CLI and that the audit does not read is different: it gives no row in diff or the PR comment, whatever is edited, and is named only as a limit in audit --host. An agent launch that gains a rule, such as a plain claude_args gaining --dangerously-skip-permissions on any of its lines, widens, and the hook announces it (#823), unless the launch may be a step of that job the audit did not read, rewritten, or the job's launch held before a ${{ }} expression or an unread argument input the rule is read from, or the rule moved in from a job the launch left, as a renamed job's does. It names each widening row once, and repeats the announcement only when the change or its rows change. A missing base ref, an incomparable inventory or unparsed output is never quiet. When host configuration changes beside other files, the Stop hook still compares the host configuration that way and advises on the other files separately, in one message. The PostToolUse hook never nudges on host configuration edits, with or without a manifest, because the Stop hook compares them. Instruction files, skills, commands, rules and policies keep the route above, and a repository with a manifest always runs verify. The hook reads configuration, not the agent's runtime permissions.

Within one Claude Code session, the hooks do not repeat an advisory the agent has already seen. Each changed path remembers the last verdict the hooks reached for it and the verdicts already announced. An edit that brings no new verdict for any path it names stays quiet, even though it is evaluated again. A new path, a different verdict, a different base or manifest, or a new session is announced. So is every result that could not read its input. This is presentation only: CI and verify evaluate the whole change regardless.

These hooks are advisory local feedback. Local setup failures such as a missing CLI or unavailable base ref are surfaced as context, and verifier output the hook cannot parse is surfaced as an explicit warning rather than treated as a pass. A verify that exits with anything other than 0 or 20 is surfaced as context naming the exit, and does not block the Stop. That includes exit 2 when the default agents-shipgate-reports is refused because it holds repository content, such as committed files or stray notes beside the reports: a change that would otherwise soft-block on agent_action_required ends the turn with that context instead, so resolve the refusal it names (see troubleshooting). They are not a trust boundary and not a replacement for CI. CI should continue to run the GitHub Action or an equivalent agents-shipgate verify command, and CI's report.json.release_decision.decision remains authoritative.

CI installation trust

The following recipes use python -P to keep the checkout off Python's implicit module search path during installation (Python 3.12 or newer). Use a trusted Python executable and installed CLI on PATH, and do not supply a checkout through PYTHONPATH. -P does not neutralize an attacker-controlled pipeline, explicit imports, environment variables, or an editable package's build backend. The source-only editable-install example above is for trusted source testing.

  • GitLab: use a protected/parent pipeline definition from a trusted source when scanning an untrusted checkout; a merge-request-controlled job can alter its own commands.
  • CircleCI: keep the configuration and any setup/continuation configuration trusted when analyzing untrusted changes; checkout isolation alone does not authenticate the job definition.
  • Jenkins: use a trusted Jenkinsfile or shared library for the scan stage; do not execute a proposed Jenkinsfile as the authority for its own review.

These examples document the installation boundary; they do not claim a hosted security test on GitLab, CircleCI or Jenkins.

GitLab CI

First-class GitLab CI recipes live in ../examples/gitlab-ci/:

  • advisory rollout;
  • strict mode with a baseline;
  • SARIF-or-artifact retention;
  • monorepo multi-config scans;
  • tool-source-change triggers.
agents-shipgate:
  stage: test
  image: python:3.12
  script:
    - python -P -m pip install --pre "agents-shipgate==1.2.0"
    - agents-shipgate scan --config shipgate.yaml --ci-mode advisory --format markdown,json,sarif
  artifacts:
    when: always
    expire_in: 1 week
    paths:
      - agents-shipgate-reports/

GitLab SARIF report ingestion is tier/version dependent. Always retain agents-shipgate-reports/ as path artifacts; enable artifacts:reports:sarif only where your GitLab instance supports it.

CircleCI

First-class CircleCI recipes live in ../examples/circleci/:

  • advisory rollout;
  • strict mode with a baseline;
  • SARIF artifact retention;
  • monorepo multi-config scans;
  • tool-source-change triggers.
version: 2.1

jobs:
  agents-shipgate:
    docker:
      - image: cimg/python:3.12
    steps:
      - checkout
      - run: python -P -m pip install --pre "agents-shipgate==1.2.0"
      - run: agents-shipgate scan --config shipgate.yaml --ci-mode advisory --format markdown,json,sarif
      - store_artifacts:
          path: agents-shipgate-reports
          destination: agents-shipgate-reports

Jenkins

stage('Agents Shipgate') {
  steps {
    sh 'python -P -m pip install agents-shipgate'
    sh 'agents-shipgate scan --config shipgate.yaml --ci-mode advisory'
    archiveArtifacts artifacts: 'agents-shipgate-reports/**', allowEmptyArchive: true
  }
}

MCP server (optional)

For coding agents without comfortable shell access (Cursor, restricted harnesses), Agents Shipgate can serve read-only static tools over an MCP stdio server. It is a thin wrapper over the same deterministic projections the CLI uses: shipgate.check, shipgate.preflight, shipgate.explain, and shipgate.capabilities. The release gate stays report.json.release_decision.decision. Claude Code users should prefer the CLI + hooks surface.

pip install 'agents-shipgate[mcp]'
// .mcp.json
{
  "mcpServers": {
    "agents-shipgate": {
      "command": "agents-shipgate",
      "args": ["mcp-serve"]
    }
  }
}

Tools: shipgate.check (caller-provided diff to shipgate.agent_boundary_result/v3), shipgate.preflight (protected surfaces, required evidence, and policy/trust root hashes), shipgate.explain (check id or fp_... fingerprint), and shipgate.capabilities (capability lock export or diff). The server is read-only: it does not run agents, call tools, write artifacts, connect to external MCP servers, or broker general MCP permissions.

Pre-commit hook (local)

Run Agents Shipgate locally on every commit that touches a tool-surface artifact. Two equivalent setups:

Canonical (let pre-commit manage the install):

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/ThreeMoonsLab/agents-shipgate
    rev: v1.0.0
    hooks:
      - id: agents-shipgate

Local (agents-shipgate already on PATH):

repos:
  - repo: local
    hooks:
      - id: agents-shipgate
        name: Agents Shipgate merge-gate verify
        entry: agents-shipgate verify --config shipgate.yaml --ci-mode advisory --format text
        language: system
        pass_filenames: false
        # pre-commit's default `types: [file]` drops a tracked symlink
        # before `files:` runs; governance paths can be symlinks.
        types: []
        types_or: [file, symlink]
        files: |
          (?ix)^(
            (.*/)?shipgate\.yaml|
            .*tools.*\.json|
            .*mcp.*\.json|
            .*n8n.*\.json|
            (.*/)?\.n8n(/.*)?|
            (.*/)?conductor/.*\.json|
            (.*/)?ai/examples/.*\.json|
            (.*/)?\.codex/(config\.toml|hooks\.json|requirements\.toml)|
            (.*/)?\.claude/(settings(\.local)?\.json|commands(/.*)?|hooks/hooks\.json)|
            (.*/)?\.cursor/(cli\.json|mcp\.json|rules(/.*)?)|
            (.*/)?\.vscode/mcp\.json|
            (.*/)?\.shipgate/agent-contract\.json|
            (.*/)?(AGENTS(\.override)?|CLAUDE)\.md|
            \.(agents|claude)/skills/.*|
            (.*/)?\.codex-plugin(/.*)?|
            (.*/)?\.agents/plugins(/.*)?|
            .*\.app\.json|
            (.*/)?SKILL\.md|
            .*openapi.*\.(yaml|yml|json)|
            .*swagger.*\.(yaml|yml|json)|
            \.agents-shipgate/.*\.json|
            (.*/)?prompts(/.*)?|
            (.*/)?policies(/.*)?|
            (.*/)?\.github/workflows/.*\.(yaml|yml)
          )$

The hook fires when a staged change touches a path-based trigger from docs/triggers.json: shipgate.yaml, MCP/OpenAPI/Swagger exports, **/*tools*.json inventories, n8n and Conductor workflow JSON, Codex repo config and static requirements, Claude settings and commands, Cursor permissions and rules, VS Code MCP, the downstream local contract, agent instructions (AGENTS.md, AGENTS.override.md, CLAUDE.md), Codex plugin package files, prompts/**, policies/**, and GitHub workflows. Matching is case-insensitive, and every clause whose catalog glob is recursive is recursive too — services/foo/policies/refund.yaml, enterprise/lib/captain/Prompts/system.md, and a nested AGENTS.md all stage the same as a repo-root copy, as do the nested protected copies covered by the boundary registry. A dir/** glob also matches a tracked path named exactly dir. Diff-only triggers (TRIGGER-FUNCTION-TOOL-DECORATOR, TRIGGER-GOOGLE-ADK-AGENT-TOOLS-CHANGED, and the diff-leg of TRIGGER-SHIPGATE-CI-WORKFLOW) are not covered by the regex pre-gate — pre-commit's files: regex is purely path-based. TRIGGER-FRAMEWORK-VERSION-BUMP needs a framework package token in the diff in addition to a changed dependency manifest, so the path-only regex cannot decide it either. Once the hook fires, the verify entry runs the full trigger evaluator (including diff rules) and base auto-detection itself. Use the GitHub Action for coverage on commits whose paths don't match the regex at all, or python -m agents_shipgate.triggers --git-diff HEAD for diff-aware local checks. The canonical hook manifest pre-commit reads from the repo root is /.pre-commit-hooks.yaml — it exposes agents-shipgate, agents-shipgate-strict, and agents-shipgate-validate. See examples/pre-commit/ for the longer write-up on advisory vs. strict modes and which hook ID to pick.