Sergeant supervises AI engineering work. A Linear issue delegated to Sergeant becomes merged code or a clear question back on the issue, without a human driving each step: Sergeant's reasoning reads the issue, GitHub, and run state; briefs one primary worker to do the engineering; starts a separate fresh-context reviewer when the change warrants it; asks a human only for genuine judgment; and merges through a deterministic Gate.
This is Sergeant 2, a TypeScript workspace (pnpm, Turborepo, oxlint, Vitest). It is currently a walking skeleton:
| Package | What it is |
|---|---|
packages/contracts (@terros/sergeant-contracts) |
Zod schemas for the conversation, Situation Report, run reports, proposed actions, and PR facts; the ports adapters implement; the pure merge Gate |
packages/reasoning (@terros/sergeant-reasoning) |
One fresh-context reasoning turn through the local claude CLI: Situation Report in, validated proposed actions out |
packages/linear, packages/github (@terros/sergeant-linear, -github) |
Live Linear and GitHub adapters (UNF-704), and GitHub App installation tokens for the control-plane and worker Apps (UNF-720) |
packages/runner (@terros/sergeant-runner) |
Local primary worker and fresh-context reviewer (UNF-705) |
packages/sergeant (@terros/sergeant) |
The app: executes proposed actions through the Gate against the ports, and the polling loop for one explicitly selected, V2-delegated issue that files follow-up issues, can ask a human and wait for the reply, posts one outcome comment after the merge, records review telemetry and audits a sample of skipped reviews, and stays within a task budget (UNF-706, UNF-724, UNF-727, UNF-728, UNF-729, UNF-730); and the long-running service that runs that loop for every delegated issue (UNF-719), with its client API (UNF-713) for Linear-authenticated callers (TECH-4938) |
packages/cli (@terros/sergeant-cli) |
sgt, a thin client of that API (UNF-714), with sgt login through Linear (TECH-4938) |
packages/mcp (@terros/sergeant-mcp) |
sgt-mcp, a read-only MCP server over stdio, another thin client of that API (TECH-4940) |
The architecture is designed in pseudocode in docs/design/ (draft 3).
Code here implements only the parts a ticket asks for. AGENTS.md is the guide for
agents and contributors.
Requires Node.js 24 LTS (.nvmrc) and pnpm (pinned via packageManager;
corepack enable). Run from the repository root:
pnpm install # install dependencies
pnpm exec turbo boundaries # one-way package dependencies
pnpm lint # turbo run lint (oxlint)
pnpm typecheck # turbo run typecheck (tsc, strict)
pnpm test # turbo run test (vitest)CI (.github/workflows/v2.yml) runs the same boundaries, lint, typecheck,
and test tasks on every pull request and every push to main.
The live commands below are manual and never run from tests or CI. They take an installation
config file that holds identifiers and secret references only (AWS Secrets Manager ids, resolved
with the configured AWS profile and region); no credential is ever printed, and no ambient gh,
Linear, or Claude login is used. Its shape is InstallationConfig in
packages/sergeant/src/config.ts:
{
"secrets": { "awsRegion": "us-west-2", "awsProfile": "<profile>" },
"linear": {
"tokenSecret": "<V2 Linear agent token secret id>",
"agentUserId": "<V2 agent user id>",
"reviewerProfiles": { "github-login": "https://linear.app/<workspace>/profiles/<user>" }
},
"github": {
"controlPlaneApp": { "appId": 1, "installationId": 2, "privateKeySecret": "<secret id>" },
"workerApp": { "appId": 3, "installationId": 4, "privateKeySecret": "<secret id>" }
},
"repositories": { "owner/name": { "mergeMethod": "squash", "mergePolicy": "sergeant" } },
"modelTokenSecret": "<Sergeant model token secret id>",
"gitIdentity": { "name": "<human name>", "email": "<human email>" }
}Each repository's mergePolicy says who merges in it (TECH-5244): sergeant approves and merges a
gated head; human never does, and a repository without one is human, so it fails safe. In a
human repository the same merge checks decide when a head is ready, and Sergeant then hands it to a
human instead: it marks the PR ready, requests review from the code owners GitHub asked or else the
issue's assignee (their Linear profile in linear.reviewerProfiles), posts the review summary on the
PR and the issue, and waits. A human's review or comment is feedback as always; their merge finishes
the task as Sergeant's own would. The GitHub adapter rereads the live policy before any approval or
merge call and refuses in a human repository whatever asked it to merge.
The Linear token must act as agentUserId (checked at startup); every Linear read and write uses it.
Optional linear.otherAgentUserIds lists other agents' users (V1's) whose comments are not human
input. Optional linear.delegatingAppIds lists Linear app ids (botActor.id), such as Linear's MCP
connector, whose delegations count as the assignee's own when the app acted for a user with the
assignee's display name; any other app's delegation is refused. Optional linear.reviewerProfiles maps GitHub logins to Linear profile URLs. Sergeant puts
the URL in a re-review request so Linear renders a real user mention and sends an Inbox notification;
a missing mapping or failed lookup leaves the existing plain @github-login text. Optional
review.auditSampleRate (0 to 1, default 0.2) is the fraction of merged heads that
skipped fresh review which get an audit review. Optional budget is the budget window a task gets
when it starts or a human answers one of its questions, for serve and canary alike: "budget": { "minutes": 45, "usd": 10 } (positive
numbers, each optional; unset, 120 minutes and $25; see the task budget under Commands). Optional
maxTasks (a positive integer, default 2) is how many task slots serve fills, and optional
waitingGraceMinutes (default 15) how long a waiting task keeps its slot; serve --max-tasks and --waiting-grace-minutes win over them (see Intake under Commands). Optional
linear.webhookSecret and github.webhookSecret are the signing secrets of the Linear app's and the control-plane App's
webhooks; serve has each webhook endpoint only when its secret is set. Optional humans says who
may use sgt and the client API, each with their own Linear login (see The sgt CLI):
{ "linearClientId": "<the Linear OAuth app's client id>", "teams": ["<team key>"], "approvers": ["<Linear user id>"] }.
Optional release makes the host update itself to green commits of main: { "channel": "main" }
or { "channel": "soaked", "soakMinutes": 90 }, with "paused": true to stop (deploy/README.md,
Automatic updates).
Optional runners chooses each role's agent (TECH-5009): { "reviewer": "codex-local" } runs reviewers
on the Codex CLI; a role not named runs Claude Code, as before. A codex-local role needs codex:
{ "model": "<Codex model>" }, which also lets people register Codex accounts; each run uses its task
owner's registered account. A Codex run's cost is estimated from its tokens at OpenAI's list price for
its model (TECH-5021); optional codex.prices, { "<model>": { "input": 1.25, "cachedInput": 0.125, "output": 10 } } in USD per million tokens, adds or replaces a model's price, and a model with none
counts as unknown cost (packages/runner/README.md).
The control-plane App reads PRs, checks, and branch rules, and approves then merges; it needs
contents and pull requests write, checks and commit statuses read, and metadata read. The worker App
needs contents and pull requests write, and checks and actions read. Neither may hold administration,
workflows, secrets, environments, deployments, or actions write, and the worker App must not be a
ruleset bypass actor. Contents write would let the worker App merge its own green PR, so each
repository's ruleset must require at least one approving review: GitHub never lets a PR's author
approve it, and the control-plane App submits that approval on the exact gated head only after every
Gate check passes, immediately before its SHA-guarded merge. A failed approval stops the merge.
Only the base branch's declared required checks count toward a merge (ruleset
required_status_checks), so a repository with none cannot be merged; a repository's
"observedChecksFallback": true instead treats every check observed on the exact head as required.
After the V2 identities and the repository ruleset are provisioned (UNF-720), the live check confirms them without writing anything or spending model money, and fails if the base branch does not require an approving review:
pnpm --filter @terros/sergeant live-check --config <file> --repo owner/name [--issue UNF-123] [--pr 45]The canary loop reads and writes live Linear and GitHub, launches real model sessions, and costs
money. It works on the issue only while the issue is delegated to agentUserId: before anything
starts, on every poll, and again from a live read before each start, merge, and the outcome comment
(Gate rule A1), and only while it is not in Backlog, Canceled, or Done (A2, TECH-4989; Done after
Sergeant's own closing merge is the normal end). A task either runs or is stopped, and every stop
takes one path: the issue undelegated or reassigned, moved by a human to Backlog, Canceled, or Done,
or sgt task cancel (a reassignment is a handoff that keeps the PRs open and puts the issue back in
Todo, undelegated; TECH-5179). The stop is recorded in <dir>/cancel.json first, and from then on the loop
takes no turn and makes no effect: it keeps retrying each run's cancellation, and treats a run whose
status it cannot read as still running, until the runner confirms every run stopped; then it closes
the task's open PRs (only those the worker App opened) with a short comment, posts one comment on the
issue saying it stopped and which PRs it closed, and sets state.json aside. Nothing resumes a stopped
task: the issue delegated and in Todo again starts a fresh one, with a new budget. After the merge, the V2 agent posts one outcome comment (PR, reviewed
head, observed required checks, merge result, known gaps), keyed by the issue and the merge so a
rerun never posts it twice; Linear's GitHub integration moves the issue to Done. Workers push their branches and open PRs with a worker-App token scoped to their run's
repositories; the merge is one control-plane action: fresh exact-head PR, check, and Linear reads,
the Gate, and GitHub's SHA-guarded merge. A blocking review finding or a failed required check
wakes a turn that may start a successor worker (one at a time, R1) on the same PR; its brief carries
the PRs with their check states and every earlier run's report and findings. The fix is a new head,
so the merge again needs a fresh approving review of it or the worker's waiver for it (M6). No
turn count bounds the iterations: the budget window does (TECH-5059). This runner cannot resume a worker's session, so every
continuation is a successor. Re-running the same command resumes from <dir>/state.json.
Reasoning may ask a human (ask_human): the V2 agent posts one question comment, keyed by the issue
and the conversation revision it was asked from, and the loop then takes no turn and makes no effect
until a human comments or edits the issue. No timeout decides for the human; STOP or undelegation
still ends the loop. The wait is never stored locally: a restart finds the question on the issue.
A follow-up issue is only for a concrete bug, required unfinished work from the task's own scope, a
real blocker, or a current operational or security problem (TECH-5186). There is no per-task quota,
but more than one from a task is exceptional. Workers suggest them in their report's followups, each
with its category and why it meets it; everything else they noticed (what made the task harder or
slower, what could be better, whether it will recur) goes in the report's feedback and its short
Feedback section, and is never filed. Reviewers' non_blocking findings and nits never become
follow-ups, nor do theoretical edge cases, future robustness, generalized cleanup, speculative rollback
hazards, or abstraction improvements. Reasoning decides whether a suggestion deserves an issue and
proposes create_followup with its category, why, and a short key naming the idea; the filed issue
opens with the category and why. The V2 agent files it in the task issue's team and project, related to the issue
(or blocked by it), with no delegate or assignee, so humans triage it. Linear's client-supplied ids,
derived from followup:<task>:<key>, make it at most one issue and one relation per key, even across
a crash or a rerun; filed follow-ups are kept in state.json, shown to every later turn, and listed
in the outcome comment. The category and why are the only limit: no count caps them. The Linear token
needs permission to create issues and issue relations.
With the closing PR's merge, reasoning may give up to three short feedback lines worth keeping, from
the workers' feedback and any non-blocking review notes worth keeping. After the outcome comment,
Sergeant posts them once as a Sergeant feedback comment on the issue, keyed by the merge like the
outcome, and adds the sergeant-feedback label. A merge reasoning did not make (a human's) uses only
the closing worker's explicit feedback. An issue completed in Linear without a recognized closing merge
gets the same from the latest worker, posted once as its stop finishes (keyed by the stop) while the
issue is still delegated to Sergeant. Delivery is the comment and the label: until both succeed, the
feedback is not marked posted and the task is neither seen through nor its stop finished, so a later
pass retries, and the key keeps the comment to one. A task with nothing worth keeping ("Nothing
notable") gets neither. A retro across tasks (TECH-5187) reads these comments; no other store holds
them. The sergeant-feedback label must already exist (a workspace label, or one in the issue's team);
Sergeant never creates it. The Linear token also needs permission to edit issue labels.
A Sergeant retro (TECH-5187, packages/sergeant/src/retro.ts) sees across tasks what a single worker
can't. With the installation config's retro (the Sergeant project's Linear id and the team its issues
are filed in), serve checks hourly and runs one when about 10 tasks got a Sergeant feedback comment since
the last retro, or at most 2 weeks after the last one if anything new happened; sgt retro runs one now
(after a big architecture change, say). Nothing else schedules it, so a healthy system that leaves less
feedback gets rarer retros. It runs on the control plane with Sergeant's own reasoning model, never a
worker or a person's model account. It reads two inputs only: the Sergeant feedback comments posted since
the last retro, and what became of the issues Sergeant filed in that time (done, canceled, or still
waiting), plus the previous retro and its issues as they stand now. Reasoning first says in one paragraph
whether the last retro's recommendations happened and their themes stopped recurring, then gives themes
with their evidence (task ids) and a recommendation each, and rarely an issue: only for repeated evidence
across tasks, a meaningful recurring cost or risk, a clear systemic defect, or a strong simplification,
preferring removing complexity, then guidance, then docs or tooling, then new machinery (at most 3).
Sergeant files those in Backlog in the Sergeant project, unassigned and not delegated, for a human to
promote, and posts the retro as one Linear document there titled Sergeant retro <date>. Linear is the
store: the newest such document is when the last window ended and what it filed. <dir>/retro.json only
keeps an answer whose filing or posting failed, so the retry pays for no second answer; issue and
document ids derive from the window, so a retry files and posts each once. A failed retro is retried a
day later, or at the next sgt retro. The Linear token needs
permission to create documents.
Review quality is telemetry, never a gate (UNF-730, design 06 §8–9). Every reviewer run that
finishes, whether or not the task ever merges, is written as a line of <dir>/reviews.jsonl: trigger
(required or audit), mode (a separate fresh run), reviewer and implementer provider and model,
whether their vendors are the same, the heads reviewed, the verdict, finding counts, the must-fix
(blocking) findings verbatim, the merged head once there is one, and the resulting change
(resultingMutation: true when a worker's report lists one of its findings fixed in
addressedFindings, false when it had none or workers answered them and fixed none, otherwise
"unknown"; a head changing after a review is not counted). A review is written again only when one
of those later facts changes, so the last line per run id holds. A merged head that no fresh review
approved skipped review; a stable hash of the head picks review.auditSampleRate of those for an
audit review, a separate fresh reviewer of exactly the merged head, started only after the merge
so it can never hold one up. After observing Done the loop waits for any review still running,
the audit included, and records it; an audit's must-fix findings on merged code also go to
<dir>/audit-followups.jsonl and an AUDIT FOLLOW-UP log line for a human to act on (nothing is
reopened or reverted, and no turn runs after the merge to propose a create_followup). To compare
review modes across runs:
cat <state dirs>/reviews.jsonl | jq -s 'reduce .[] as $f ({}; .[$f.runId] = $f) | [.[]] | group_by([.trigger, .vendor])
| map({trigger: .[0].trigger, vendor: .[0].vendor, reviews: length, withMustFix: map(select(.findings.blocking > 0)) | length,
ledToChange: map(select(.resultingMutation == true)) | length, unknown: map(select(.resultingMutation == "unknown")) | length})'Each task has a budget window, saved in state.json with the task's start before anything else
happens. It is the installation config's budget ({ "minutes": 45, "usd": 10 }, say; TECH-4964),
each field unset defaulting to 120 minutes and $25; canary's --budget-minutes and --budget-usd
override the config's for its task. A restart keeps the stored window and logs that it ignores a
different one, so changing the config's budget affects only tasks that start afterward, and their
next fresh window. A human's answer to any of the V2 agent's questions gives the task a fresh window
(TECH-5059): from the answer's time, with zero spend (runs of earlier windows no longer count) and the
config's current budget. Wall time is hard and runs from the window's start, including time spent
waiting for a human before an answer, and for a task slot after one. Spend is best-effort: the cost runs and reasoning turns report when they end (a turn's cost counts before its proposals run), so a running or
canceled run's cost is unknown and the wall time is the backstop; there is no billing ledger. Once
either is exhausted, no run, message, follow-up, or merge happens (Gate rule B1, checked before every
effect and again right after its live reads), running runs are canceled until the runner confirms it,
and the V2 agent asks one Question for you, summarizing spend, runs, and PRs, with the options to
extend or accept as-is. It is posted like any question, under a key of the task and the window, so a
restart finds it on Linear instead of asking again, and nothing happens until a human replies after it.
A reply opens a fresh window like any answer, so "extend" needs nothing more: the next turn carries on
with the work, and a steer is carried out in it. On "accept as-is" (the option's number, or the same in
the human's own words) reasoning proposes accept_as_is instead (TECH-5118): the loop ends accepted
with nothing more asked, and state.json is set aside so intake does not resume it, leaving its PRs
and the issue for a human. One line on the issue says so (TECH-5120), posted under a key of the
accepting reply, so a retried turn posts no second one. Gate rule Q2 refuses it unless a human replied to the budget question, and
unless the live conversation is still the one reasoning read, so a reply posted while it reasoned is
read by the next turn instead. An accepted.json marker keeps intake from starting the issue afresh
while it stays delegated and in Todo; moving the issue out of Todo (or undelegating it, or
sgt task wake) clears it, so back in Todo it starts a fresh task, as after a stop. When a
worker's verification shows nothing to change and the task opened no PR, reasoning proposes
close_issue instead of asking (TECH-5232): the same ending posts its evidence in one comment and
moves the issue to Done (main already covers it) or Canceled (obsolete), and the loop ends accepted;
Gate rules C1–C4 refuse it with any PR in the task, without evidence, before a worker finished, or once
the conversation changed. A task
that runs away in its fresh window is stopped at that window's end and asked once more. The outcome comment after a merge is the one
effect B1 does not hold back: it reports a merge that already happened, and withholding it would hide
the merge from the human. A run's id is saved before the runner starts it, so a crash in between still
leaves a run the loop cancels; one the runner never started is dropped once a cancel confirms it.
docker build -t sergeant-runner:local packages/runner/container
pnpm --filter @terros/sergeant canary --config <file> --issue UNF-123 --repo owner/name --dir <state dir>serve runs Sergeant unattended (UNF-719): one process that works every open issue delegated to the
V2 agent, in every repository the installation config enrolls, with nobody starting an issue by hand.
It is a thin shell over the canary's per-task loop, not a workflow engine:
- Intake lists open (not completed or canceled) issues delegated to
agentUserIdevery--intake-seconds(120) and runs each one's loop in one ofmaxTasks(2) task slots (TECH-5008). A free slot goes to the highest-ordered task that wants one: by Linear status, In Review, then In Progress, then Todo; then by priority, Urgent to none; then newest first. Finishing work beats starting it. A task asked to wake (sgt task wake) goes first. A task holds its slot while it runs a worker, a reviewer, or a reasoning turn, and while it waits on anything (a question, a budget reply, CI, mergeability, a human merge, an unreadable runner) for up towaitingGraceMinutes(15). Answered within the grace, it continues at once; past it, its slot goes to the next task quietly, with nothing posted in Linear, and its loop keeps polling without a slot (TECH-5015). Once it has work again, it queues for a slot in the same order as new work, so In Progress goes ahead of Todo.GET /statuslists the released tasks underreleased. Linear's list only discovers new work, and a task starts only from Todo (TECH-4989): an issue in Triage or Backlog waits until a human moves it there. A Todo issue with a Linear "blocked by" issue that is neither completed nor canceled waits too (TECH-5066), logging<issue> waiting on blocker <issue>, and starts at the first intake after its last blocker finishes; one with more than 20 inverse relations waits as if blocked, since Sergeant reads only 20. A task already under way is not held back. Starting it moves the issue from Todo to In Progress, the only state Sergeant moves an issue to. Every intake also resumes each local task (state.json) with no loop, whether or not Linear lists it, into a free slot in the same order, or with no slot until it has work to do. Its loop's own live checks then continue it, stop it (undelegated, or in Backlog, Canceled, or Done without its closing PR merged), or see it through after the merge; a stop never needs a task slot, and a task seen through is not resumed again. - Each task loop is the canary's: every
--poll-seconds(60) it re-reads its runs, the PRs Linear links to the issue or a worker reported, with their checks, and the Linear conversation, and takes a reasoning turn only when they changed, so a missed webhook costs only latency. Its state is<state dir>/tasks/<issue>/; runs live under<state dir>/runs/. - Webhooks (TECH-4937) are a latency optimization, never the source of truth.
POST /webhooks/linearandPOST /webhooks/githubrefuse a delivery whose signature (HMAC-SHA256 of the body,Linear-SignatureorX-Hub-Signature-256) does not verify, before parsing it, and a Linear delivery whose signedwebhookTimestampis over a minute from now. A verified event about an issue's delegation, state, title, description, labels, comments, attachments, or relations, or a PR's changes, pushes, reviews, check runs and suites, or statuses, names the issue or PR it is about; each task loop watching that issue, PR, or head ends its wait and rereads, and a delegated issue with no loop, or a delegation change to or from the V2 agent, runs an intake now. Nothing else happens: no event owes a turn or is recorded, repeated events coalesce (each loop, and intake, wakes at most once per 5 seconds), an issue or PR no task watches is ignored, and the polls still find every change. - Post-merge feedback (TECH-4985). While a task is active, a human's comment is part of its
conversation, as above. Once its completing PR (
Fixes) merged or its issue is Done, the loop takes no more turns, so a sweep every 10 minutes reads those issues (delegated toagentUserIdand completed in the last 14 days, open and delegated, or merged by this host in that time): human comments on the issue after the work landed, and comments, review comments, and reviews on its merged worker-App PRs after their merge from the repository's owners, members, and collaborators. A reasoning call judges each one; an acknowledgement or discussion files nothing. Actionable feedback becomes one ordinary follow-up issue (Backlog, assigned to the origin's owner, not delegated; see the design's §11), related to the origin, whose description is the delta reasoning wrote, the feedback verbatim, and links to the original issue and the merged PRs; a comment on the origin says so. A human starts it like any issue, by moving it to Todo and delegating it to Sergeant. Its Linear id is derived from the feedback, so the same feedback never files a second issue.<state dir>/feedback.jsonrecords what was judged and failed attempts, and itssince(the first sweep) keeps feedback from before the rollout out. Per issue, at most 3 follow-ups are filed and 10 pieces of feedback judged; past that, or after 3 failed attempts at one piece, Sergeant says so in a comment on the issue rather than dropping it silently. - A loop that ends (idle, a failed read) is admitted again on a later intake while the issue is still delegated: an unchanged task takes no turn, a changed one does. A failed intake is logged and retried next interval.
- SIGINT or SIGTERM stops intake and ends each loop at its next poll, never mid-turn; a second signal
exits at once. A restart rereads Linear, GitHub, the runner, and each task's
state.json, and continues, accepting some repeated work. One process serves a state directory: it holds an OS file lock on<state dir>/service.lock(released when it exits, however it exits), and a second start is refused while it does.GET /healthon--host(127.0.0.1) and--port(8080) reports the process alive, the tasks running, and the last intake. - The same port answers the client API under
/v1(UNF-713, the slice of design 11 §2sgtuses): task and run reads,POST /v1/tasks/:ref/wake,/v1/tasks/:ref/cancel, and/v1/runs/:id/cancel. A wake ends the task loop's wait and owes it one turn, behind every hold the loop keeps. A task cancel removes the V2 agent's delegation and then takes the task's one stop (above); it answers with any run the runner has not yet confirmed, andservekeeps driving the recorded stop at each intake, across a restart, until it has. A run cancel is the runner's confirmed cancel, noted on the issue so the next turn does not just restart it. Before anything else, every/v1call must name a client version (Sergeant-Cli-Version) no older than the oldestsgtthisservesupports, or it is refused with400and "Runsgt update" (TECH-5188), so a too-oldsgtchanges nothing. Every/v1call names its caller and fails closed without one (TECH-4938): a bearer Linear access token fromsgt login, whichservereads back from Linear on every call and admits only for an active user of the agent's own Linear workspace, not an agent, in one ofhumans.teams; those listed inhumans.approversare approvers too./v1/whoamisays whether the caller is an approver; only an approver (or the loopback operator) mayPOST /v1/accounts/remove-personto remove everything one person registered (TECH-5130), restart or update the Sergeant host through/v1/admin(TECH-5195, deploy/README.md), orPOST /v1/repositories/addor/removeto change the enrolled repositories, which every caller mayGET /v1/repositories(TECH-5193). A change is refused unless both GitHub Apps reach the repository; it rewrites onlyrepositoriesin the installation-config parameter (--config-parameter, elseSERGEANT_CONFIG_PARAMETER), naming the caller in the parameter version's description, andservetakes it in place. With a parameter,servereads its enrolled repositories from it at startup rather than from--config, so a restart keeps every change. Wakes and cancels are logged with the caller's name, and a cancel's note on the issue names them. OnlyGET /v1/auth/config, the public client idsgt loginstarts with, needs no caller. For development on one machine,--trust-loopbackalso admits a caller on the host with no login, as an operator; it is refused unless--hostis127.0.0.1or::1(not a name such aslocalhost), and never covers a request relayed by a proxy or naming a non-loopbackHost. Posts must be JSON, so a cross-site form cannot post.
pnpm --filter @terros/sergeant serve --config <file> --state-dir <dir> [--port 8080] [--max-tasks 2] [--waiting-grace-minutes 15] [--config-parameter <SSM name>] [--trust-loopback]To run serve on one AWS host behind an HTTPS endpoint, see deploy/: Terraform,
the host install and update scripts, and the runbook.
New to sgt? docs/sgt.md is the user guide: install, log in, everyday commands.
This section is the reference.
sgt (UNF-714) only calls the client API and prints the answer: concise lines by default, the API's
own JSON with --json (errors too, as {"error":{"code","message"}}; exit 1 for an API error, 2 for
usage). It needs no AWS credentials. Run sgt --help for the commands: login, logout, whoami,
task list | show | wake | cancel, run list | show | report | cancel, account …, repo list, admin repo add | remove, update. Keep it current: the
installation refuses every request from an sgt older than it supports, before acting on it, and says
to run sgt update, which needs no API.
alias sgt="node $PWD/packages/cli/src/sgt.ts" # from the repository root
sgt login # once: sign in with Linear in the browser
sgt whoami
sgt task show UNF-123 # serve on this machine: http://127.0.0.1:8080The API is --api <url>, else SGT_API_URL, else http://127.0.0.1:8080. sgt login signs you in
as yourself with Linear OAuth (PKCE, no client secret) through the installation's Linear OAuth app,
whose client id the API serves, and keeps the resulting Linear token, renewed before it expires, for
that API URL in ~/.config/sergeant/credentials.json (or under XDG_CONFIG_HOME), readable only by
you. It is the only credential sgt holds, and it is sent only over HTTPS or to loopback. The
browser returns to http://localhost:4546/callback (SGT_LOGIN_PORT changes the port, and the app
must list the result). sgt logout forgets it on this machine; revoking the app in your Linear
account settings ends it at Linear.
The installation's Linear OAuth app (the V2 agent's own app; one app per installation is enough) needs
http://localhost:4546/callback among its callback URLs; sgt login asks for the read scope as the
user (actor=user), and its client id goes in the installation config's humans.linearClientId.
The hosted API is available through the installation's permanent HTTPS endpoint. From a laptop—no
AWS account or session is needed—set that endpoint, sign in, and list tasks:
export SGT_API_URL=https://<hostname>
sgt login && sgt task listThe host operator must first set all three humans fields: the app's public linearClientId, the
allowed Linear team keys in teams, and the allowed teams' Linear user ids in approvers; the app
must have the callback above. Until then the API fails closed. See deploy/README.md
for rollout and unauthenticated live checks.
sgt-mcp (TECH-4940) gives MCP clients the read-only half of sgt over stdio: task_list,
task_show(ref), run_list(task?), run_show(run), run_report(run), and health. Each tool is
one GET to the same client API, and returns the API's JSON unchanged as structured content (errors
as {"error":{"code","message"}} tool errors). It cannot wake or cancel anything. It reaches the API
at the same URL as sgt (--api, else SGT_API_URL, else http://127.0.0.1:8080), and sends the
Linear login sgt login saved for that URL, read and renewed on every call as sgt does (TECH-5123),
so it works from a laptop against the hosted API. With no login saved it sends none, which only a
serve --trust-loopback on the same machine answers. Register it with a client as a stdio server
(docs/sgt.md §9):
{ "mcpServers": { "sergeant": { "command": "node", "args": ["<repo>/packages/mcp/src/sgt-mcp.ts", "--api", "https://<installation hostname>"] } } }Once the V2 Linear agent app, both GitHub Apps, and the canary repository's ruleset exist:
- Put the V2 agent's Linear user id in
linear.agentUserIdand V1's agent user inlinear.otherAgentUserIds. - Create a small controlled issue in the canary repository's team and delegate it to the V2 agent.
live-check --issue <it>must pass, includingissue is delegated to the V2 agent. - Undelegated or delegated to V1's agent,
canary --issue <it>stops at once withCANARY RESULT {"outcome":"stopped",...}and nothing is started; V1 never sees an issue delegated to V2. - Delegated to the V2 agent, run
canarythrough to the merge. Expect exactly one comment, authored by the V2 agent, with the PR link, reviewed head, required checks, merge SHA, and known gaps, then the issue moving to Done through the GitHub integration (CANARY RESULT {"outcome":"done",...}). Running the same command again posts nothing more. - On a second controlled issue, reassign it to V1's agent (or remove the delegate) while a turn is deciding or a worker is running: the loop stops, cancels the running run, merges nothing, and posts nothing.
With the identities from UNF-720, on a controlled issue in the canary repository whose objective invites a fixable mistake (or after a human pushes a deliberately failing commit to the worker's PR):
- Run
canaryuntil a reviewer returnschanges_requestedwith a blocking finding, or a required check fails on the PR head. The next turn starts one successor worker; its<dir>/runs/<runId>/workspace/sergeant-brief.mdlists the PR, the failing check, and the finding. - The successor pushes to the same PR (no second PR, no second worker running), and the loop merges
only after a fresh approving review of the new head, or the successor's own waiver for that exact
head, with required checks green (
CANARY RESULT {"outcome":"done",...}).
On a controlled issue delegated to the V2 agent whose description leaves a real product choice open (for example "retain or purge X; the captain decides"):
- Run
canary. Expect one comment from the V2 agent headed Question for you, then onlywaiting:log lines: no run started, no merge, and noidlestop. - Stop the loop (
touch <dir>/STOP), removeSTOP, and run it again: still exactly one question comment, and it keeps waiting. - Reply on the issue in your own words. The next poll takes a turn that interprets the reply and continues the work (or asks one short clarifying question, which waits the same way).
With the identities from UNF-720, on a controlled issue in the canary repository whose change invites a minor, out-of-scope observation:
- Run
canaryuntil a reviewer approves with anon_blockingfinding, or the worker's report lists afollowupsentry. In the turn that merges (or earlier),turns.jsonlshows acreate_followup <key>: done {"identifier":...}outcome. - In Linear, exactly one new issue exists for it: in the origin issue's team and project, with no
delegate or assignee, a description that stands alone and links back, and a
related(or blocked-by) relation to the origin. The outcome comment lists it under "Follow-ups filed". - Rerun the same command, and once more with
state.jsonrestored from before that turn: no second issue or relation appears.
With the identities from UNF-720 and "review": { "auditSampleRate": 1 } in the installation config,
on a controlled issue in the canary repository whose change is small enough that the worker skips
review (a one-line docs fix, say):
- Run
canarythrough the merge. The worker's report saysreview.required: falsewith a reason, and the merge relies on it (not_required). - After the outcome comment, the log shows
audit review run_audit-<head> started, and the loop waits for it after the issue reaches Done. Its brief (<dir>/runs/run_audit-<head>/workspace/sergeant-brief.md) names the merged head and the skip reason. <dir>/reviews.jsonlends with one"trigger":"audit"line for it; any blocking finding is also in<dir>/audit-followups.jsonland anAUDIT FOLLOW-UPlog line.CANARY RESULTstays"done".- With
auditSampleRate0, or on an issue whose head a fresh review approved, no audit starts, and the lastreviews.jsonlline of each reviewer run is"trigger":"required"withmergedset. - On an issue whose first review requests changes, the successor's report lists the finding in
addressedFindings, and that review's last line says"resultingMutation":true.
On a controlled issue delegated to the V2 agent whose work takes more than a few minutes:
- Run
canary --budget-minutes 5. Once a worker is running and five minutes have passed, expectbudget exhausted (wall time exhausted at ...)andcanceled run_...log lines,docker psshowing nosergeant-run_*container, and one Question for you comment from the V2 agent with the spend, the runs (the workercanceled), any PR, and the extend / accept-as-is options. No run starts and nothing merges while it waits. - Reply "extend". The loop logs
a human answered (...): a fresh budget window ..., and the next turn resumes the work in it. Replying "accept as-is" instead ends the loopaccepted, with no further question, the PR left open, the issue left in its state, and one comment saying Sergeant has stopped and the PR and issue are the human's to merge or close. - With
--budget-usd 1, the first finished run's reported cost exhausts the spend instead (spent $... of $1.00), with the same question. - On another issue, undelegate it while a worker runs: the loop logs
canceled run_...and stops only after that; stopping Docker first (or makingdocker stopfail) keeps it retrying, not stopped. - Kill the canary (Ctrl-C) while the budget question is unanswered and rerun it with
--budget-minutes 120: it logsignoring the budget options, posts no second question, and still waits for the reply.
Sergeant 1, the Rust implementation, has been removed. Its final source, including its runbooks and
architectural decision records, is the annotated tag v1-final (git switch --detach v1-final), and
its last installable release is v0.1.0+aad6046.
Sergeant is open source, licensed under the Apache License 2.0.
Copyright 2026 Terros Inc.
For contributions, see CONTRIBUTING.md. To report a security vulnerability, see SECURITY.md.
