Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 17 additions & 1 deletion .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -7,11 +7,27 @@ on:
workflow_dispatch:
jobs:
test:
runs-on: ubuntu-latest
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
# Windows is not incidental coverage. Claude Code reports `file_path`
# with the host separator, so a Linux-only matrix once let a policy
# ship that matched nothing on Windows — every secret-read deny was
# silently allowed there. Path handling is platform-specific enough
# that it has to be exercised on both.
os: [ubuntu-latest, windows-latest]
python-version: ["3.10", "3.11", "3.12", "3.13"]
exclude:
# Windows runs the full version sweep on one interpreter; the
# platform-specific code paths do not vary by Python version, and
# eight Windows runners per PR buys nothing.
- os: windows-latest
python-version: "3.10"
- os: windows-latest
python-version: "3.11"
- os: windows-latest
python-version: "3.12"
steps:
- uses: actions/checkout@v5
- name: Setup uv
Expand Down
97 changes: 97 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,85 @@ This project follows [Semantic Versioning](https://semver.org/).

## [Unreleased]

### Fixed

- **`janus init` could write a deployment that enforced nothing, and report seven PASS
lines over it.** Claude Code runs hooks through its own shell, with the `PATH` of the
session the operator starts later. `janus init` is normally run as `uv run janus init`,
which puts the project venv's `bin` on `PATH`, so `shutil.which("janus-hook")` succeeded
and the wizard wrote a bare `janus-hook` command. A `claude` session started from an
ordinary shell could not find it: the hook exited 127, and Claude Code treats a
non-zero-but-not-2 hook exit as a **non-blocking** error, so the tool ran and nothing was
enforced — confirmed live, a `Read` of `.env` that the policy denies succeeded with no
error shown. Now: the private scopes (`user`, `project-local`) get the resolved absolute
path, since those settings files are machine-specific anyway; the shared `project` scope
keeps the portable bare name — an absolute venv path would be wrong for teammates — and
verification warns loudly when it will not resolve in a plain shell. A `_hook_is_reachable`
warning for exactly this failure already existed and never fired, because it inspected the
wizard's `PATH` rather than the one that matters.
- **`janus init` verified the policy instead of the deployment.** The closing checks called
`handle_cli_payload` in process, which answers "would this policy deny this payload" —
never "does the command just written to settings.json deny it". That is why the bug above
was invisible. `verify()` now executes the exact command string with `CLAUDE_PROJECT_DIR`
set as the CLI sets it, feeding payloads on stdin, and fails on a non-zero exit,
unparseable stdout (which also makes the shim's stdout-isolation property a standing
check), or a wrong decision. One additional probe re-runs with this interpreter's venv
stripped from `PATH`, standing in for the shell a real session gets.
- **`Bash` could read the secrets `Read` was denied.** `.env` and `*.pem` were in
`SECRET_READ_PATTERN` but missing from `BASH_EXFIL_PATTERN`, so `Read` on `.env` was
denied while `cat .env` was allowed — the same secret, one tool apart. Found by a live
agent, which reached for `Bash` the moment `Read` was refused and returned the contents,
against a policy whose own verification had just printed `reading a .env file is denied`.
Both are now in the `Bash` deny in command-line form (`\.pem\b`, not the end-anchored
`\.pem$` a file path uses), the `.env.example` exemption is preserved, and probes now
cover the `Bash` route to each secret. Probe labels name the tool they tested
(`.env is denied (Read)` / `(Bash)`), because the old wording read as coverage the
deployment did not have. `docs/claude-code-deployment.md` now states plainly what shell
argument-matching can and cannot promise.

- **Subagents were broken under `mode="policy"`, and subagent output silently stopped
tainting.** Two defects, found by re-running the payload-shape capture against CLI
2.1.278 (the fixtures were pinned at 2.1.233). First: `SubagentHandback` — the tool a
subagent uses to deliver its report to its caller — is a CLI-internal transport tool
that strict default-deny blocked, stranding every subagent's work. It is now in
`DEFAULT_CLI_PASSTHROUGH_TOOLS` alongside `ToolSearch`. Second, and worse: **the taint
source for subagent output moved.** On 2.1.233 the subagent's report came back in the
parent's `PostToolUse[Agent].tool_response.content`; on 2.1.278 that field is a
placeholder pointing at the `SubagentHandback` call (new `handback: "send"` key), and
the content travels in that call's `tool_input.message` — its *input*, while its
response is only a delivery receipt. A `TaintTracker` sourcing `Agent` therefore
recorded a fixed placeholder sentence and lost the subagent's content entirely, with
no error and no failing test, leaving every downstream sink open. The new
`CLI_INPUT_SOURCE_TOOLS` mapping makes the recording seam read the input for such
tools, gated on `PostToolUse` so a pre-decision or failed handback records nothing.
**Deployments that want subagent output to taint must now list `SubagentHandback` as a
source — naming `Agent` alone no longer reaches that content.** Three fixtures captured
from the live 2.1.278 session pin all of it.

- **Path policies did not match on Windows — every secret-read deny was silently allowed
there.** Claude Code reports `tool_input.file_path` with the *host's* separator
(`C:\Users\...\.env`, verified against a live CLI 2.1.246 session), while the starter
policy anchored on `/`. Against the previous starter, reads of `.env`, `~/.ssh/id_rsa`,
`~/.aws/credentials` and `~/.claude/.credentials.json` were all **allowed** on Windows, as
were writes to `.claude/settings.json` — the rule meant to stop an agent disarming the
guard. Only `\.pem$` held, being the one pattern needing no separator. Path patterns now
use a separator class (`janus.cli.starter_policy.SEP`, `[/\\]`), user-supplied paths are
normalized the same way, and `examples/claude_code/policy.starter.json` is regenerated to
match. A Windows payload fixture captured from a live session
(`tests/fixtures/claude_code_payloads/pretooluse.windows-read.json`) pins it. The bug was
invisible because every prior fixture, and the whole CI matrix, was Linux.
- **`janus init` verification reported PASS against paths the CLI never sends.** Its probes
built paths with `as_posix()`, so on Windows they exercised forward slashes while the
deployment received backslashes — seven green checks on a policy that was allowing `.env`
reads. Probes now use the host's native separator.
- **The `janus-hook` deadline was inert on Windows.** `_deadline` needs `SIGALRM`, so on
Windows it degraded to no deadline at all, and a wedged decision ran until the CLI's own
hook timeout — which fails **open**. A worker-thread fallback restores the property: the
shim reaches its own limit and emits a deny while it still can. This also fixes the one
test that had been failing on Windows.
- CI now runs the suite on `windows-latest` as well as `ubuntu-latest`. All three bugs above
were platform-specific and a Linux-only matrix could not see any of them.

### Changed

- **BREAKING — `openai` and `jinja2` moved out of core dependencies** into the new `generate`
Expand Down Expand Up @@ -35,6 +114,24 @@ This project follows [Semantic Versioning](https://semver.org/).

### Added

- **`janus init` — an onboarding wizard, behind a new `janus` console script.** Setting Janus
up on the Claude Code CLI previously meant hand-writing a policy, pasting a hooks block into
a settings file, and merging the backstop by hand; a guard nobody finishes installing
protects nothing. `janus init` asks a handful of questions (scope, what to protect, network
posture, git posture, MCP servers, strictness), shows the exact settings diff, and on
confirmation writes the policy, the `PreToolUse` entry — with the explicit `timeout` the docs
always asked for and no example ever showed — and the `permissions.deny` backstop. It then
verifies by feeding synthetic payloads through `handle_cli_payload` with the flags it just
wrote, so a `PASS` reflects the deployed decision path rather than the wizard's intent.
Re-running updates the existing hook in place; foreign hooks, foreign deny entries, and
unrelated settings are never touched, and the previous file is backed up. `--dry-run`,
`--yes` (CI; a non-TTY without it is refused rather than defaulted), `--scope`, `--force`.
Optional: with the `generate` extra and an API key, it can draft argument-level rules for
review — accepting *replaces* a tool's blanket allow, since generated priority-100 rules
would otherwise sit unreachable behind it.
The `janus` script is deliberately separate from `janus-hook`, which stays a pure
decision process with no interactive surface. `janus doctor` delegates to the same
`janus.cli.hook.run_doctor` (renamed from `_doctor`) that `janus-hook doctor` uses.
- **Claude Code CLI adapter** (`janus.adapters.claude_code` + the `janus-hook` console script,
core install — no extra): enforce a Janus policy on the *interactive* `claude` CLI via its
`PreToolUse`/`PostToolUse` hooks. Unlike the SDK path, Janus does not construct the session
Expand Down
25 changes: 22 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -732,19 +732,33 @@ guarded = guard_tool_body("fetch_page", my_async_body, TOOL_POLICY,

The security model is genuinely weaker than the SDK path's, and the docs say so up front: on the CLI, **Janus is a policy monitor over a session it does not own, backstopped by `permissions.deny` — not a reachability lockdown.** The human constructs the session, so the SDK path's `tools=[]`/`strict_mcp_config`/`allowed_tools` layers are simply gone.

Wire the `janus-hook` shim into a settings file:
The fastest way in is `janus init` — it asks a few questions (what to protect, network and
git posture, how strict), writes the policy, the hook wiring, and the `permissions.deny`
backstop, shows you the settings diff before touching anything, and then verifies the
result through the real decision path:

```bash
pip install janus-guard
janus init # --dry-run to preview, --yes for CI
```

Or wire the `janus-hook` shim into a settings file yourself:

```json
{
"hooks": {
"PreToolUse": [
{ "hooks": [{ "type": "command",
"command": "janus-hook pre --policy /etc/janus/policy.json --mode gate" }] }
"command": "janus-hook pre --policy /etc/janus/policy.json --mode gate",
"timeout": 10 }] }
]
}
}
```

Keep the hook's `timeout` above the shim's `--deadline` (default 5s): the CLI's own hook
timeout fails **open**, so the shim must reach its deadline first and deny while it can.

`--mode gate` (default) enforces the tools the policy has an opinion about and abstains to the CLI's own permission flow elsewhere; `--mode policy` is strict default-deny. Gate mode auto-promotes to policy mode under `bypassPermissions`, where abstention would be a silent allow — so bypass sessions need the policy to enumerate their tool surface. The shim fails **closed** (unreadable policy, internal error, or its own `--deadline` all deny), which matters because the CLI's hook dispatch fails **open** on timeout. `janus-hook doctor` self-tests the install; `janus-hook backstop` prints the `permissions.deny` block that holds even if hooks stop running.

Phase 1 is deliberately stateless — static policy per call, no taint or cross-call state (the phase-2 daemon restores those). See the [adapters guide](https://agentic-ai-risk-mitigation.github.io/Janus/adapters/) for the full security model, gate/policy semantics, and the verified `ask`/`escalate` probe results.
Expand Down Expand Up @@ -883,7 +897,12 @@ janus/
│ └── claude_code.py # Claude Code CLI hook adapter (interactive `claude`)
│
└── cli/
└── hook.py # `janus-hook` — the CLI hook shim (fails closed)
├── hook.py # `janus-hook` — the CLI hook shim (fails closed)
├── main.py # `janus` — operator commands (init, doctor)
├── init.py # the `janus init` onboarding wizard
├── starter_policy.py # starter-policy builder + Claude Code tool table
├── claude_settings.py # settings.json read/merge/backup/write
└── _console.py # stdlib prompts (no TUI dependency)

examples/ # Demo scenario framework + FastAPI web app + docker-compose.yml for SpiceDB
tests/ # Offline regression suite (+ tests/smoke/, opt-in live SDK checks)
Expand Down
33 changes: 32 additions & 1 deletion docs/adapters.md
Original file line number Diff line number Diff line change
Expand Up @@ -276,6 +276,11 @@ doc.

### Wiring it (phase 1: settings file, stateless)

`janus init` does all of the below interactively — policy, hook entry, and backstop — and
verifies the result through this same decision path; see
[Claude Code Deployment → Wizard setup](claude-code-deployment.md#wizard-setup-janus-init).
By hand:

```bash
janus-hook backstop > /tmp/backstop.json # the permissions.deny block; merge into settings
```
Expand All @@ -285,7 +290,8 @@ janus-hook backstop > /tmp/backstop.json # the permissions.deny block; merge i
"hooks": {
"PreToolUse": [
{ "hooks": [{ "type": "command",
"command": "janus-hook pre --policy /etc/janus/policy.json --mode gate" }] }
"command": "janus-hook pre --policy /etc/janus/policy.json --mode gate",
"timeout": 10 }] }
]
}
}
Expand Down Expand Up @@ -326,6 +332,31 @@ one of these behaviours is pinned by verbatim payload captures in
`tests/fixtures/claude_code_payloads/` — where the fixtures and the docs disagree, the fixtures
win.

### Subagents: `SubagentHandback`

Two things to know if your agents spawn subagents.

**It is never policy-gated.** `SubagentHandback` is how a subagent delivers its final report
to its caller. It reaches no resource, so denying it accomplishes nothing except stranding the
subagent's work — `mode="policy"` used to do exactly that. It now sits in
`DEFAULT_CLI_PASSTHROUGH_TOOLS` next to `ToolSearch`.

**But unlike `ToolSearch`, it carries content — and it is where subagent output enters the
parent turn.** This changed under us. On CLI 2.1.233 the subagent's report came back in the
parent's `PostToolUse[Agent]` result; on 2.1.278 that field is a placeholder pointing at the
handback call, and the report travels in `SubagentHandback`'s **`tool_input.message`** (its
response is only a delivery receipt). So:

```python
TaintTracker(sources={"SubagentHandback": "subagent", ...})
```

**Listing `Agent` alone no longer reaches that content** — it records a fixed placeholder
sentence, with no error and nothing failing, while every downstream sink stays open. The
recording seam reads the input for tools named in `CLI_INPUT_SOURCE_TOOLS`, gated on
`PostToolUse` so a handback that was denied or failed records nothing. Both CLI versions'
shapes are pinned in `tests/fixtures/claude_code_payloads/`.

`claude_code_resolve_name(name, known_servers=...)` maps `mcp__<server>__<tool>` (and the
plugin form `mcp__plugin_<plugin>_<server>__<tool>`) to the bare policy key. Supply
`known_servers`: the CLI has no `strict_mcp_config`, so an unsanctioned server would otherwise
Expand Down
Loading
Loading