diff --git a/AGENTS.md b/AGENTS.md index efb8e0e..7a57446 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -10,15 +10,20 @@ | Doing | Command / skill | |---|---| | Full audit | `/axguard-audit` or skill `axguard-audit` | +| MCP-first review / investigate / verify fix | skill `axguard-security` → [docs/mcp.md](docs/mcp.md) | | Security lead pass | skill `axguard-cso` | | Triage | `/axguard-triage` | | Fix | `/axguard-fix` / skill `axguard-remediate` | | Report | `/axguard-report` | +`axguard-security` teaches when to call AXGuard and which MCP tool to use (`axguard_security_review`, `axguard_investigate`, `axguard_verify_fix`). Prefer MCP when available; CLI as fallback. + CLI: ```bash pip install -e . +pip install -e '.[mcp]' # agent MCP interface axguard help axguard audit . +axguard mcp doctor ``` diff --git a/CLAUDE.md b/CLAUDE.md index fc644a7..20f8e42 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -19,6 +19,7 @@ pip install -e . | What you are doing | Start here | |---|---| | About to publish | `/axguard-audit` | +| Security Diff on a change | `/axguard-diff` · `axguard diff` | | Quick check while coding | `/axguard-scan` | | New / unknown codebase | `/axguard-threat-model` → `/axguard-audit` | | Secrets | `/axguard-secrets` | @@ -46,8 +47,11 @@ threat-model → audit → triage → fix → report → ci | Skill | Role | |---|---| | `axguard-audit` | Pre-ship Lead — full A→Z | +| `axguard-security` | MCP-first agent security skill (when to call / which tool / how to read verdicts) | | `axguard-cso` | Chief Security Officer — confidence-gated lead pass | | `axguard-preship` | Focused checklist | | `axguard-triage` | False-positive filter | | `axguard-remediate` | Patch + re-audit | | `axguard-report` | HTML/MD deliverables | + +MCP (AI coding agents): `pip install -e '.[mcp]'` → `axguard mcp` · skill `axguard-security` · [docs/mcp.md](docs/mcp.md) diff --git a/README.md b/README.md index a486d5f..e79eef6 100644 --- a/README.md +++ b/README.md @@ -43,6 +43,21 @@ Pre-ship security gate — not a full pentest platform. Scan source, triage nois --- +## Pre-Ship Security + +Find → Explain → Fix → Verify → Ship. + +```bash +axguard preship . +``` + +AXGuard analyzes security-sensitive changes, verifies findings, checks attack paths and security regressions, and tells you whether the application is ready to ship. + +- Pre-Ship: [docs/preship.md](docs/preship.md) +- Security Diff: [docs/security-diff.md](docs/security-diff.md) + +--- + ## What is AXguard? **AXguard is a pre-ship security gate.** @@ -73,6 +88,8 @@ Use AXGuard as the security layer for your coding agent. ```text AI Agent ↓ +AXGuard Agent Skill + ↓ AXGuard MCP ↓ AXGuard Security Engine @@ -88,7 +105,13 @@ Agent Skills GitHub ``` -Primary agent tool: `axguard_security_review`. Install: `pip install -e '.[mcp]'` → `axguard mcp doctor` → configure your host ([docs/mcp-config.md](docs/mcp-config.md)). Overview: [docs/mcp.md](docs/mcp.md) · Tools: [docs/mcp-tools.md](docs/mcp-tools.md) · Security: [docs/mcp-security.md](docs/mcp-security.md). +| Layer | Role | +|---|---| +| **CLI** | Human security interface | +| **MCP** | AI-agent security interface (`axguard mcp`) | +| **Agent Skill** | Teaches agents *when* to use AXGuard (`skills/axguard-security`) | + +Primary agent tool: `axguard_security_review`. After a fix: `axguard_verify_fix`. Install: `pip install -e '.[mcp]'` → `axguard mcp doctor` → configure your host ([docs/mcp-config.md](docs/mcp-config.md)). Overview: [docs/mcp.md](docs/mcp.md) · Tools: [docs/mcp-tools.md](docs/mcp-tools.md) · Security: [docs/mcp-security.md](docs/mcp-security.md) · Skill roadmap: [docs/mcp-skill-roadmap.md](docs/mcp-skill-roadmap.md). --- @@ -222,6 +245,8 @@ open .findings/axguard/axguard-report.html | Security Memory | `axguard memory record .` → [docs/memory](docs/memory/README.md) | | Investigation Agent | `axguard investigate .` → [docs/investigation](docs/investigation/README.md) | | Predictive security risk | `axguard predict .` → [docs/predictive](docs/predictive/README.md) | +| Security Diff | `axguard diff` → [docs/security-diff.md](docs/security-diff.md) | +| Pre-Ship gate | `axguard preship .` → [docs/preship.md](docs/preship.md) | | Local Security Intelligence API | `axguard api start` → [docs/api](docs/api/overview.md) | | MCP for AI coding agents | `axguard mcp` → [docs/mcp.md](docs/mcp.md) | | GitHub PR bot (self-host) | `axguard github setup` → [docs/github](docs/github/README.md) | @@ -357,6 +382,16 @@ axguard predict --agent axguard predict --mcp axguard predict --what-if +# Security Diff (security-aware comparison of two versions) +axguard diff +axguard diff HEAD~1 +axguard diff main...HEAD +axguard diff --base main --head HEAD +axguard diff --json +axguard diff --verbose +axguard diff --fail-on high +axguard diff baseline save + # Training-data pipeline (no model training) axguard data discover axguard data inspect diff --git a/docs/mcp-skill-roadmap.md b/docs/mcp-skill-roadmap.md index d994e0b..892b5b3 100644 --- a/docs/mcp-skill-roadmap.md +++ b/docs/mcp-skill-roadmap.md @@ -1,14 +1,14 @@ # AXGuard MCP → Agent Skill Roadmap **Date:** 2026-09-17 -**Status:** Foundation only — **not** a full Agent Skill implementation. -**Related:** [mcp-research.md](./mcp-research.md), [mcp-threat-model.md](./mcp-threat-model.md) +**Status:** Skill implemented — `skills/axguard-security/` (behavioral wrapper over MCP; no duplicated engines). +**Related:** [mcp-research.md](./mcp-research.md), [mcp-threat-model.md](./mcp-threat-model.md), [mcp.md](./mcp.md) --- ## Purpose -Document how a future **AXGuard Agent Skill** should sit **above** MCP without duplicating security logic, tool schemas, or finding/evidence formats. +Document how the **AXGuard Agent Skill** sits **above** MCP without duplicating security logic, tool schemas, or finding/evidence formats. ```text Agent Skill ← teaches when/how to use AXGuard @@ -22,23 +22,25 @@ Do **not** implement the Skill as a replacement for MCP. --- -## Future skill YAML (from product brief) +## Skill package ```yaml name: axguard-security -description: Scan applications for security vulnerabilities, investigate findings, and help verify fixes before deployment. +description: Analyze code for security vulnerabilities, investigate findings, verify fixes, and assess security risk before deployment. ``` -The skill should instruct the agent roughly: +Path: `skills/axguard-security/SKILL.md` + +The skill instructs: ```text Before shipping security-sensitive code: -1. Use AXGuard security review. +1. Use AXGuard security review (MCP). 2. Investigate suspicious findings. 3. Review evidence and counter-evidence. 4. Check attack paths and regressions. 5. Separate verified findings from predictive risk. -6. Verify fixes before declaring an issue resolved. +6. Verify fixes with axguard_verify_fix before declaring resolved. ``` --- @@ -51,7 +53,7 @@ Before shipping security-sensitive code: | **MCP** | Tool/resource/prompt surface; structured outputs; policy gates; path/network limits | Host-agent pedagogy beyond tool descriptions | | **Core** | All security reasoning | Client-specific UX copy | -Reuse MCP tool names (`axguard_security_review`, `axguard_investigate`, …) and schemas so the Skill is a thin behavioral wrapper. +Reuse MCP tool names (`axguard_security_review`, `axguard_investigate`, `axguard_verify_fix`, …) and schemas so the Skill is a thin behavioral wrapper. --- @@ -86,8 +88,8 @@ MCP | Phase | Deliverable | Notes | |---|---|---| -| **Now** | Native MCP server (stdio), `axguard_security_review`, policy, docs | This initiative | -| **Next** | Agent Skill package (`axguard-security`) wrapping MCP tools | No duplicated engines; YAML + playbook only | +| **Done** | Native MCP server (stdio), `axguard_security_review`, policy, docs | PR #22 | +| **Done** | Agent Skill (`axguard-security`) + `axguard_verify_fix` | This initiative | | **Later** | GitHub Actions invoking the same core/API | MCP stays independent of GitHub | | **Later** | GitHub Security Review / App comments & checks | Reuse review engine; do not couple MCP transport to GitHub | @@ -112,11 +114,4 @@ All converge on the **same** AXGuard security engine. 2. Keep MCP output free of marketing; Skill may add human-facing onboarding separately. 3. Local-first: Skill install must not require AwareXone cloud. 4. Prefer teaching agents to call `axguard_security_review` before inventing ad-hoc scanner chains. - ---- - -## Non-goals (this document) - -- Full `SKILL.md` body or installer -- Shipping Skill files under `skills/` in this pass -- GitHub App / Actions implementation +5. Prefer `axguard_verify_fix` after remediations — never resolve on path rename alone. diff --git a/docs/mcp-tools.md b/docs/mcp-tools.md index 0fbfafd..af1b0d8 100644 --- a/docs/mcp-tools.md +++ b/docs/mcp-tools.md @@ -73,9 +73,12 @@ Does **not** modify source, execute exploits, or treat predictive risk as a veri | `axguard_list_findings` | AUTO | RO, idempotent | Summaries after scan/review. Use progressive disclosure. | | `axguard_get_finding` | AUTO | RO, idempotent | One finding: severity, confidence, location, verdict. | | `axguard_verify_finding` | AUTO / APPROVAL_REQUIRED (deep) | RO, idempotent | Hunter→Judge style verification for a candidate. | +| `axguard_verify_fix` | APPROVAL_REQUIRED | RO | After a fix: re-scan and return `RESOLVED` / `STILL_PRESENT` / `REGRESSED` by fingerprint. Never resolve on path rename alone. | Verdicts remain AXGuard-owned (`VERIFIED` · `LIKELY` · `UNVERIFIED` · `FALSE_POSITIVE` · `REQUIRES_REVIEW`). Agents must not “declare vulnerable” without this evidence path. +**Agent Skill:** prefer MCP tools via `skills/axguard-security` rather than inventing scan chains. + --- ## Evidence diff --git a/docs/mcp.md b/docs/mcp.md index e6faed8..ff97638 100644 --- a/docs/mcp.md +++ b/docs/mcp.md @@ -76,12 +76,30 @@ Typical agent result shape: decision, risk, verified findings, evidence, attack | Possible vulnerability to dig into | `axguard_investigate` | | Authz / agent / MCP permission changes | `axguard_security_review` | | “What attack paths does this create?” | `axguard_find_attack_paths` / review | -| Fix applied — confirm resolved | `axguard_security_review` or `axguard_verify_finding` | +| Fix applied — confirm resolved | `axguard_verify_fix` (never mark resolved on file edit alone) | Full catalog and approval tiers: [mcp-tools.md](mcp-tools.md). --- +## Agent Skill + +The behavioral layer above MCP (no duplicated scanners): + +```text +AI Coding Agent + ↓ +AXGuard Agent Skill (`skills/axguard-security`) + ↓ +AXGuard MCP + ↓ +AXGuard Security Engine +``` + +Skill teaches **when** to call AXGuard, **which** tool, and **how** to interpret VERIFIED / UNKNOWN / FALSE_POSITIVE / PREDICTIVE_RISK. Install via `./install.sh --agent agents` (or Cursor/Claude skill install). Roadmap: [mcp-skill-roadmap.md](mcp-skill-roadmap.md). + +--- + ## Quick start ```bash diff --git a/engines/mcp/policy.py b/engines/mcp/policy.py index 45e96b1..89de49f 100644 --- a/engines/mcp/policy.py +++ b/engines/mcp/policy.py @@ -33,6 +33,7 @@ class ApprovalTier(str, Enum): "axguard_list_findings": ApprovalTier.AUTO, "axguard_get_finding": ApprovalTier.AUTO, "axguard_verify_finding": ApprovalTier.APPROVAL_REQUIRED, + "axguard_verify_fix": ApprovalTier.APPROVAL_REQUIRED, # Evidence "axguard_get_evidence": ApprovalTier.AUTO, "axguard_get_evidence_chain": ApprovalTier.AUTO, diff --git a/engines/mcp/server.py b/engines/mcp/server.py index 44eb356..9f2a91a 100644 --- a/engines/mcp/server.py +++ b/engines/mcp/server.py @@ -244,6 +244,20 @@ def axguard_get_finding(finding_id: str, approved: bool = False) -> dict[str, An def axguard_verify_finding(finding_id: str | None = None, approved: bool = False) -> dict[str, Any]: return wrap_fn(HANDLERS["axguard_verify_finding"])(finding_id=finding_id, approved=approved) + @mcp.tool(name="axguard_verify_fix", description=by_name["axguard_verify_fix"]["description"], annotations=ann_fn("axguard_verify_fix")) + def axguard_verify_fix( + finding_id: str | None = None, + fingerprint: str | None = None, + path: str | None = None, + approved: bool = False, + ) -> dict[str, Any]: + return wrap_fn(HANDLERS["axguard_verify_fix"])( + finding_id=finding_id, + fingerprint=fingerprint, + path=path, + approved=approved, + ) + @mcp.tool(name="axguard_get_evidence", description=by_name["axguard_get_evidence"]["description"], annotations=ann_fn("axguard_get_evidence")) def axguard_get_evidence(finding_id: str | None = None, approved: bool = False) -> dict[str, Any]: return wrap_fn(HANDLERS["axguard_get_evidence"])(finding_id=finding_id, approved=approved) diff --git a/engines/mcp/tools/catalog.py b/engines/mcp/tools/catalog.py index 6a5235a..5437c06 100644 --- a/engines/mcp/tools/catalog.py +++ b/engines/mcp/tools/catalog.py @@ -199,6 +199,23 @@ "openWorldHint": False, }, ), + ( + "axguard_verify_fix", + ( + "Re-analyze after a remediation and classify the prior finding as " + "RESOLVED, STILL_PRESENT, or REGRESSED using fingerprints — never " + "mark resolved solely because a file path changed. Call after the " + "agent applies a fix for a known finding. Read-only re-scan. " + "Requires approved=true. Does not modify source." + ), + { + "title": "Verify Fix", + "readOnlyHint": True, + "destructiveHint": False, + "idempotentHint": False, + "openWorldHint": False, + }, + ), ( "axguard_get_evidence", ( diff --git a/engines/mcp/tools/handlers.py b/engines/mcp/tools/handlers.py index 3284452..3ed383f 100644 --- a/engines/mcp/tools/handlers.py +++ b/engines/mcp/tools/handlers.py @@ -275,6 +275,23 @@ def axguard_verify_finding(finding_id: str | None = None, approved: bool = False ) +def axguard_verify_fix( + finding_id: str | None = None, + fingerprint: str | None = None, + path: str | None = None, + approved: bool = False, +) -> dict[str, Any]: + from engines.mcp.tools.verify_fix import run_verify_fix + + return run_verify_fix( + finding_id=finding_id, + fingerprint=fingerprint, + path=path, + approved=approved, + session=_sess(), + ) + + def _ensure_evidence(sess: McpSession) -> dict[str, Any]: if sess.last_evidence: return sess.last_evidence @@ -650,6 +667,7 @@ def axguard_security_review_tool( "axguard_list_findings": as_tool(axguard_list_findings), "axguard_get_finding": as_tool(axguard_get_finding), "axguard_verify_finding": as_tool(axguard_verify_finding), + "axguard_verify_fix": as_tool(axguard_verify_fix), "axguard_get_evidence": as_tool(axguard_get_evidence), "axguard_get_evidence_chain": as_tool(axguard_get_evidence_chain), "axguard_get_counter_evidence": as_tool(axguard_get_counter_evidence), diff --git a/engines/mcp/tools/verify_fix.py b/engines/mcp/tools/verify_fix.py new file mode 100644 index 0000000..b1c1a1b --- /dev/null +++ b/engines/mcp/tools/verify_fix.py @@ -0,0 +1,196 @@ +"""Fix verification — re-analyze and compare finding fingerprints. + +Never mark RESOLVED solely because a path string changed. +""" + +from __future__ import annotations + +from typing import Any + +from engines.mcp import engines_bridge as bridge +from engines.mcp.policy import enforce +from engines.mcp.schemas.errors import McpError +from engines.mcp.schemas.results import success_result, truncate_result +from engines.mcp.session import McpSession, get_session + +_SEV_RANK = { + "critical": 5, + "high": 4, + "medium": 3, + "med": 3, + "moderate": 3, + "low": 2, + "info": 1, + "informational": 1, + "unknown": 0, +} + + +def _severity_rank(value: Any) -> int: + return _SEV_RANK.get(str(value or "unknown").strip().lower(), 0) + + +def _fingerprint_of(finding: dict[str, Any]) -> str: + explicit = finding.get("fingerprint") or finding.get("finding_fingerprint") + if explicit: + return str(explicit) + try: + from engines.memory.fingerprints import finding_fingerprint + + return finding_fingerprint(finding) + except Exception: # noqa: BLE001 + for key in ("id", "finding_id", "rule_id"): + if finding.get(key): + return str(finding[key]) + return "" + + +def _ids_of(finding: dict[str, Any]) -> set[str]: + out: set[str] = set() + for key in ("id", "finding_id", "rule_id", "fingerprint", "finding_fingerprint"): + v = finding.get(key) + if v: + out.add(str(v)) + return out + + +def _match_finding( + findings: list[dict[str, Any]], + *, + finding_id: str | None, + fingerprint: str | None, +) -> dict[str, Any] | None: + fp_target = (fingerprint or "").strip() + id_target = (finding_id or "").strip() + for f in findings: + if not isinstance(f, dict): + continue + fp = _fingerprint_of(f) + if fp_target and fp == fp_target: + return f + ids = _ids_of(f) + if id_target and (id_target in ids or any(id_target in x for x in ids)): + return f + if fp_target and fp_target in ids: + return f + return None + + +def _compare( + before: dict[str, Any], + after: dict[str, Any] | None, +) -> str: + if after is None: + return "RESOLVED" + if _severity_rank(after.get("severity")) > _severity_rank(before.get("severity")): + return "REGRESSED" + return "STILL_PRESENT" + + +def run_verify_fix( + *, + finding_id: str | None = None, + fingerprint: str | None = None, + path: str | None = None, + approved: bool = False, + session: McpSession | None = None, +) -> dict[str, Any]: + """Re-scan and classify fix status for a previously observed finding.""" + sess = session or get_session() + sess.begin_tool() + enforce("axguard_verify_fix", approved=approved) + + if not finding_id and not fingerprint: + raise McpError( + "INVALID_INPUT", + "finding_id or fingerprint is required for fix verification.", + ) + + prior = _match_finding( + list(sess.last_findings or []), + finding_id=finding_id, + fingerprint=fingerprint, + ) + if prior is None and sess.last_review: + review_findings = list(sess.last_review.get("verified_findings") or []) + prior = _match_finding( + review_findings, + finding_id=finding_id, + fingerprint=fingerprint, + ) + + if prior is None: + raise McpError( + "INSUFFICIENT_EVIDENCE", + "No prior finding in session cache to verify against. " + "Run axguard_security_review or axguard_list_findings first.", + details={ + "finding_id": finding_id, + "fingerprint": fingerprint, + "hint": "Do not mark RESOLVED from a file edit alone.", + }, + ) + + prior_fp = fingerprint or _fingerprint_of(prior) + target = bridge.require_path(sess, path) + scan_root = target if target.is_dir() else sess.project_root + result = bridge.run_scan_engine(sess, scan_root) + after_findings = list(result.get("findings") or []) + # Keep session findings current after re-analysis + sess.last_findings = after_findings + + matched = _match_finding( + after_findings, + finding_id=finding_id or str(prior.get("id") or prior.get("finding_id") or ""), + fingerprint=prior_fp, + ) + # If id lookup fails but fingerprint matches any row, use that + if matched is None and prior_fp: + for f in after_findings: + if isinstance(f, dict) and _fingerprint_of(f) == prior_fp: + matched = f + break + + status = _compare(prior, matched) + payload = { + "status": status, + "verification_status": status, + "finding_id": finding_id or prior.get("id") or prior.get("finding_id"), + "fingerprint": prior_fp, + "before": { + "severity": prior.get("severity"), + "title": prior.get("title") or prior.get("message"), + "file": prior.get("file") or (prior.get("location") or {}).get("file") + if isinstance(prior.get("location"), dict) + else prior.get("file"), + }, + "after": None + if matched is None + else { + "severity": matched.get("severity"), + "title": matched.get("title") or matched.get("message"), + "file": matched.get("file") + or ( + (matched.get("location") or {}).get("file") + if isinstance(matched.get("location"), dict) + else matched.get("file") + ), + "fingerprint": _fingerprint_of(matched), + }, + "recommended_action": { + "RESOLVED": "Finding no longer present after re-analysis. Confirm with review if high-impact.", + "STILL_PRESENT": "Issue remains after re-scan. Continue remediation; do not mark closed.", + "REGRESSED": "Severity worsened or issue intensified. Stop and re-investigate.", + }.get(status, "Re-run axguard_security_review."), + "note": "Never mark resolved solely because a file path changed.", + } + trimmed = truncate_result( + payload, + max_bytes=sess.limits.max_output_bytes, + max_list=sess.limits.max_list_items, + ) + return success_result( + trimmed if isinstance(trimmed, dict) else {"data": trimmed}, + state="OBSERVED", + confidence="HIGH" if status == "RESOLVED" else "MEDIUM", + ) diff --git a/fixtures/mcp_benchmark/cases/08_fix_verification/case.json b/fixtures/mcp_benchmark/cases/08_fix_verification/case.json new file mode 100644 index 0000000..67e8a2d --- /dev/null +++ b/fixtures/mcp_benchmark/cases/08_fix_verification/case.json @@ -0,0 +1,30 @@ +{ + "id": "08_fix_verification", + "category": "fix_verification", + "prompt": "I applied a fix for missing object-level authorization. Confirm whether the finding is resolved.", + "expected": { + "label": "SHOULD_CALL", + "primary_tool": "axguard_verify_fix", + "outcomes": ["RESOLVED", "STILL_PRESENT", "REGRESSED"], + "forbidden": ["mark resolved solely because path changed"] + }, + "cases": [ + { + "name": "fingerprint_gone", + "expected_status": "RESOLVED" + }, + { + "name": "same_fingerprint", + "expected_status": "STILL_PRESENT" + }, + { + "name": "severity_worse", + "expected_status": "REGRESSED" + }, + { + "name": "path_rename_same_fingerprint", + "expected_status": "STILL_PRESENT", + "must_not": "RESOLVED" + } + ] +} diff --git a/fixtures/mcp_benchmark/cases/09_skill_tool_selection/case.json b/fixtures/mcp_benchmark/cases/09_skill_tool_selection/case.json new file mode 100644 index 0000000..72b0d0e --- /dev/null +++ b/fixtures/mcp_benchmark/cases/09_skill_tool_selection/case.json @@ -0,0 +1,11 @@ +{ + "id": "09_skill_tool_selection", + "category": "selection_accuracy", + "prompt": "Agent Skill guidance: which MCP tool after a security-sensitive authz change, and after applying a fix?", + "expected": { + "label": "EXPECTED_TOOL", + "after_security_sensitive_change": "axguard_security_review", + "after_fix": "axguard_verify_fix", + "avoid_tools": ["shell", "network", "browser"] + } +} diff --git a/fixtures/mcp_benchmark/expected.json b/fixtures/mcp_benchmark/expected.json index a052084..7603d94 100644 --- a/fixtures/mcp_benchmark/expected.json +++ b/fixtures/mcp_benchmark/expected.json @@ -17,7 +17,8 @@ "when_not_to_call", "unknown_cases", "reject_malicious", - "security_regressions" + "security_regressions", + "fix_verification" ], "cases": [ "01_tool_discovery", @@ -26,7 +27,9 @@ "04_when_not_to_call", "05_unknown_cases", "06_reject_malicious", - "07_security_regressions" + "07_security_regressions", + "08_fix_verification", + "09_skill_tool_selection" ], "metrics": [ "tool_discovery", diff --git a/skills/axguard-security/SKILL.md b/skills/axguard-security/SKILL.md new file mode 100644 index 0000000..1f2feb3 --- /dev/null +++ b/skills/axguard-security/SKILL.md @@ -0,0 +1,59 @@ +--- +name: axguard-security +description: Analyze code for security vulnerabilities, investigate findings, verify fixes, and assess security risk before deployment. Use when reviewing security-sensitive changes, before ship/deploy, or when an agent should call AXGuard MCP instead of inventing its own scanner. +--- + +# AXguard Security (Agent Skill) + +This skill does **not** implement security analysis. It teaches when to use AXGuard, which MCP tool to call, and how to interpret results. + +**Primary interface:** AXGuard MCP tools (prefer over ad-hoc scans). +**Fallback:** if MCP is unavailable, `axguard scan .` / `axguard audit .` — then still apply the same verdict rules below. +Tool names: `references/mcp-tools.md` · catalog: [docs/mcp-tools.md](../../docs/mcp-tools.md). + +## When to call AXGuard + +**Call** after security-sensitive edits, or before ship/deploy, when changes touch: + +authn / authz / tenant / identity / permissions · DB · external HTTP · file / cmd · uploads · webhooks · secrets · OAuth · GraphQL · cloud · crypto · deps · AI agents / LLM tools / MCP · privileged ops + +**Do not call** for trivial non-security edits (typos, comments, pure renames with no security surface). + +## Workflow + +```text +Code change + → security-sensitive? + NO → skip AXGuard + YES → axguard_security_review + → findings? + NO → done (respect UNKNOWN / predictive separately) + YES → axguard_investigate (suspicious) + → fix + → axguard_verify_fix +``` + +| Step | Tool | Notes | +|------|------|--------| +| Default entry | `axguard_security_review` | Prefer this over chaining engines manually | +| Dig deeper | `axguard_investigate` | Suspicious / incomplete candidates | +| After a fix | `axguard_verify_fix` | Do not mark resolved on edit alone | +| Evidence (optional) | `axguard_get_evidence` / `axguard_get_counter_evidence` | When debating a finding | + +## Interpret results + +| Label | Agent rule | +|-------|------------| +| **VERIFIED** | Do not dismiss without **new** counter-evidence. Treat as real until AXGuard says otherwise. | +| **UNKNOWN** | Do **not** call the change safe. State uncertainty; investigate or escalate. | +| **FALSE_POSITIVE** | Do **not** report as a vulnerability. | +| **PREDICTIVE_RISK** | Risk signal only — **not** a confirmed vulnerability. Do not merge into verified counts. | + +Never invent evidence. Prefer AXGuard’s structured output over model speculation. + +## Response discipline + +- Prefer MCP when available; do not build a second scanner. +- Keep predictive risks and verified findings separate in summaries. +- Ship/go-no-go: blockers = VERIFIED (and policy-gated LIKELY if the review returns them) — not predictive-only noise. +- After fixes, call `axguard_verify_fix` (or re-review) before declaring closed. diff --git a/skills/axguard-security/references/mcp-tools.md b/skills/axguard-security/references/mcp-tools.md new file mode 100644 index 0000000..5893a7c --- /dev/null +++ b/skills/axguard-security/references/mcp-tools.md @@ -0,0 +1,17 @@ +# MCP tools (names only) + +Prefer these over inventing scanner chains. Full contracts: [docs/mcp-tools.md](../../../docs/mcp-tools.md). + +| Tool | Role | +|------|------| +| `axguard_security_review` | Primary entry — security impact of code / change | +| `axguard_investigate` | Deepen a suspicious finding | +| `axguard_verify_fix` | Confirm a remediation resolved the issue | +| `axguard_verify_finding` | Hunter→Judge verify a candidate (pre-fix) | +| `axguard_get_evidence` | Supporting evidence for a finding | +| `axguard_get_counter_evidence` | FP / blocked signals | +| `axguard_find_attack_paths` | Credible attack paths for a change / finding | +| `axguard_find_regressions` | Prior control removed or weakened | +| `axguard_predict_security_risks` | Predictive risk only — not confirmed vulns | + +**Skill preference:** MCP first. CLI fallback: `axguard scan .` / `axguard audit .`. diff --git a/skills/index.yaml b/skills/index.yaml index b04c1b8..c7ae59b 100644 --- a/skills/index.yaml +++ b/skills/index.yaml @@ -25,6 +25,10 @@ skills: - name: axguard-knowledge domain: operations path: skills/axguard-knowledge/SKILL.md + - name: axguard-security + domain: operations + path: skills/axguard-security/SKILL.md + related: [axguard-preship, axguard-audit, axguard-remediate] # Domain skills — discovery (3) - name: threat-modeling diff --git a/tests/test_mcp_benchmark.py b/tests/test_mcp_benchmark.py index a19919f..148d405 100644 --- a/tests/test_mcp_benchmark.py +++ b/tests/test_mcp_benchmark.py @@ -94,3 +94,20 @@ def test_security_regression_scenarios_listed(): assert "authorization_removed" in names assert "new_mcp_tool" in names assert case["expected"]["predictive_is_not_verified"] is True + + +def test_fix_verification_fixture(): + case = _load(FIXTURE / "cases" / "08_fix_verification" / "case.json") + assert case["expected"]["primary_tool"] == "axguard_verify_fix" + outcomes = set(case["expected"]["outcomes"]) + assert outcomes >= {"RESOLVED", "STILL_PRESENT", "REGRESSED"} + pytest.importorskip("engines.mcp.tools.catalog") + from engines.mcp.tools.catalog import tool_names + + assert "axguard_verify_fix" in tool_names() + + +def test_skill_tool_selection_fixture(): + case = _load(FIXTURE / "cases" / "09_skill_tool_selection" / "case.json") + assert case["expected"]["after_security_sensitive_change"] == "axguard_security_review" + assert case["expected"]["after_fix"] == "axguard_verify_fix" diff --git a/tests/test_mcp_skill.py b/tests/test_mcp_skill.py new file mode 100644 index 0000000..bbcba17 --- /dev/null +++ b/tests/test_mcp_skill.py @@ -0,0 +1,125 @@ +"""Contract tests for skills/axguard-security (Agent Skill over MCP). + +The Skill teaches when/which MCP tool to call — it must not embed a second scanner. +""" + +from __future__ import annotations + +from pathlib import Path + +import pytest + +ROOT = Path(__file__).resolve().parents[1] +SKILL = ROOT / "skills" / "axguard-security" / "SKILL.md" + +_SCANNER_ALGORITHM_MARKERS = ( + "def run_scan", + "def scan_", + "class Scanner", + "semgrep", + "bandit.run", + "regex = r\"", + "AST visitor", + "taint_propagate(", + "own scanner algorithm", +) + + +def _parse_frontmatter(text: str) -> tuple[dict[str, str], str]: + if not text.startswith("---"): + return {}, text + end = text.find("\n---", 3) + assert end != -1, "SKILL.md frontmatter not closed" + raw = text[3:end].strip() + body = text[end + 4 :] + meta: dict[str, str] = {} + for line in raw.splitlines(): + if ":" not in line: + continue + key, _, val = line.partition(":") + meta[key.strip()] = val.strip().strip("\"'") + return meta, body + + +@pytest.fixture(scope="module") +def skill_parts() -> tuple[dict[str, str], str, str]: + assert SKILL.is_file(), f"missing skill file: {SKILL}" + text = SKILL.read_text(encoding="utf-8") + meta, body = _parse_frontmatter(text) + return meta, body, text + + +def test_skill_file_exists(): + assert SKILL.is_file() + + +def test_skill_frontmatter_name_and_description(skill_parts): + meta, _body, _text = skill_parts + assert meta.get("name") == "axguard-security" + desc = meta.get("description") or "" + assert len(desc) >= 40 + low = desc.lower() + assert "security" in low + assert any( + tok in low + for tok in ("axguard", "mcp", "vulnerabilit", "review", "verify") + ) + + +def test_skill_teaches_when_to_use_and_not_use(skill_parts): + _meta, body, text = skill_parts + low = text.lower() + assert any( + tok in low + for tok in ( + "auth", + "authorization", + "ship", + "deploy", + "security-sensitive", + "meaningful", + ) + ), "skill should teach when to call AXGuard" + assert any( + tok in low + for tok in ("do not", "don't", "skip", "trivial", "typo", "rename", "comment") + ), "skill should teach when not to call AXGuard" + assert "when" in low + + +def test_skill_mentions_security_review_and_verify_fix(skill_parts): + _meta, _body, text = skill_parts + assert "axguard_security_review" in text + assert "axguard_verify_fix" in text + + +def test_skill_verdict_rules(skill_parts): + _meta, _body, text = skill_parts + upper = text.upper() + for label in ("VERIFIED", "UNKNOWN", "FALSE_POSITIVE"): + assert label in upper, f"skill must teach verdict rule {label}" + assert "PREDICTIVE" in upper, "skill must teach PREDICTIVE / PREDICTIVE_RISK" + + +def test_skill_does_not_embed_second_scanner(skill_parts): + _meta, body, text = skill_parts + low = text.lower() + assert any( + phrase in low + for phrase in ( + "does not implement", + "do not implement", + "not implement security", + "prefer mcp", + "primary interface", + "does **not** implement", + "thin behavioral", + "instead of inventing", + "second scanner", + "does not reimplement", + ) + ), "skill should state it wraps MCP rather than scanning itself" + for marker in _SCANNER_ALGORITHM_MARKERS: + assert marker.lower() not in low, f"skill embeds scanner marker: {marker}" + assert body.count("```") <= 6 + assert len(text) < 20_000 diff --git a/tests/test_mcp_verify_fix.py b/tests/test_mcp_verify_fix.py new file mode 100644 index 0000000..b7c77ce --- /dev/null +++ b/tests/test_mcp_verify_fix.py @@ -0,0 +1,212 @@ +"""Unit tests for axguard_verify_fix — mocked engines, no live network.""" + +from __future__ import annotations + +from pathlib import Path +from typing import Any +from unittest.mock import patch + +import pytest + +pytest.importorskip("engines.mcp.tools.handlers") +pytest.importorskip("engines.mcp.session") +pytest.importorskip("engines.mcp.policy") + +from engines.mcp.policy import ApprovalTier, enforce, tier_for +from engines.mcp.schemas.errors import McpError +from engines.mcp.session import reset_session +from engines.mcp.tools import handlers as handlers_mod +from engines.mcp.tools.handlers import HANDLERS + +TOOL = "axguard_verify_fix" + + +def _resolve_handler(): + fn = getattr(handlers_mod, TOOL, None) + if callable(fn): + return fn, True + wrapped = HANDLERS.get(TOOL) + if callable(wrapped): + return wrapped, False + pytest.fail(f"{TOOL} handler not registered") + + +def _status_of(out: dict[str, Any]) -> str: + if not isinstance(out, dict): + raise AssertionError(f"expected dict result, got {type(out)}") + if out.get("ok") is False: + err = out.get("error") or {} + code = err.get("code") if isinstance(err, dict) else None + raise AssertionError(f"tool returned error: {code or out}") + candidates: list[Any] = [ + out.get("status"), + out.get("verification_status"), + out.get("result"), + out.get("outcome"), + out.get("verdict"), + ] + nested = out.get("verification") or out.get("fix_verification") or out.get("data") + if isinstance(nested, dict): + candidates.extend( + [ + nested.get("status"), + nested.get("verification_status"), + nested.get("result"), + nested.get("outcome"), + ] + ) + for key in ("payload", "result", "review"): + blob = out.get(key) + if isinstance(blob, dict): + candidates.extend( + [ + blob.get("status"), + blob.get("verification_status"), + blob.get("result"), + blob.get("outcome"), + ] + ) + for c in candidates: + if isinstance(c, str) and c.strip(): + return c.strip().upper() + blob = str(out).upper() + for token in ("STILL_PRESENT", "REGRESSED", "RESOLVED"): + if token in blob: + return token + raise AssertionError(f"could not find verification status in: {out!r}") + + +def _finding( + *, + fid: str = "f-authz-1", + fingerprint: str = "fp.authz.idor.v1", + severity: str = "high", + path: str = "api.py", + title: str = "Missing object-level authorization", +) -> dict[str, Any]: + return { + "id": fid, + "finding_id": fid, + "fingerprint": fingerprint, + "rule_id": "auth-idor", + "title": title, + "severity": severity, + "status": "CONFIRMED", + "file": path, + "path": path, + "line": 2, + "message": "user_id reaches db without ownership check", + } + + +@pytest.fixture() +def project(tmp_path: Path) -> Path: + root = tmp_path / "app" + root.mkdir() + (root / "api.py").write_text( + "def get_user(user_id):\n return db.users[user_id]\n", + encoding="utf-8", + ) + return root + + +@pytest.fixture() +def session(project: Path): + return reset_session(project, mode="BALANCED") + + +def _call_verify(fn, raises: bool, **kwargs): + try: + out = fn(**kwargs) + except McpError: + raise + if isinstance(out, dict) and out.get("ok") is False and not raises: + err = out.get("error") or {} + code = err.get("code") if isinstance(err, dict) else "ANALYSIS_FAILED" + msg = err.get("message") if isinstance(err, dict) else str(out) + details = err.get("details") if isinstance(err, dict) else {} + raise McpError(code or "ANALYSIS_FAILED", msg or "error", details=details or {}) + return out + + +def _invoke(session, **kwargs): + fn, raises = _resolve_handler() + reset_session(session.project_root, mode="BALANCED") + from engines.mcp.session import get_session + + sess = get_session() + if getattr(session, "last_findings", None): + sess.last_findings = list(session.last_findings) + kwargs.setdefault("approved", True) + try: + return _call_verify(fn, raises, **kwargs), sess + except TypeError: + slim = dict(kwargs) + for drop in ("fingerprint", "path", "mode"): + slim.pop(drop, None) + try: + return _call_verify(fn, raises, **slim), sess + except TypeError: + continue + raise + + +def test_verify_fix_requires_approval_without_flag(): + tier = tier_for(TOOL) + assert tier == ApprovalTier.APPROVAL_REQUIRED + with pytest.raises(McpError) as ei: + enforce(TOOL, approved=False) + assert ei.value.code == "APPROVAL_REQUIRED" + + +def test_verify_fix_handler_rejects_unapproved(session): + fn, raises = _resolve_handler() + sess = reset_session(session.project_root) + sess.last_findings = [_finding()] + with pytest.raises(McpError) as ei: + _call_verify(fn, raises, finding_id="f-authz-1", approved=False) + assert ei.value.code == "APPROVAL_REQUIRED" + + +def test_resolved_when_fingerprint_gone_after_rescan(session): + session.last_findings = [_finding()] + empty = {"findings": [], "finding_count": 0} + with patch("engines.mcp.tools.verify_fix.bridge.run_scan_engine", return_value=empty) as scan: + out, _sess = _invoke(session, finding_id="f-authz-1", fingerprint="fp.authz.idor.v1") + assert _status_of(out) == "RESOLVED" + scan.assert_called() + + +def test_still_present_when_same_finding_remains(session): + prior = _finding() + session.last_findings = [prior] + same = {"findings": [dict(prior)], "finding_count": 1} + with patch("engines.mcp.tools.verify_fix.bridge.run_scan_engine", return_value=same) as scan: + out, _sess = _invoke(session, finding_id="f-authz-1", fingerprint="fp.authz.idor.v1") + assert _status_of(out) == "STILL_PRESENT" + scan.assert_called() + + +def test_regressed_when_severity_worsens(session): + prior = _finding(severity="medium") + session.last_findings = [prior] + worse = _finding(severity="critical") + rescan = {"findings": [worse], "finding_count": 1} + with patch("engines.mcp.tools.verify_fix.bridge.run_scan_engine", return_value=rescan) as scan: + out, _sess = _invoke(session, finding_id="f-authz-1", fingerprint="fp.authz.idor.v1") + assert _status_of(out) == "REGRESSED" + scan.assert_called() + + +def test_never_resolved_solely_from_path_string_change(session): + prior = _finding(path="api.py") + session.last_findings = [prior] + renamed = _finding(path="handlers/api.py") + rescan = {"findings": [renamed], "finding_count": 1} + + with patch("engines.mcp.tools.verify_fix.bridge.run_scan_engine", return_value=rescan) as scan: + out, _sess = _invoke(session, finding_id="f-authz-1", fingerprint="fp.authz.idor.v1") + status = _status_of(out) + assert status != "RESOLVED" + assert status in {"STILL_PRESENT", "REGRESSED"} + scan.assert_called()