Change the repository type filter
All
Repositories list
22 repositories
Evalgate
PublicPrompt and agent regression CI. The build fails when your prompt gets dumber. GitHub Action with PR delta comments.Agentrace
PublicObservability for Claude Code subagents. Reads session transcripts, flags the results you should not trust. Checks derived from real agent failures.agentpostmortem
PublicEvery AI agent failure, documented. Public case registry..github
PublicVaultRAG
PublicMCP-audit
PublicSecurity scanner and linter for MCP servers. npx mcp-audit <target>: 18 rules, SARIF output for GitHub code scanning.Voiceeval
PublicEvaluation for voice agents. Catches what text evals cannot see: mis-hearing, missing confirmation, latency, barge-in. Everyone can demo a voice agent; this tel…Ctxlens
PublicAnswerproof
PublicVerifiable, tamper-evident receipts for RAG answers. Merkle inclusion proofs and Ed25519 signatures.Ctxtrim
PublicTrim what bloats your AI coding context — find the files ballooning your Claude Code / Cursor / Codex token cost and write ignore files to cut it. Zero-dep. npx…tokencut
PublicSkill-audit
PublicSecurity scanner for agent skills — flags prompt-injection, dangerous shell, secret access, and exfiltration before you install a Claude/agent Skill. 31 rules, …Tenantq
PublicMulti-tenant hybrid-search reference on Qdrant: tenant-isolated retrieval, dense+sparse RRF fusion, HNSW tuning, Recall@K/p95 benchmarks, batch ingestion, Docke…RelayG
PublicA support ticket triage agent built as a LangGraph state machine. LLM classification, refund policy as pure Python, and a human-in-the-loop interrupt that pause…Injection-arena
PublicCasebook-Chat
PublicA streaming AI chat UI that investigates AI-agent failures. Searches the live AgentPostmortem case registry over MCP, pulls full case files, and answers with ci…Webhands
PublicA computer-use agent for the tools that have no usable API. Drives the real dashboard via Cloudflare Browser Rendering, returns clean structured data, and refus…Greenlite
PublicMobile command and approval cockpit for AI agents. Agents escalate a proposed action with its context; you approve or deny in one tap and it routes back to the …Resolvd
PublicAn end-to-end inbox operator. Triages, drafts, and acts within policy on inbound support messages: auto-resolves order lookups and refunds under the limit, esca…Casebook-MCP
PublicA remote MCP server that turns AgentPostmortem, a public registry of documented AI-agent failures, into tools any agent can query. Ships with a companion invest…Bridgekit
PublicA scoped MCP server exposing company tools (Shopify, Triple Whale, Postgres) to an AI stack with per-client permission boundaries and an append-only audit log. …Tracecase
PublicCI for AI agents. Record agent runs, replay them against prompt and model changes, and catch regressions and unsafe tool calls before they ship. Diffs each suit…
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.