The system turns a single model-driven exploration of an application into a reusable capability that runs with no model involved. A model explores the application once and records what it did as a typed, versioned capability artifact. Every later run executes from that artifact deterministically, and when a run cannot proceed it pauses and hands the live browser session to a human.
Three principles define the design:
- The model discovers. Exploration is the only stage that uses an LLM. The model reads a numbered list of on-screen elements and returns one action at a time.
- The artifact is the capability. It is a reviewable JSON document, not a transcript or script. It declares parameters, outputs, per-step risk, and machine-checkable checkpoints for every state-changing step and for the run as a whole.
- Replay is deterministic. The replay engine imports no LLM client, and a test enforces this. The same artifact with the same parameters produces the same behaviour.
flowchart TD
subgraph D["Discovery (model in the loop, runs once)"]
G[Goal] --> P[Perceive] --> DM["Decide (model)"] --> A[Act]
A -- next action --> P
A --> R[Recorder]
end
R --> ART[(artifact.json)]
subgraph RP["Replay (no model, deterministic)"]
PA[Params] --> B[Bind] --> RL[Resolve ladder] --> AC[Act + checkpoint] --> C[Classify]
end
ART -- loaded with parameters --> B
C --> BO["Business outcome (exit 2)"]
C --> RC["Recoverable (cleared and retried)"]
C --> E["Escalate (exit 3)"]
C --> HF["Hard failure (exit 1)"]
E --> H["Human handoff: same live session, run resumes"]
Discovery and replay share one output, the artifact, and nothing else: the model never runs during replay. Both reach the browser through a single SurfaceDriver interface, where every action passes one choke point for the allowlist, risk policy and control token. A Windows UIA stub (desktop_stub.py) shows the interface is not tied to Playwright, and a test fails if mechanism-specific terms appear in the schema package.
Each step in an artifact names an action, a target, an optional value, and a risk level. Values that come from parameters are stored as parameter references, never as literals.
{
"id": "type_username",
"index": 1,
"action": "type",
"target": {
"description": "the Username field in the login form",
"candidates": [
{"strategy": "role_name", "role": "textbox", "name": "Username", "confidence": 0.95},
{"strategy": "label", "label": "Username", "confidence": 0.85},
{"strategy": "dom_path", "raw": "input[name='username']", "confidence": 0.4}
],
"fingerprint": {"role": "textbox", "name": "Username", "tag": "input"}
},
"value": {"kind": "param", "param": "username"},
"risk": "safe"
}Locator ladder. Candidates are tried in order, most portable first. The top rung describes the control the way a person would (role and accessible name). The last rung is a CSS path, kept only as a fallback; when it is the rung that matches, the run reports drift. Fingerprints detect when a selector still matches but lands on a different element than the one recorded.
Checkpoints. State-changing steps carry assertions about what must become true, not just confirmation that an action fired:
"checkpoint": {
"description": "the member is logged in and looking at their accounts",
"assertions": [{"kind": "text_present", "value": "Accounts Overview"}]
}Inferred names. Many legacy applications publish no accessible name for their controls. ParaBank's login fields, for example, have no label, aria-label or placeholder. The driver derives a name from the visible caption in the preceding block or the adjacent table cell, and marks it with "name_inferred": true. The flag is advisory and does not change matching. capability validate warns on any step that leads with an inferred name or a CSS path.
Every step result is classified in a fixed order, and the order is deliberate:
- BUSINESS_OUTCOME. The application declined, and the artifact declared that answer in advance (for example, "no transactions in range"). Checked first so a declared answer is never reported as a failure.
- RECOVERABLE. A known nuisance the artifact can clear, such as an interstitial. Retried up to twice per step and five times per run, then reclassified.
- ESCALATE. The run is safe but stuck: an ambiguous target, a risky step in unattended mode, or a blocked action.
- HARD_FAILURE. Anything else. The run fails with evidence.
Fault paths can be exercised with --inject (interstitial, session_timeout, slow_load, server_error, element_missing; repeatable).
Human handoff. An escalated run writes an intervention request, releases a control token, and waits. A human attaches to the same live browser over CDP:
capability attach --session ws://127.0.0.1:9222/devtools/browser/...The control token is a state machine: while a human holds it, any automated action raises an error instead of competing for input. On resume, the engine re-perceives the page, retries the step once, and continues the run rather than restarting it. The default --operator none writes the intervention and exits as escalated; --operator console blocks in the terminal until the operator types r.
All guardrails are enforced at the driver's single action choke point and configured in three YAML files under config/.
| Guardrail | Behaviour | Config |
|---|---|---|
| Allowlist | Restricts domains, URL path patterns and action types. Also applied as a browser route guard, so in-page redirects cannot leave the permitted origin. Fails closed. | allowlist.yaml |
| Risk policy | Every step has a risk level. Submit-like actions with no matching rule default to risky. A risky step in --mode unattended escalates. |
risk_policy.yaml |
| Redaction | Declared-sensitive parameters are redacted by identity before any write; regex patterns catch undeclared values. Covers run logs, results and model transcripts. | redaction.yaml |
Redaction does not apply to screenshots. Evidence directories should be treated as containing anything that was visible on screen.
| Command | Purpose | Needs a model |
|---|---|---|
discover |
Explores an application and records an artifact | Yes |
replay |
Executes a recorded artifact | No |
validate |
Checks an artifact without opening a browser | No |
attach |
Joins the live session of an escalated run | No |
fixture |
Serves the offline ParaBank stand-in | No |
Capabilities are called by software, so exit codes are part of the contract:
| Code | Meaning |
|---|---|
| 0 | Success, result returned |
| 2 | Business outcome (success, nothing to return) |
| 3 | Escalated |
| 1 | Hard failure |
| 4 | Bad invocation |
Callers can branch on 0 versus 2 without parsing output.
Configuration is intentionally minimal. replay, validate, attach and fixture read no environment variables. Policy lives in config/*.yaml and CLI flags, so the policy behind any run can be reconstructed from the artifact, config and run log.
Only discover needs environment setup, and only when using a real model. Copy .env.example to .env and set one key. Exported variables take precedence over .env. LLM_PROVIDER and LLM_MODEL optionally override --provider and --model.
| Provider | Key | Default model |
|---|---|---|
anthropic |
ANTHROPIC_API_KEY |
claude-3-5-sonnet-latest |
openai |
OPENAI_API_KEY |
gpt-4o |
groq |
GROQ_API_KEY |
openai/gpt-oss-120b |
The groq provider works with any OpenAI-compatible endpoint (vLLM, Ollama and others) via base_url. Use a plain model rather than groq/compound, since its server-side web search and code execution conflict with discovery's rule that only on-screen elements exist.
Model choice. Use the strongest model available. Discovery runs once per capability, but its judgment, especially each checkpoint it writes, is frozen into an artifact that may run unattended thousands of times.
The reference capability logs in to the live ParaBank demo as john, stays on Accounts Overview, and reads the account number and balance of the first three rows. Discovery completed in 9 steps, and replay of the same artifact with no model completed 9 of 9 steps.
| Item | Value |
|---|---|
| Artifact | artifacts/parabank.log_as_member_accounts.v1.json |
| Parameters | params-login.json (username, password) |
| Discovery model | Groq openai/gpt-oss-120b |
| Discovery evidence | evidence/discovery-20260819T233525-8169b3 |
| Replay evidence | evidence/replay-success-20260819T233946-17582a |
Replay output:
| Row | Account | Balance |
|---|---|---|
| 1 | 12345 | -$2,300.00 |
| 2 | 12456 | $10.45 |
| 3 | 12567 | $100.00 |
The recorded output names are inconsistent (for example 12345_2300_00_0_00), but the values are correct.
Offline test suite. All 241 tests run without network, Docker or an API key, against a bundled ParaBank stand-in in fixtures/parabank/. Tests and scripts/demo.py use a separate fixture artifact, tests/fixtures/parabank.read_account_activity.v1.json, which takes an account_id parameter. Fault injection is exercised against this artifact so every error path is actually executed.
- Install:
python -m venv .venv && .venv/Scripts/activate # use bin/activate on macOS/Linux pip install -e ".[dev,llm]" # omit llm for replay only playwright install chromium
- Validate the reference artifact (no browser):
capability validate --artifact artifacts/parabank.log_as_member_accounts.v1.json
- Replay it against live ParaBank:
capability replay --artifact artifacts/parabank.log_as_member_accounts.v1.json --params-file params-login.json --verbose
- Run the offline suite:
pytest
- Optionally, rediscover with a model (requires a key in
.env):capability discover --url https://parabank.parasoft.com/parabank/index.htm --params-file params-login.json --goal "Log in as the member. On Accounts Overview, stay on that page and read only the first 3 rows of the accounts table: for each of those 3 rows, read the account number and the Balance. Do not click into an account. Do not read more than 3 accounts. Then done."
Prefer --params-file over inline JSON, since PowerShell strips quotes from inline arguments. Never distribute .venv or .env.