crates.io · Docs · API · GitHub · DeepWiki · Changelog
The agent loop for Rust. Stream from any of 7 LLM protocols, run tools, loop until done.
A loop library, not an agent: it powers yoyo, a coding agent evolving its own source since March 2026, and runs from a laptop to a Cloudflare Worker.
git clone https://github.com/yologdev/yoagent && cd yoagent
ollama serve &
ollama pull llama3.1:8b # or pass --model <any pulled model>
cargo run --example cli -- --provider ollamaThat's a working coding agent in your terminal — file read/write/edit, shell, ripgrep search, streaming output, skills. No signup, no key, nothing to configure.
yoagent cli — mini coding agent
Type /quit to exit, /clear to reset
model: llama3.1:8b
cwd: /home/user/my-project
> find all TODO comments in src/
▶ search 'TODO' ✓
Found 3 TODOs:
src/main.rs:42: // TODO: handle edge case
src/lib.rs:15: // TODO: add tests
src/utils.rs:8: // TODO: optimize this
tokens: 1250 in / 89 out
Point it at a hosted model instead by swapping the flag:
ANTHROPIC_API_KEY=sk-... cargo run --example cli
GROQ_API_KEY=... cargo run --example cli -- --provider groq --model openai/gpt-oss-120b
cargo run --example cli -- --api-url http://localhost:1234/v1 --model my-model # LM Studio, llama.cpp, vLLM[dependencies]
yoagent = "0.25"
tokio = { version = "1", features = ["full"] }Building for wasm32? Use default-features = false (see
WebAssembly & Cloudflare Workers).
On a native target keep the default native feature: without it reqwest has no TLS, and every
HTTPS provider call fails at runtime.
An agent that actually uses a tool — the thing the crate exists for:
use yoagent::provider::ModelConfig;
use yoagent::{tools, Agent, AgentEvent, StreamDelta};
#[tokio::main]
async fn main() {
// The provider is selected from the config's protocol and the key is read
// from ANTHROPIC_API_KEY. Call `.with_api_key(k)` to pass one explicitly.
let mut agent = Agent::from_config(ModelConfig::claude_sonnet_5())
.with_system_prompt("You are a coding assistant.")
.with_tools(tools::default_tools());
let mut events = agent.prompt("Find every TODO in src/ and summarise them").await;
while let Some(event) = events.recv().await {
match event {
AgentEvent::MessageUpdate { delta: StreamDelta::Text { delta }, .. } => print!("{delta}"),
AgentEvent::ToolExecutionStart { tool_name, .. } => println!("\n▶ {tool_name}"),
AgentEvent::AgentEnd { .. } => break,
_ => {}
}
}
agent.finish().await;
}Swap the model by swapping the config — the provider follows, and the key is read from that provider's conventional env var:
Agent::from_config(ModelConfig::groq("openai/gpt-oss-120b", "GPT-OSS 120B")); // GROQ_API_KEY
Agent::from_config(ModelConfig::google("gemini-3.8-flash", "Gemini 3.8 Flash")); // GEMINI_API_KEY
Agent::from_config(ModelConfig::ollama("http://localhost:11434/v1", "llama3.1:8b")); // no keyEvery agent runs the same loop, and almost every hand-written one gets the same things wrong: streams that end badly, tools that panic or hang, context that runs out, runs that never stop, spend nobody can account for. yoagent writes that loop once, with each guarantee pinned by a test, and keeps one rule for what goes in: does every agent need it, the same way? Anything else is an extension, a feature or a companion crate (design philosophy).
So it ships no vector stores, embedding pipelines, or task-graph layer — if your problem is retrieval or orchestration, one of these is the better fit:
| If you need | Look at |
|---|---|
| RAG pipelines, vector stores, embeddings, transcription and image generation | rig — "Build modular and scalable LLM Applications in Rust" |
| Typed task graphs and streaming RAG indexing alongside agents | swiftide — "Composable LLM agents and harness, typed task graphs, and streaming RAG pipelines in Rust" |
| A tool-calling loop you host, gate, steer, branch, and record | yoagent |
What that focus bought:
- The loop is a free function.
agent_loop()is stateless and takes everything it needs as arguments.Agentis an optional wrapper that adds history and queues. You can drive the loop yourself without adopting our state model. - 7 native wire protocols, not one OpenAI-compat shim with adapters bolted on. Anthropic Messages, OpenAI Completions, OpenAI Responses, Azure, Gemini, Vertex, and Bedrock each have a real implementation, so provider-specific features (thinking budgets, prompt-cache breakpoints, reasoning deltas) survive instead of being flattened away.
- One plug-in contract for the whole run. An
Extensioncan add tools, check input, allow, modify or deny each tool call, redact results, and check the final answer, with state that starts fresh each run. Install it as host policy and it governs every sub-agent too. Hooks that guard a call fail closed: a policy that cannot run denies, an input check rejects, a redaction that fails withholds. - Steer a run that's already going. Inject guidance mid-flight; it's picked up between tool batches without restarting the turn.
- History is a tree, not a list.
Sessionforks, checkpoints, and seeks. Edit an earlier turn and re-run it without destroying the original branch. - Runs are recordable. With
features = ["gasp"], a run becomes an append-only semantic event log in a git repo — restore is clone + replay. Conformance-checked in CI. - The whole loop is testable offline.
MockProviderscripts multi-turn tool-calling conversations and honours cancellation, so abort and steering paths are testable with no network or key.
Everything that changes what the loop does goes through one contract, Extension: its run's
RunHooks see on_input, before_model, before_tool, after_tool, on_stop, every event,
and finish. A tool policy is a few lines:
use yoagent::extension::{ClonedHooks, RunHooks};
use yoagent::{ToolCallRequest, ToolDecision};
#[derive(Clone)]
struct NoForcePush;
#[async_trait::async_trait] // the `async-trait` crate
impl RunHooks for NoForcePush {
async fn before_tool(&self, call: &ToolCallRequest<'_>) -> ToolDecision {
let command = call.args["command"].as_str().unwrap_or_default();
if command.contains("push --force") {
ToolDecision::Deny("force-push is not allowed here".into())
} else {
ToolDecision::Allow
}
}
}
let agent = agent.with_extension(ClonedHooks::new("no-force-push", NoForcePush));yoagent's own budget, tool gate and input guard are built the same way. Six runnable
extension_* examples cover policies, redaction, verifiers, budgets, sub-agent policy and
audit logs (guide).
Plugins and other ecosystems. yoagent-rutis turns
rutis plugins — Rust, TypeScript or Python, loaded,
reloaded and unloaded at runtime — into one Extension. Through small adapters it also runs many
pi extensions and DSH tool plugins unchanged, with no
change to yoagent's core: tools, tool policies, input checks, prompt additions and images cross
over; commands and UI don't. An extension that hooks something the adapter can't enforce is
refused, not half-run.
The same loop builds natively and for wasm32-unknown-unknown. In yoyo's Cloudflare Worker we
measured about 4 ms to start and a median of ~45 ms of CPU per agent run — an agent
spends most of a run waiting on the model, which Workers do not bill as CPU. A minimal Worker
(one provider via Agent::from_provider, one tool, no decision model) is about 330 KiB gzipped. HTTP MCP works over the host's fetch;
yoagent-workers turns the Workers AI binding into a decision
model. See WebAssembly & Cloudflare Workers.
yoyo-evolve — a coding
agent that evolves its own source in public. It began as 200 lines of Rust; every commit since has
been agent-written and gated on tests. It runs on this loop with the
openapi feature enabled.
Also built on yoagent:
| Project | What it is |
|---|---|
rab |
A lightweight, extensible Rust coding agent |
greatsage |
"Rimuru's Unique Skill, you know the one" |
yoclaw |
OpenClaw reborn in Rust — a single-binary agent that remembers you |
Built something on yoagent? Open a PR and add it here — we'd like to see it.
| The loop | Full event stream; parallel, sequential or batched tools; steering and follow-ups; execution limits; retry with backoff and jitter | Agent loop · Events · Retry |
| Extensions | One plug-in contract for the run; Budget; the older hooks (ToolMiddleware, input filters, TurnHook, callbacks) still work |
Extensions · Callbacks |
| Providers | 7 native protocols reaching 20+ providers; thinking controls; prompt-cache hints; one context-overflow classifier | Providers · Prompt caching |
| Tools | bash, file read/write/edit, list_files, search; custom tools via one trait; MCP over stdio or HTTP; OpenAPI specs; per-run ToolSources |
Tools · MCP · OpenAPI |
| Sub-agents | Child loops with their own model and tools; large artifacts passed by reference; spend rolled up | Sub-agents |
| Context | Usage-calibrated tracking, tiered compaction, optional LlmCompaction, loop detection |
Context management |
| Sessions, skills, structured output | Branching session trees with JSONL; AgentSkills SKILL.md; typed prompt_structured::<T>() |
Sessions · Skills · Structured outputs |
Decision models (decision) |
Typed yes/no, one-of-N and score judgments in a few hundred ms (Jev, Clef, OpenAI Decisions, any logprobs server); a tool gate and an input guard | Decision models |
| Cost and telemetry | Opt-in pricing (nothing priced by default); SessionStats per run incl. sub-agents; tracing spans with tokens and cost |
Pricing · Telemetry |
Recording (gasp) |
serde on every core type; runs recorded into a GASP repo, plugin logs too (opt-in) | Persistence · GASP |
| WebAssembly | --no-default-features for wasm32-unknown-unknown, e.g. Cloudflare Workers |
WebAssembly & Workers |
Seventeen of the 22 runnable examples in examples/ are below; ten need no API key at all (eleven counting cli with a local model). The rest are live-provider harnesses and offline evaluation sweeps.
| Example | What it shows | Key needed |
|---|---|---|
cli |
A ~400-line coding agent — all tools, skills, streaming, colored output. Like a baby Claude Code | optional¹ |
rlm |
An LLM that explores a codebase on its own by spawning sub-agents | yes |
code_review |
Three sub-agents reviewing a diff in parallel, results merged | yes |
shared_state |
Passing a large artifact between sub-agents by reference | yes |
sub_agent |
Delegation basics with a per-sub-agent model | yes |
basic |
The smallest possible agent | yes |
callbacks |
Lifecycle hooks and a custom tool | no |
persistence |
Save and restore a session | no |
telemetry |
tracing spans with token and cost fields |
no |
gasp_emit |
Recording a run into a GASP repo | no |
decision |
Decision-model questions in one line, and attaching a model to an agent (feature decision) |
yes |
extension_policy, _redact, _verifier, _budget, _tree, _audit |
Extensions: a tool policy, redaction, a verifier, budgets, policy over sub-agents, an audit log (guide) | no² |
¹ --provider ollama or --api-url needs no key; hosted providers read their conventional env var.
² Scripted offline by default; -- --live uses DEEPSEEK_API_KEY or ANTHROPIC_API_KEY.
The companion crates have their own: language_plugins, pi_extensions and dsh_tools
(TypeScript, Python, pi and DSH plugins) and a Clef tool-gate Worker.
MockProvider scripts a whole multi-turn tool-calling conversation with no network, and honours
cancellation, so abort and steering paths are testable too. See
Testing Your Agent; how the crate itself is tested and what CI runs is
in CONTRIBUTING.
- The book — concepts, guides, a page per provider, and the architecture and module map (source)
- Design philosophy — why yoagent is a loop library, the one extension contract, and what that costs
- API reference — built with all features enabled
- CHANGELOG — every release
- CONTRIBUTING — how to build, test, and send a PR
MSRV is 1.86, enforced in CI. Raising it is a minor-version change.
Bug reports, ideas and PRs are welcome — open an issue or pick one labelled help wanted. CONTRIBUTING covers building, the checks CI runs, and the PR checklist. Report security issues privately as described in SECURITY.
rutis, the plugin runtime behind yoagent-rutis, and the
contributors who send fixes and ideas — including the ones who suggest them on X.
MIT — see LICENSE.