Context engineering best practices for Claude Code. Budget zones, degradation detection, token-saving hooks, model switching, and more.
10 techniques packaged as one installable plugin - no configuration needed.
| Technique | Type | Purpose |
|---|---|---|
| Budget Zones | Skill | GREEN/YELLOW/ORANGE/RED context usage zones |
| Degradation Detection | Skill | Self-diagnosis when Claude starts forgetting |
| Code Intelligence | Skill (background) | Efficient tool selection for code navigation |
| Model Switching | Skill (background) | Task-to-model mapping (Haiku/Sonnet/Opus) |
| Prompt Caching | Skill (background) | Cache-friendly workflow - cache reads cost ~0.1x |
| Thinking Control | Skill (background) | Calibrated reasoning depth per task type |
| Test Output Filter | Hook | ~80% token savings on test runs |
| Build Output Filter | Hook | ~90% token savings on builds |
| Lint Output Filter | Hook | ~70% token savings on lints |
| Fresh Context Pattern | Command | TASK.md + PROGRESS.md handoff files |
- macOS, Linux, or Windows (via WSL or Git Bash - both required by Claude Code)
jq(pre-installed on most dev machines;brew install jq/apt install jqif missing)
# Add the marketplace
claude plugin marketplace add silvesterdivas/context-engineer
# Install the plugin
claude plugin install context-engineer@context-engineer-marketplacePrefer the interactive UI? Run /plugin inside Claude Code and use the Discover and Marketplaces tabs to do the same thing.
context-engineer is a third-party marketplace, so updates are not automatic by default. When a new version ships, refresh the marketplace and reinstall from inside Claude Code:
/plugin marketplace update context-engineer-marketplace
/plugin install context-engineer@context-engineer-marketplace
/reload-plugins
/reload-plugins applies the new version in your current session without a full restart. To update automatically on startup instead, run /plugin, open the Marketplaces tab, select context-engineer-marketplace, and toggle Enable auto-update.
After installing, run the setup command to add context engineering rules to your project:
/context-engineer:setup
Then run a health check:
/context-engineer:diagnose
That's it. The background skills and hooks activate automatically.
Real situations and exactly what to do. Every command and signal below ships with the plugin.
Symptom: Responses feel slow and costly, or you are watching spend climb on a long session.
What to run:
/context-engineer:diagnose
The scorecard ends with a live session line, and the budget-warning hook already prints a zone note after each tool call. To actively cut cost, ask Claude "how do I cut token costs?" to trigger the prompt-caching skill.
What you see:
Session: ~120K ctx tokens, 78% cached, est. $0.34/turn (Opus 4.8 rates)
A high cache-hit rate is good: cache reads cost about 0.1x. If it is low, that is your lever. Keep CLAUDE.md and early context stable, batch edits to the same file, and avoid re-reading files you opened earlier so the prompt cache stays warm.
Symptom: Claude drops earlier decisions, repeats a failed fix, or invents APIs. Context is filling up.
What to run:
/context-engineer:fresh-context "the work still left to do"
What you see: The budget hook escalates through zones on its own as context fills:
[context-budget] YELLOW ZONE - Context: 65% [650000/1000000 tokens, 78% cached]. Be selective: prefer Grep over Read, summarize before processing large files.
In YELLOW and ORANGE the budget-zones and degradation-detection skills make Claude conserve automatically. At RED, fresh-context writes TASK.md + PROGRESS.md so you can open a clean session and continue with full awareness instead of fighting a degraded one.
Symptom: Lots of MCP tools you never use load on every turn, eating context budget up front.
What to run:
/context-engineer:audit-mcp
What you see: A per-server table with tool counts, estimated token overhead, and a keep/remove recommendation:
| Server | Tools | Est. Tokens | Recommendation |
|----------|-------|-------------|----------------|
| comfyui | ~110 | ~38,000 | Remove / scope |
| stitch | ~14 | ~4,900 | Remove |
Move heavy, project-specific servers into that project's own .mcp.json instead of a workspace-wide one, then restart so the change takes effect (MCP config loads at startup).
Symptom: Reaching for Opus on everything, or unsure whether a job could run cheaper and faster.
What to run: Nothing to invoke; the model-switching skill guides this in the background. Ask "which model should I use for this?" to surface it directly.
What you see: A task-to-model mapping you can act on:
- Haiku for file search, grep, quick lookups, and simple edits
- Sonnet for code review, refactors, and multi-file changes
- Opus for architecture, complex debugging, and security review
Opus is about 5x Haiku, so reserve it for judgment-heavy work. Delegate broad searches to the investigator (Haiku) subagent and reviews to the reviewer (Sonnet) subagent to keep your main context clean.
Starting fresh? Here's the complete workflow.
mkdir my-new-project
cd my-new-project
git initOpen Claude Code in your project directory and run:
/context-engineer:setup
This creates a CLAUDE.md in your project root with the full Context Engineering Rules section: budget zones, fresh context pattern, tool efficiency guidelines, model switching, and output filtering confirmation.
/context-engineer:diagnose
This runs a health scorecard checking 6 areas:
| Area | What It Checks |
|---|---|
| CLAUDE.md configuration | Rules section exists and is complete |
| Token-saving hooks | All 3 hooks installed and functional |
| MCP server hygiene | No unnecessary servers wasting tokens |
| Fresh context files | TASK.md/PROGRESS.md templates ready |
| Git hygiene | Repository is clean and well-structured |
| Project structure | Files organized for efficient AI navigation |
That's it. Here's what happens behind the scenes:
When you run tests (npm test):
- Before: Claude sees all 200 lines of test output (196 passed, 4 failed)
- After: Claude sees only the 4 failures + summary - ~80% token savings
When you run a build (npm run build):
- Before: Claude reads 500 lines of webpack output
- After: Claude sees only errors and warnings - ~90% token savings
When you run a linter (npx eslint src/):
- Before: Claude processes 150 lines of lint output
- After: Claude sees only the problems - ~70% token savings
Hooks only activate when output exceeds 30 lines. Short outputs pass through unchanged.
- Let Claude explore freely at first. You're in GREEN zone (< 60% context). Read full files, make architecture decisions, explore freely.
- Use Opus for early decisions. Model switching routes architecture decisions to Opus where it matters most.
- Create TASK.md early for complex features. Run
/context-engineer:fresh-context "building the auth system"before starting multi-session work.
You have a codebase - maybe it already has a CLAUDE.md, maybe it doesn't. Context-engineer fits right in.
Navigate to your project and open Claude Code:
/context-engineer:setup
- No CLAUDE.md? One gets created with the complete Context Engineering Rules section.
- Existing CLAUDE.md? The setup command reads your file and appends the rules. Your existing content stays untouched.
- Already have the rules? Setup detects it and skips. No duplication.
Existing projects often accumulate MCP servers over time. Each one adds token overhead:
/context-engineer:audit-mcp
You'll get a table showing each server's name, tool count, estimated token overhead, and a recommendation (keep, review, or remove).
/context-engineer:diagnose
For existing projects, pay attention to:
- Token-saving hooks - are they matching your build tools?
- CLAUDE.md configuration - did the rules integrate cleanly with your existing content?
- Project structure - the scorecard may flag reorganization opportunities.
Your workflow doesn't change. You just get dramatically better sessions.
- Your first session will feel different. The hooks silently save thousands of tokens per test/build/lint cycle, and the budget zones now track against Claude Code's ~1M-token context window, so sessions run much longer before handoff.
- Watch the budget zones on large codebases. You'll hit YELLOW and ORANGE faster because there's more code to read. The budget zone skill automatically guides Claude to be selective.
- Use fresh context for big refactors. Run
/context-engineer:fresh-context "refactoring the payment module"before starting. When context gets heavy, TASK.md and PROGRESS.md are ready for a clean handoff. - Don't fight the RED zone. When context exceeds 85%, the plugin guides Claude to wrap up. A fresh session with handoff files outperforms a degraded session every time.
Context usage determines how Claude behaves:
| Zone | Context Used | Behavior |
|---|---|---|
| GREEN | < 60% | Full exploration. Read entire files, no constraints. |
| YELLOW | 60-75% | Selective. Prefers Grep over Read. Summarizes before processing. |
| ORANGE | 75-85% | Conservation mode. Line ranges only, targeted searches, delegates to subagents. |
| RED | > 85% | Wrap up. Finishes current task, creates TASK.md + PROGRESS.md, suggests new session. |
Percentages are measured against Claude Code's context window (~1M tokens on Opus 4.8 / Sonnet 4.6). The budget hook reads real token usage from the transcript; override the assumed window with the CONTEXT_ENGINEER_BUDGET environment variable. Budget zones activate automatically. No configuration needed.
Complex tasks survive across sessions using handoff files.
- A feature will take more than one session
- You're approaching ORANGE/RED zone with more work to do
- You want to hand off work to yourself tomorrow
/context-engineer:fresh-context "implementing JWT authentication"
Creates two files:
- TASK.md - Goal, constraints, key files, decisions made, current context
- PROGRESS.md - Steps completed, current state, next steps, blockers
Open Claude Code in the same directory. Claude reads CLAUDE.md automatically. Then:
Read TASK.md and PROGRESS.md and continue where the last session left off.
Fresh context, full awareness, zero token baggage.
The plugin guides Claude to use the right model for each task:
| Task Type | Model | Price (in/out per 1M) | Relative cost |
|---|---|---|---|
| File search, grep, quick lookups | Haiku 4.5 | $1 / $5 | 1x |
| Code review, refactoring, multi-file changes | Sonnet 4.6 | $3 / $15 | ~3x |
| Architecture decisions, complex debugging, security review | Opus 4.8 | $5 / $25 | ~5x |
This works automatically through the model switching skill. Note that Opus is now ~5x Haiku (it was 25x under older pricing), so model choice is a smaller cost lever than it used to be - prompt caching (cache reads ~0.1x) saves far more.
Four stages of context degradation with automatic detection:
| Stage | Context | Signs |
|---|---|---|
| 1 | ~60% | Forgetting details, asking for re-confirmation |
| 2 | ~75% | Contradicting decisions, missing imports |
| 3 | ~85% | Hallucinating APIs, looping on same fix |
| 4 | ~90%+ | Incoherent reasoning, nonsensical code |
When degradation is detected, the skill flags it in real time so you can act before quality collapses.
Starting a session:
- Open Claude Code in your project - CLAUDE.md loads automatically
- Hooks are already active - no action needed
- Continuing previous work? "Read TASK.md and PROGRESS.md and pick up where we left off"
During a session:
- Tests/builds/linting are filtered automatically
- Budget zones adapt silently as context fills
- Degradation detection watches in the background
- Model switching guides subagent usage
Ending a session:
- Task done? Commit your code
- Task continues? Run
/context-engineer:fresh-context "remaining work" - Hit RED zone? Claude suggests creating handoff files automatically
Periodic maintenance:
/context-engineer:diagnoseif sessions feel slow/context-engineer:audit-mcpafter adding new MCP servers- Update CLAUDE.md if project architecture changes
Hooks only activate on output exceeding 30 lines. Run a test suite with enough tests to generate verbose output.
Run /context-engineer:diagnose to check. The setup command checks for existing rules before adding. If duplication occurred, remove the duplicate section manually.
Expected for large codebases. Budget zones and model switching extend your session, but use fresh context more frequently. Focus sessions on specific subsystems.
Edit the hook scripts at hooks/scripts/. Each script has a pattern matcher at the top.
Set CONTEXT_ENGINEER_FILTER_MIN_LINES, or edit FILTER_MIN_LINES in hooks/scripts/filter-common.sh.
Adds context engineering rules (budget zones, fresh context pattern, tool efficiency guidelines) to your project's CLAUDE.md.
Creates TASK.md and PROGRESS.md with current progress so you can continue in a new conversation with full context. Use when your context is getting heavy or before a complex task.
Lists all configured MCP servers, estimates their token overhead, and flags unused ones that are wasting context space.
Runs a health scorecard checking: CLAUDE.md config, hooks, MCP hygiene, fresh context files, git state, and project structure. Produces a pass/warn/fail table with specific fix recommendations. As of v1.2.0 it also prints a live session line (context size, cache-hit rate, and estimated cost per turn at Opus 4.8 rates) when an active transcript and jq are available.
Two kinds of skills ship with the plugin. User-invocable skills (Budget Zones, Degradation Detection) can be called any time. Background skills (Code Intelligence, Model Switching, Prompt Caching, Thinking Control) activate automatically based on what you are doing; you can also trigger one on demand by describing the situation. For example, "how do I cut token costs?" or "make this cache-friendly" activates Prompt Caching, and "which model should I use?" activates Model Switching.
Defines four context usage zones - GREEN (free), YELLOW (selective), ORANGE (conserve), RED (wrap up). Activates automatically when context fills, or invoke manually to check your current zone.
Four stages of context degradation with specific signs and corrective actions. Helps Claude recognize when it's starting to forget, hallucinate, or loop.
Guides efficient tool selection: when to Grep vs Read, how to navigate imports and type hierarchies, token cost comparisons for different navigation approaches.
Maps task types to optimal models. Haiku for searches, Sonnet for reviews, Opus for architecture. Includes current pricing (Haiku $1/$5, Sonnet $3/$15, Opus $5/$25 per 1M) and delegation guidance.
Cache-friendly workflow guidance: caching is a prefix match, so keep CLAUDE.md and early context stable, batch edits, and avoid re-reading early files. Cache reads cost ~0.1x, the biggest token-cost lever there is.
Calibrates reasoning depth: minimal for mechanical tasks, deep for security/architecture. Prevents over-thinking simple tasks and under-thinking complex ones.
Fast, cheap codebase search agent. Use via the Task tool for broad searches that would clutter your main context. Tools: Read, Grep, Glob.
Code review agent with fresh context. Reads code thoroughly and provides structured feedback (critical/warning/suggestion). Tools: Read, Grep, Glob, Bash (git only).
Three PostToolUse hooks automatically filter verbose Bash output:
| Hook | Matches | Keeps | Savings |
|---|---|---|---|
filter-test-output.sh |
jest, vitest, pytest, go test, cargo test, etc. | FAIL/ERROR lines + summary | ~80% |
filter-build-output.sh |
tsc, gradle, xcodebuild, cargo build, etc. | Error/warning lines | ~90% |
filter-lint-output.sh |
eslint, pylint, clippy, biome, etc. | Problem lines only | ~70% |
Hooks only activate on output > 30 lines. Short output passes through unfiltered.
Remove the corresponding entry from hooks/hooks.json.
Set CONTEXT_ENGINEER_FILTER_MIN_LINES, or edit FILTER_MIN_LINES in hooks/scripts/filter-common.sh.
Set CONTEXT_ENGINEER_FILTER_OFF=1 to turn the output filters off entirely. Every filtered summary also prints this hint, so you can always get the full output back when a filter trims a line you needed.
Run /context-engineer:setup then edit the generated section in your CLAUDE.md.
Copy the agent files to your project's .claude/agents/ directory and modify the model frontmatter.
- New:
/context-engineer:statuscommand - live budget zone, session cost, and a menu of every command, skill, and hook. - New: Model-aware session cost - the diagnose and scorecard cost line prices Opus, Sonnet, and Haiku sessions correctly instead of always at Opus rates.
- New: Filter escape hatch - set
CONTEXT_ENGINEER_FILTER_OFF=1to pass raw output through; every filtered summary advertises the switch. - New: Expanded filter coverage - gradle and xcodebuild (build); biome, ruff, oxlint (lint); bun test, deno test, phpunit, rspec (test).
- New: Hook test harness (
tests/run.sh, 31 assertions) and CI on Linux and macOS. - New: Auto-handoff reads the RED-zone sentinel JSON, recording real token and cache numbers in TASK.md.
- Fix: Portability - replaced GNU-only
\swith[[:space:]], so the budget hook and lint filter no longer silently fail on stock BSD grep (macOS). - Fix: Validate
CONTEXT_ENGINEER_BUDGET- an empty, non-numeric, or zero override falls back to the default instead of crashing the hook. - Fix: Corrected the filter threshold docs to 30 lines (overridable via
CONTEXT_ENGINEER_FILTER_MIN_LINES). - Refactor: Scorecard logic deduplicated into one source of truth (
scripts/scorecard.sh);/context-engineer:diagnoseruns it instead of embedding a copy.
- Refactor: The three output filters (test, build, lint) now share a single sourced scaffold (
hooks/scripts/filter-common.sh) instead of duplicating it three times. Behavior is unchanged and verified byte-for-byte. - New: Filter line-count threshold is centralized and overridable via
CONTEXT_ENGINEER_FILTER_MIN_LINES(default 30). - Fix: Removed em dashes from the budget hook's zone messages to match the repo style rule.
- Docs: Refreshed the scorecard preview image version to v2.0.0.
- Docs: Added an Updating section with the marketplace refresh and reinstall flow, plus auto-update guidance.
- Docs: Documented how to invoke skills on demand, including the background Prompt Caching and Model Switching skills.
- Docs: Noted the v1.2.0 session-cost line in the diagnose command reference.
- New: Prompt Caching skill - cache-friendly workflow guidance (the biggest token-cost lever; cache reads cost ~0.1x)
- New: Cache-hit reporting - the budget warning hook now surfaces
% cachedin zone warnings and the RED handoff sentinel - New: Session cost estimate -
/context-engineer:diagnoseand the scorecard show current context size, cache-hit rate, and estimated cost at current Opus 4.8 rates - Update: Context window budget raised from 200K to ~1M to match current Claude Code (Opus 4.8 / Sonnet 4.6); now overridable via
CONTEXT_ENGINEER_BUDGET - Update: Corrected model pricing - Haiku $1/$5, Sonnet $3/$15, Opus $5/$25 per 1M (Opus is ~5x Haiku, was listed as 25x)
- Update: Rescaled heuristic fallback constants in the budget hook for the larger window
- Fix: Cap
FILE_SIZE_PCTat 100% - previously could exceed 100, distorting composite scores - Fix: Narrow compression detection to system markers only - no longer triggers on user content discussing "compressed" or "summarized" data
- Fix: Tighten message count grep to match JSON structure (
"role":) instead of any"role"string - Fix: Guard against empty transcript files in budget warning hook
- Fix: Atomic sentinel file creation to prevent race conditions on concurrent hook invocations
- Fix: Add
jqavailability check to all hook scripts - graceful exit instead of confusing errors - Fix: Add
jqerror handling in filter scripts for malformed JSON input
- Fix: Plugin root discovery for marketplace cache installs
- Fix: Windows compatibility note in README
- New: Auto-pilot handoff system - context budget warning hook, auto-handoff skill, enhanced fresh-context templates
- New: Composite scoring with 4 signals (file size, message count, compression, tool density)
- New: Sentinel file on RED zone to trigger automatic handoff
- Initial release: budget zones, degradation detection, token-saving hooks, model switching, thinking control, code intelligence
Built by Silvester Divas. Based on 10 context engineering techniques for Claude Code.
MIT