Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -361,7 +361,8 @@
"sdk/guides/llm-image-input",
"sdk/guides/llm-error-handling",
"sdk/guides/llm-fallback",
"sdk/guides/llm-profile-store"
"sdk/guides/llm-profile-store",
"sdk/guides/llm-prompt-composition"
]
},
{
Expand Down
89 changes: 89 additions & 0 deletions sdk/guides/llm-prompt-composition.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,89 @@
---
title: Prompt Token Composition
description: Break down prompt tokens per LLM call into system prompt, tool schemas, conversation history, and the latest message.
---

## Overview

When enabled, every LLM call records a per-call decomposition of its prompt tokens, so you can see where the input budget goes on each step of an agent run:

- `system_prompt_tokens` - Tokens in system messages
- `tool_tokens` - Tokens in the tool schemas included with the call
- `history_tokens` - Tokens in conversation history (all non-system messages except the latest one)
- `latest_message_tokens` - Tokens in the latest observation or user message

Recording is **opt-in**: set `enable_prompt_composition=True` on the LLM. When off (the default), no tokenization pass runs and no records are appended. When on, the decomposition is computed on every call - both the Chat Completions and Responses API paths - and recorded into the LLM's [Metrics](/sdk/guides/metrics). On the Responses API path, the decomposition is computed on the finalized payload (instructions plus input items), so the record reflects what the provider received.

```python icon="python" focus={4}
llm = LLM(
model="anthropic/claude-sonnet-4-5-20250929",
api_key=SecretStr(os.getenv("LLM_API_KEY")),
enable_prompt_composition=True,
)
```

## Accessing the Composition

With the flag enabled, each call appends one `PromptComposition` record to `llm.metrics.prompt_compositions`. The most recent record is available as `llm.metrics.latest_prompt_composition` (`None` when the flag is off):

```python icon="python"
conversation.run()

composition = llm.metrics.latest_prompt_composition
assert composition is not None
print(f"System prompt: {composition.system_prompt_tokens}")
print(f"Tool schemas: {composition.tool_tokens}")
print(f"History: {composition.history_tokens}")
print(f"Latest message: {composition.latest_message_tokens}")
```

Each record carries the `response_id` of its call. Join records with the provider-reported usage in `llm.metrics.token_usages` by `response_id` rather than by position - the two lists can diverge when a composition is skipped or a response carries no usage:

```python icon="python"
usage_by_id = {u.response_id: u for u in llm.metrics.token_usages}

for composition in llm.metrics.prompt_compositions:
usage = usage_by_id.get(composition.response_id)
if usage is None:
continue
estimated = (
composition.system_prompt_tokens
+ composition.tool_tokens
+ composition.history_tokens
+ composition.latest_message_tokens
)
print(
f"{composition.response_id}: estimated {estimated}, "
f"provider reported {usage.prompt_tokens}"
)
```

Agent steps that send tools as native function-calling schemas have `tool_tokens > 0`. Auxiliary calls that pass no tools - for example the [context condenser](/sdk/guides/context-condenser) or title generation - are recorded with `tool_tokens == 0`.

<Note>
On models without native function calling, the SDK renders tool schemas into the prompt text instead of sending them as tool parameters. Those agent steps are recorded with `tool_tokens == 0` and the schema tokens appear in the message buckets instead.
</Note>

## Estimates vs Provider-Reported Usage

Composition counts are **client-side estimates**, computed with the model's tokenizer before the request is sent. The provider-reported `TokenUsage` remains the authoritative accounting:

- `is_estimate` is `True` on records produced by the client-side estimator.
- Each component is counted independently, so per-message framing overhead is included in every component and the components do not necessarily sum exactly to the provider-reported `prompt_tokens`.
- Tool schema counts follow litellm's `token_counter` serialization convention for tools, which can differ from the provider's wire-format tokenization.
- For models litellm has no tokenizer mapping for, counts use litellm's fallback tokenizer and may deviate more from the provider's counts.
- When token counting fails or is disabled (for example `litellm.disable_token_counter`), composition recording is skipped for that call; the call itself is unaffected.
- Buckets follow the wire, not the logical prompt: in subscription mode the system prompt is folded into the first user message before transport, so those tokens are counted in `history_tokens` or `latest_message_tokens` rather than `system_prompt_tokens` on that path.

<Tip>
Treat the composition as a breakdown of *where* prompt tokens go, and `token_usages` as the record of *how many* tokens the provider billed.
</Tip>

## Performance

The estimator's cost scales linearly with prompt size: roughly 10-20 ms per call on a typical agent payload, measured at ~31 ms for a ~100K-token prompt and ~61 ms for ~190K tokens (19 tools, gpt-4o tokenizer) - negligible next to network latency. Because the feature is opt-in, this cost is only paid when `enable_prompt_composition=True`.

## Next Steps

- **[Metrics Tracking](/sdk/guides/metrics)** - Token usage, costs, and latency metrics for your agents
- **[Context Condenser](/sdk/guides/context-condenser)** - How OpenHands keeps history within the context window
2 changes: 2 additions & 0 deletions sdk/guides/metrics.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,7 @@ The `llm.metrics` object is an instance of the [Metrics class](https://github.co
- `costs` - List of individual cost records per API call
- `token_usages` - List of detailed token usage records per API call
- `response_latencies` - List of response latency metrics per API call
- `prompt_compositions` - List of per-call [prompt token composition](/sdk/guides/llm-prompt-composition) estimates (system prompt, tool schemas, history, latest message); populated only when `enable_prompt_composition=True` on the LLM

<Tip>
For more details on the available metrics and methods, refer to the [source code](https://github.com/OpenHands/software-agent-sdk/blob/main/openhands-sdk/openhands/sdk/llm/utils/metrics.py).
Expand Down Expand Up @@ -409,5 +410,6 @@ for usage_id, metrics in conversation.conversation_stats.usage_to_metrics.items(

## Next Steps

- **[Prompt Token Composition](/sdk/guides/llm-prompt-composition)** - Break down prompt tokens per call into system prompt, tool schemas, history, and latest message
- **[Context Condenser](/sdk/guides/context-condenser)** - Learn about context management and how it uses separate LLMs
- **[LLM Routing](/sdk/guides/llm-routing)** - Optimize costs with smart routing between different models
Loading