Reuse a per-session context across chat exchanges - #215
Conversation
ba9cd2a to
59f035f
Compare
Image segments threw unsupportedFeature because the backend had no multimodal path, even though the prebuilt llama.cpp binaries ship the mtmd library and its helpers. Accept an mmprojPath at initialization and load the projector next to the model. When a projector is present, prompt formatting replaces each image segment with the mtmd media marker and collects payloads in order, then generation tokenizes the marker-annotated prompt with mtmd_tokenize and evaluates text and image chunks through mtmd_helper_eval_chunks before sampling continues from the resulting position. Both respond and streaming support images, and models without a projector keep rejecting image input. Adds live tests generating from an embedded test image through both paths.
Gemma 4's canonical chat template no longer contains the start_of_turn marker that llama_chat_apply_template keys its Gemma detection on, so formatting threw encodingFailed for every Gemma 4 GGUF. When template application fails and the embedded template carries the Gemma 4 turn syntax, render it directly: turns open with a turn marker and role, close with the reverse marker, the assistant role is named model, and generation opens a model turn. The BOS token is applied during tokenization, and thinking is opt-in in this format so no suppression is needed.
Every generation created a fresh llama_context and prefilled the full rendered conversation from token zero, so multi-turn chat cost grew with the square of the transcript and long conversations spent most of their time re-decoding history. Keep one context alive per session for plain chat generations. Each exchange tokenizes the rendered prompt, keeps the longest token prefix shared with the context's recorded state, removes diverged state with llama_memory_seq_rm, and decodes only the remainder. Backends that cannot rewind, such as recurrent models, fall back to clearing memory and decoding the full prompt, and appends need no rewind on any backend. The final prompt token is always re-decoded so sampling has fresh logits, generated tokens extend the recorded state as they decode, and any generation error discards the cached context. Structured generation, image prompts, and encoder models keep single-use contexts, and clearCachedContext lets consumers free the cached state under memory pressure. Adds a live test asserting prefix reuse on the second turn of a session.
59f035f to
e79f8f1
Compare
There was a problem hiding this comment.
🟡 Changes recommended
Decode failures can leave invalid cached state, concurrent diagnostics race, and caching does not persist per session as described.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Pull request overview
Adds reusable llama.cpp session contexts to reduce repeated prefill work, alongside stacked Gemma 4 and vision support.
Changes:
- Reuses cached token prefixes across chat turns.
- Adds multimodal projector and image generation support.
- Adds Gemma 4 prompt rendering and integration tests.
File summaries
| File | Description |
|---|---|
LlamaLanguageModel.swift |
Implements caching, multimodal generation, and Gemma 4 formatting. |
LlamaLanguageModelTests.swift |
Tests context reuse and vision behavior. |
LlamaGemma4TemplateTests.swift |
Tests Gemma 4 prompt rendering. |
Review details
Suppressed comments (1)
Sources/AnyLanguageModel/Models/LlamaLanguageModel.swift:1615
- A multimodal decode failure is treated as normal completion, so callers receive a successful truncated response and streaming finishes without an error. Propagate
decodingFailed, consistent with the prompt-evaluation failure above.
guard llama_decode(context, batch) == 0 else {
break
}
- Files reviewed: 3/3 changed files
- Comments generated: 3
- Review effort level: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| /// frees a context that is still decoding: it runs on a transient context | ||
| /// instead and leaves the cache untouched. | ||
| private let sessionContextLock = NSLock() | ||
| private var cachedSessionContext: CachedSessionContext? |
| lastReusedTokenCount = startIndex | ||
| lastPrefillTokenCount = promptTokens.count - startIndex |
| decodedTokens.append(nextToken) | ||
| } | ||
|
|
||
| recordCachedTokens(decodedTokens, context: context) |
|
Hi @james-333i. Two small things on this one, written up at the end of #213 (review) so you can do the whole stack in one pass: guard the |
Every generation created a fresh llama_context and prefilled the full conversation from token zero, so multi-turn cost grew with the square of the transcript. This keeps one context alive per session, reuses the longest shared token prefix, and decodes only the remainder. Backends that cannot rewind fall back to a full decode, and any generation error discards the cached context. Stacks on the Gemma 4 PR.