Skip to content

fix(compiler): retry transient LLM API timeouts with bounded backoff #229

Description

@sebastianbraun25

Problem

During batch document ingest (openkb add with multiple documents), the LLM compilation pipeline sometimes fails to generate complete concept/entity pages due to transient Gateway Timeouts from the Anthropic API.

Example output:

[147/148] Adding: jira-SSMPA-9592.md
  [WARN] 8 concept(s) planned but only 6 written for jira-SSMPA-9592 (Timeout).
  [WARN] 6 entity(ies) planned but only 5 written for jira-SSMPA-9592 (Timeout).

This results in an incomplete knowledge base where specific concepts/entities are missing entirely.

Reproduktion

# Run batch ingest with multiple documents
openkb add ~/documents/batch/*.md

# Observe logs: After ~150 seconds of concurrent API calls (concurrency=5),
# last few concept/entity generations fail with:
# litellm.Timeout: AnthropicException Timeout - Gateway Timeout

Occurs when:

  • Document count is high (multiple parallel compilations)
  • Anthropic API is under load
  • Concurrency=5 sustained for ~150+ seconds

Workaround: Re-run the ingest command (transient nature means retry would succeed).

Kontext

  • Python 3.12
  • openkb add with high-concurrency compilation
  • Anthropic Claude API timeout (503 Gateway Timeout)
  • Not a client-side timeout — server-side load-induced

Root Cause

High-concurrency API requests (5 parallel via semaphore) over extended duration cause temporary gateway overload on Anthropic's infrastructure. Individual API calls are successful; only sustained parallel load triggers 503 responses on the N-th request.

Proposed Solution

Implement exponential backoff retry for transient errors (litellm.Timeout, 5xx API errors, rate limits) in the LLM call pipeline. Only retry errors that are stateless and temporary; skip permanent errors (4xx, validation, auth).

  • Use LiteLLM's built-in retries parameter (exponential backoff by default)
  • Filter exception types: only retry transient errors
  • Bound retries to 2 attempts + 4s total wait
  • Better logging to distinguish transient vs. permanent failures

Activity

sebastianbraun25 commented on Aug 27, 2026

@sebastianbraun25
Author

Note: the originally proposed solution referenced LiteLLM's 'retries' parameter — this was incorrect. LiteLLM only recognizes 'num_retries'/'max_retries' as internal retry-control kwargs; a bare 'retries' kwarg is passed through unrecognized to the provider request body, which strict-mode proxies reject (400 'Extra inputs are not permitted'). The implementation in PR #230 has been corrected to use 'num_retries'.

sebastianbraun25 commented on Aug 28, 2026

@sebastianbraun25
Author

Withdrawing this issue for now (see closing comment on PR #230 for the full
reasoning).

Summary

Root-cause investigation in downstream production use points to these
Gateway Timeout failures correlating with response duration/size — the pages
that fail do so deterministically on every attempt (including after a
retry), and are specifically the largest existing concept/entity pages in the
KB. This is consistent with a fixed-duration timeout enforced by an
intermediary (a corporate/self-hosted LLM gateway or proxy in front of the
provider, common in enterprise deployments of OpenKB), not transient upstream
load — a client-side retry policy does not fix that, and simply retrying the
exact same slow request reproduces the exact same timeout.

Closing for now; may reopen once there's a fix that addresses the actual
root cause (e.g. streaming responses for large page rewrites, or bounding how
large a single page update can grow) rather than only retrying.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions