Skip to content

fix: batch 216-223 — model-list cache, output channel, tool-call pairing, zero-part parsing, 429 retry, session headers - #224

Merged
ltmoerdani merged 12 commits into
mainfrom
fix/issues-216-223-batch
Sep 8, 2026
Merged

ltmoerdani merged 12 commits into
mainfrom
fix/issues-216-223-batch

Conversation

@ltmoerdani

Copy link
Copy Markdown
Owner

Summary

Batch of fixes for issues #216, #217, #218, #220, #221, #222, #223. Each commit carries closes #N, so merging this PR auto-closes all seven.

Changes

Issue Fix Commit
#222 GET /models served from cache (was refetching every UI poll); Refresh Models still forces a real fetch; Models registered logs only on change 9f0eba9
#220 One shared output channel instead of one per request (was leaking a channel every request) 5d2b525
#216 Strict 1:1 pairing of function_call/function_call_output; no more fabricated call_id ed53ecb
#217 Nested Responses event payloads parse; dedicated zero-part error with diagnostics 573f757
#221 Honor Retry-After on upstream 429 (single transparent retry, capped 30s) 8d8e9f0
#218 / #223 Triage (VS Code dropdown limitation / misattributed stack trace) 6c0f184
follow-up Suppress [diag-empty-response] false positive on healthy tool-call turns ac7a9a1
follow-up Regression suites for #216/#217 + header-capture tests b99f449, 25f0301
follow-up x-opencode-session now sent on all auxiliary gateway requests (models, inline completions, usage, test-connection) 418c4c4

Verification

  • 462/462 unit tests pass; full lint gate (editorconfig, eslint, markdown, prettier, shellcheck, typecheck, tests) passes.
  • Manual verification by maintainer: model-list cache, single output channel after 8 requests, 9 chained luna tool-call turns all 200, per-model thinking + Go usage tracking intact.
  • Version bumped 0.7.4 → 0.7.5; CHANGELOG + devlog + issue docs 94–99 updated.

Notes

VS Code polls provideLanguageModelChatInformation every few hundred ms;
ModelListFetcher.fetch() performed a live GET /models on every poll because
MODEL_LIST_CACHE_TTL_MS was only consulted on the failure path.

- consult the fresh cached snapshot first; stale snapshots fall through
  to the live fetch + retry path
- add ModelListFetcher.invalidate() and wire it into Refresh Models so a
  manual refresh still performs a real upstream fetch
- log the 'Models registered' summary only when its signature changes
prepareChatRequest() called vscode.window.createOutputChannel("OpenCode")
once per request; every call registers a new Output-tab entry, so long
sessions accumulated dozens of duplicate channels and leaked the old ones.

- drop the per-request channel from chatPrep; transports now receive the
  provider's lazy singleton channel (already disposed via subscriptions)
The Console Go upstream rejects the whole request with 'No tool output
found for function call <id>' when a function_call item has no matching
function_call_output — e.g. after history trimming dropped the tool
result, or when VS Code replays a tool message whose tool_call_id was
lost. The old converter even fabricated a 'tool-<timestamp>' call_id in
that case, which could never match its call.

- never fabricate a function_call_output.call_id; drop the item instead
- add pairResponsesFunctionCallItems() enforcing strict 1:1 pairing:
  orphaned outputs are dropped, calls without an output are dropped,
  duplicated outputs keep only the first
- wire the pairing pass into buildResponsesRequestBody + unit tests
…ors (closes #217)

gpt-5.6-luna streams could end with zero extractable parts while the
gateway billed completion tokens; VS Code then surfaced its
empty-response-loop guard. Verified against the published 0.7.4 VSIX:
neither error string exists in our build, so the loop guard is Copilot
Chat reacting to a zero-part stream from our provider.

- accept record-shaped payloads on response.output_text.delta
  (delta.text / delta.content / text.value) and nested reasoning deltas
  so new gateway payload shapes still produce parts
- throw a dedicated zero-part error that names the token count, event
  stats, and the [diag-sse-event-*] diagnostic path after retries are
  exhausted
Console Go upstream rate limits ('Upstream request failed:
[rate_limit_exceeded]') come from the model provider, not the gateway —
the extension cannot lift them, but it can stop turning them into hard
failures when the upstream names a wait.

- parse Retry-After (delta-seconds and HTTP-date) via parseRetryAfterMs
- on a 429 with a Retry-After within RATE_LIMIT_MAX_RETRY_AFTER_WAIT_MS
  (30s), wait once and transparently retry — nothing has been streamed
  at that point, so no content can be duplicated
- longer waits still surface the existing rate-limit error untouched
…loses #218, closes #223)

- docs/issues/94-98: full root-cause records for #222, #220, #216, #217, #221
- docs/issues/99: triage for #218 (VS Code dropdown limitation, use
  Configure Utility Models) and #223 (stack trace belongs to
  vizards.deepseek-v4-for-copilot; our requests always send
  x-opencode-session)
- devlog batch entry + CHANGELOG 0.7.5
…thy tool-call turns

[diag-empty-response] plus the raw [diag-sse-event-*] dump fired on every
tool-call-only Responses turn (gpt-5.6-luna): tool calls are flushed in the
transport's finally block AFTER the diagnostic runs, so extractedPartCount is
legitimately 0 there while the stream is healthy. Skip the dump when a real
finish_reason was extracted — genuine format mismatches (issue #217 class)
never produce one.
…ent shapes)

Automates the manual verification checklist against the exact event shapes
captured from a live gpt-5.6-luna /v1/responses session:

- a 40-tool-call-group session that triggers history trim must still yield
  strictly paired function_call/function_call_output items after the pairing
  pass (the gateway 400s on any orphan)
- the captured luna tool-call event sequence (response.created → reasoning
  items → function_call added/delta/done → response.completed) must extract
  into a complete tool call despite zero text parts
- output_text.delta must extract for flat, delta:{text}, and text:{value}
  payload shapes
OpenCode's Go docs (updated 2026-09-07) now require a stable
x-opencode-session on every request, and requests missing it may start
erroring from 2026-09-06. Only the main chat path sent the header; four
auxiliary call sites hit the gateway without it:

- GET /models (ModelListFetcher)
- inline completions POST /chat/completions (ChatCompletionEngine)
- GET /usage (fetchGoUsage via GoUsageTracker server sync)
- Manage Provider test-connection POST

These have no conversation, so they share one persisted per-installation
session id (auxiliarySessionId(), globalState-backed) instead of the
per-conversation hash the chat path uses.
Header-capture tests stubbing global fetch prove the session header is sent
by ModelListFetcher (GET /models), ChatCompletionEngine (inline completions),
and fetchGoUsage (GET /usage); plus persistence semantics of
auxiliarySessionId (stable per installation, distinct across installs).
Also teaches the shared vscode mock about extensions.getExtension so
getUserAgent resolves in tests.
…y download

CI failed on the Editorconfig lint step: editorconfig-checker@6.1.1 could not
download the 'ec-linux-amd64' binary on the Linux runner. 6.2.0 (2026-08-27)
adds 'support editorconfig-checker binary name for v4 with legacy ec fallback'
(#435), which fixes the binary-name resolution. Node >=20.11 requirement stays
compatible with the CI Node 20 matrix.
@ltmoerdani
ltmoerdani merged commit 7612e12 into main Sep 8, 2026
2 checks passed
@ltmoerdani
ltmoerdani deleted the fix/issues-216-223-batch branch September 23, 2026 14:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant