Your question
Environment
@livekit/agents 1.6.4, @livekit/agents-plugin-google 1.6.4
Model: gemini-3.1-flash-live-preview
Node 24, Ubuntu 25.04
Voice agent over SIP, Dutch (nl-BE), audio in and out
What happens
During a normal conversation the agent occasionally restarts as if the call
had just begun: it greets the user again, mid-sentence, and invents details
it cannot know. The conversation history is gone from the model's point of
view, while the WebSocket itself never reported an error to the application.
Concrete example from production (25 Aug 2026, 19:04 UTC). The user was
talking about the market in Hasselt. The agent replied:
"Good evening Mario. It's Tuesday evening, ten past half six, what are you
up to?"
The user had to correct it. The time was also wrong , the prompt supplied
21:04, the model produced 18:40.
What the token metrics show
We record per-turn token usage. The audio context collapses while the text
context stays roughly intact:
time text_in audio_in
19:04:01 17707 2016
19:04:17 17754 2367
19:04:36 17816 2758
19:04:52 17210 351
In three days of production traffic this happened twice across 68 turns, in
two different conversations. In one case a Google Search grounding call was
in flight in the same second; in the other there was no search at all.
Where it appears to come from
On reconnect, realtime_api sends the session's own chat context back:
const [turns] = await this._chatCtx.copy({ excludeFunctionCall: true })
.toProviderFormat("google", false);
toProviderFormat routes through llm/provider_format/google.js, which
starts with:
var toGoogleParts = (content) => {
const parts = [];
for (const c of content) {
if (typeof c === "string") {
parts.push({ text: c });
}
}
return parts;
};
Only strings survive. Audio content is silently dropped. For a native-audio
model, restoring a voice conversation as text-only appears to leave it
without a recognisable conversation, and it opens a new one.
_chatCtx itself is populated correctly - we verified 129 output
transcriptions against 9 input transcriptions in that call, and our own
stored transcript has all 20 turns from both sides. The loss happens in the
conversion...
Possibly related
#3386 - update_chat_ctx() not propagating to Gemini realtime
#6002 - mentions "reconnect/session-resumption complexity" as a tradeoff
Question
Is restoring a native-audio session as text-only the intended behaviour? If
so, is there a supported way to detect that a resumption has occurred, so
the application can end the call gracefully rather than let the model start
over? Right now the only signal we have is the token counts.
Your question
Environment
@livekit/agents 1.6.4, @livekit/agents-plugin-google 1.6.4
Model: gemini-3.1-flash-live-preview
Node 24, Ubuntu 25.04
Voice agent over SIP, Dutch (nl-BE), audio in and out
What happens
During a normal conversation the agent occasionally restarts as if the call
had just begun: it greets the user again, mid-sentence, and invents details
it cannot know. The conversation history is gone from the model's point of
view, while the WebSocket itself never reported an error to the application.
Concrete example from production (25 Aug 2026, 19:04 UTC). The user was
talking about the market in Hasselt. The agent replied:
"Good evening Mario. It's Tuesday evening, ten past half six, what are you
up to?"
The user had to correct it. The time was also wrong , the prompt supplied
21:04, the model produced 18:40.
What the token metrics show
We record per-turn token usage. The audio context collapses while the text
context stays roughly intact:
time text_in audio_in
19:04:01 17707 2016
19:04:17 17754 2367
19:04:36 17816 2758
19:04:52 17210 351
In three days of production traffic this happened twice across 68 turns, in
two different conversations. In one case a Google Search grounding call was
in flight in the same second; in the other there was no search at all.
Where it appears to come from
On reconnect, realtime_api sends the session's own chat context back:
const [turns] = await this._chatCtx.copy({ excludeFunctionCall: true })
.toProviderFormat("google", false);
toProviderFormat routes through llm/provider_format/google.js, which
starts with:
var toGoogleParts = (content) => {
const parts = [];
for (const c of content) {
if (typeof c === "string") {
parts.push({ text: c });
}
}
return parts;
};
Only strings survive. Audio content is silently dropped. For a native-audio
model, restoring a voice conversation as text-only appears to leave it
without a recognisable conversation, and it opens a new one.
_chatCtx itself is populated correctly - we verified 129 output
transcriptions against 9 input transcriptions in that call, and our own
stored transcript has all 20 turns from both sides. The loss happens in the
conversion...
Possibly related
#3386 - update_chat_ctx() not propagating to Gemini realtime
#6002 - mentions "reconnect/session-resumption complexity" as a tradeoff
Question
Is restoring a native-audio session as text-only the intended behaviour? If
so, is there a supported way to detect that a resumption has occurred, so
the application can end the call gracefully rather than let the model start
over? Right now the only signal we have is the token counts.