Conversation
update_tools() restarts the Gemini Live socket mid-turn, and two orderings dropped the function result owed to the model: - _mark_restart_needed() swapped _msg_ch for a fresh channel and discarded everything queued on it, including an unsent LiveClientToolResponse. - update_chat_ctx() returned early while _active_session was None (between the old socket closing and the new one connecting), leaving the result in _chat_ctx only; the reconnect replays that with exclude_function_call=True, which drops function_call_output too. Gemini holds the turn open until every call it emitted is answered, so either path hangs the conversation for good. Carry queued tool responses over to the new channel on restart, and track the call ids this session emitted but has not answered so a result produced while disconnected is queued for the reconnected session. Filtering by id keeps a handoff's historical outputs from being replayed as stray tool responses. Fixes livekit#6479
There was a problem hiding this comment.
Devin Review found 2 potential issues.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| carried: list[ClientEvents] = [] | ||
| while not self._msg_ch.empty(): | ||
| msg = self._msg_ch.recv_nowait() | ||
| if isinstance(msg, types.LiveClientToolResponse): | ||
| carried.append(msg) | ||
| elif ( | ||
| not on_error | ||
| and isinstance(msg, types.LiveClientContent) | ||
| and msg.turn_complete is True | ||
| ): | ||
| logger.warning( | ||
| "discarding client content for turn completion, may cause generate_reply timeout", | ||
| extra={"lk.pii.content": str(msg)}, | ||
| ) | ||
|
|
||
| self._msg_ch = utils.aio.Chan[ClientEvents]() |
There was a problem hiding this comment.
🔴 Connection errors discard queued input
When _send_task requests an error restart, this loop discards every queued non-tool event. Audio, video, chat content, and turn completion no longer reach the replacement session. User input disappears, and pending replies can time out.
Prompt for agents
_mark_restart_needed previously skipped channel replacement when on_error=True, so queued events remained available to the replacement session. The new unconditional drain drops every queued event except LiveClientToolResponse during send and receive error recovery. Preserve the prior retry semantics for non-tool messages on error while still guaranteeing tool responses survive deliberate restarts. Add tests for queued audio/content and turn-complete events across an on_error restart.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
This one is based on a misreading of the original — no behavior changed for non-tool events.
On main, self._msg_ch = utils.aio.Chan[ClientEvents]() is at 12-space indent, inside if not self._session_should_close.is_set(): — not inside if not on_error: (16 spaces):
def _mark_restart_needed(self, on_error: bool = False) -> None:
if not self._session_should_close.is_set():
self._session_should_close.set()
# reset the msg_ch, do not send messages from previous session
if not on_error:
while not self._msg_ch.empty():
msg = self._msg_ch.recv_nowait()
if isinstance(msg, types.LiveClientContent) and msg.turn_complete is True:
logger.warning(...)
self._msg_ch = utils.aio.Chan[ClientEvents]()So the channel was replaced unconditionally, on both paths, and queued events were never available to the replacement session — on on_error they were abandoned with the old channel rather than drained, which is the same outcome minus the warning. on_error only gated the logging pass, and this PR keeps that gating exactly as it was.
Queued audio/content is still discarded on restart, as before. The only difference is that LiveClientToolResponse is now carried over, because an unanswered call blocks the turn permanently while a dropped audio chunk does not.
_send_task takes _session_lock after pulling a message off the channel, so a restart marked in that gap left the response neither queued for the carry loop nor on the wire. A send that fails with the socket lost it the same way. Hand it back to the channel the next session drains instead of dropping it.
|
|
||
| self._msg_ch = utils.aio.Chan[ClientEvents]() | ||
| for msg in carried: | ||
| self._msg_ch.send_nowait(msg) |
There was a problem hiding this comment.
Thanks, I think the diagnosis is right. The three recovery points here, in update_chat_ctx and in _send_task are all the same event: a tool response that missed the socket.
Rather than re-queueing from each of them, could they append to one _pending_tool_responses list, and _main_task send it in one place after the connection is up, right after the chat context is sent? That keeps the ordering in one spot and removes _requeue_tool_response. When _session_resumption_handle is None the new session has no open call, so drop the list with a warning instead of sending it.
|
#6000 has since landed a different recovery that covers the same three cases without re-queueing. The plugin now skips the history replay on a resumed session and instead sends the diff between the chat context the resumption handle stands for and the current one. An output that arrives through |
update_tools()restarts the Gemini Live socket mid-turn, and two orderings dropped the function result owed to the model:_mark_restart_needed()swapped_msg_chfor a fresh channel and discarded everything queued on it, including an unsentLiveClientToolResponse.update_chat_ctx()returned early while_active_sessionwasNone— the window between the old socket closing and the new one connecting — leaving the result in_chat_ctxonly. The reconnect replays that withexclude_function_call=True, which dropsfunction_call_outputtoo.Gemini holds the turn open until every call it emitted is answered, so either path hangs the conversation for good.
Fix: carry queued tool responses over to the new channel on restart, and track the call ids this session emitted but has not answered (
_pending_tool_call_ids, matching the existing sets in the Phonic and AWS realtime plugins) so a result produced while disconnected is queued for the reconnected session. Filtering by id keeps a handoff's historical outputs from being replayed as stray tool responses.Sending a tool response on the reconnected socket is not new behavior — when the restart is marked before the result arrives, that is already what happens today. This only closes the two orderings where the response is dropped instead.
Three tests in
tests/test_plugin_google_realtime.py; the two regression ones fail onmain.Fixes #6479
The second half of #6479 (generation-completion handling for newer Live preview models) is already resolved on
main—generation_completecloses the output streams,_is_new_generation()gates straymodel_turnframes, and the model-name check now only computes themutable_chat_contextcapability once at construction. #6785 (answer every tool call) fixes the agent layer so every call produces exactly one output, but not its delivery across a socket restart, which is what this addresses. Supersedes #6481.