Skip to content

fix(google): keep the realtime tool result across a session restart - #7153

Closed
Darshak03 wants to merge 2 commits into
livekit:mainfrom
Darshak03:fix/gemini-realtime-tool-result-across-restart
Closed

Darshak03 wants to merge 2 commits into
livekit:mainfrom
Darshak03:fix/gemini-realtime-tool-result-across-restart

Conversation

@Darshak03

Copy link
Copy Markdown

update_tools() restarts the Gemini Live socket mid-turn, and two orderings dropped the function result owed to the model:

  • _mark_restart_needed() swapped _msg_ch for a fresh channel and discarded everything queued on it, including an unsent LiveClientToolResponse.
  • update_chat_ctx() returned early while _active_session was None — the window between the old socket closing and the new one connecting — leaving the result in _chat_ctx only. The reconnect replays that with exclude_function_call=True, which drops function_call_output too.

Gemini holds the turn open until every call it emitted is answered, so either path hangs the conversation for good.

Fix: carry queued tool responses over to the new channel on restart, and track the call ids this session emitted but has not answered (_pending_tool_call_ids, matching the existing sets in the Phonic and AWS realtime plugins) so a result produced while disconnected is queued for the reconnected session. Filtering by id keeps a handoff's historical outputs from being replayed as stray tool responses.

Sending a tool response on the reconnected socket is not new behavior — when the restart is marked before the result arrives, that is already what happens today. This only closes the two orderings where the response is dropped instead.

Three tests in tests/test_plugin_google_realtime.py; the two regression ones fail on main.

Fixes #6479

The second half of #6479 (generation-completion handling for newer Live preview models) is already resolved on main — generation_complete closes the output streams, _is_new_generation() gates stray model_turn frames, and the model-name check now only computes the mutable_chat_context capability once at construction. #6785 (answer every tool call) fixes the agent layer so every call produces exactly one output, but not its delivery across a socket restart, which is what this addresses. Supersedes #6481.

update_tools() restarts the Gemini Live socket mid-turn, and two orderings
dropped the function result owed to the model:

- _mark_restart_needed() swapped _msg_ch for a fresh channel and discarded
  everything queued on it, including an unsent LiveClientToolResponse.
- update_chat_ctx() returned early while _active_session was None (between
  the old socket closing and the new one connecting), leaving the result in
  _chat_ctx only; the reconnect replays that with exclude_function_call=True,
  which drops function_call_output too.

Gemini holds the turn open until every call it emitted is answered, so either
path hangs the conversation for good.

Carry queued tool responses over to the new channel on restart, and track the
call ids this session emitted but has not answered so a result produced while
disconnected is queued for the reconnected session. Filtering by id keeps a
handoff's historical outputs from being replayed as stray tool responses.

Fixes livekit#6479
@Darshak03
Darshak03 requested a review from a team as a code owner September 7, 2026 10:23
@CLAassistant

CLAassistant commented Sep 7, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

Comment on lines +555 to 570
carried: list[ClientEvents] = []
while not self._msg_ch.empty():
msg = self._msg_ch.recv_nowait()
if isinstance(msg, types.LiveClientToolResponse):
carried.append(msg)
elif (
not on_error
and isinstance(msg, types.LiveClientContent)
and msg.turn_complete is True
):
logger.warning(
"discarding client content for turn completion, may cause generate_reply timeout",
extra={"lk.pii.content": str(msg)},
)

self._msg_ch = utils.aio.Chan[ClientEvents]()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Connection errors discard queued input

When _send_task requests an error restart, this loop discards every queued non-tool event. Audio, video, chat content, and turn completion no longer reach the replacement session. User input disappears, and pending replies can time out.

Prompt for agents
_mark_restart_needed previously skipped channel replacement when on_error=True, so queued events remained available to the replacement session. The new unconditional drain drops every queued event except LiveClientToolResponse during send and receive error recovery. Preserve the prior retry semantics for non-tool messages on error while still guaranteeing tool responses survive deliberate restarts. Add tests for queued audio/content and turn-complete events across an on_error restart.
Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This one is based on a misreading of the original — no behavior changed for non-tool events.

On main, self._msg_ch = utils.aio.Chan[ClientEvents]() is at 12-space indent, inside if not self._session_should_close.is_set(): — not inside if not on_error: (16 spaces):

    def _mark_restart_needed(self, on_error: bool = False) -> None:
        if not self._session_should_close.is_set():
            self._session_should_close.set()
            # reset the msg_ch, do not send messages from previous session
            if not on_error:
                while not self._msg_ch.empty():
                    msg = self._msg_ch.recv_nowait()
                    if isinstance(msg, types.LiveClientContent) and msg.turn_complete is True:
                        logger.warning(...)

            self._msg_ch = utils.aio.Chan[ClientEvents]()

So the channel was replaced unconditionally, on both paths, and queued events were never available to the replacement session — on on_error they were abandoned with the old channel rather than drained, which is the same outcome minus the warning. on_error only gated the logging pass, and this PR keeps that gating exactly as it was.

Queued audio/content is still discarded on restart, as before. The only difference is that LiveClientToolResponse is now carried over, because an unanswered call blocks the turn permanently while a dropped audio chunk does not.

_send_task takes _session_lock after pulling a message off the channel, so a
restart marked in that gap left the response neither queued for the carry loop
nor on the wire. A send that fails with the socket lost it the same way.

Hand it back to the channel the next session drains instead of dropping it.

self._msg_ch = utils.aio.Chan[ClientEvents]()
for msg in carried:
self._msg_ch.send_nowait(msg)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, I think the diagnosis is right. The three recovery points here, in update_chat_ctx and in _send_task are all the same event: a tool response that missed the socket.

Rather than re-queueing from each of them, could they append to one _pending_tool_responses list, and _main_task send it in one place after the connection is up, right after the chat context is sent? That keeps the ordering in one spot and removes _requeue_tool_response. When _session_resumption_handle is None the new session has no open call, so drop the list with a warning instead of sending it.

@longcw

longcw commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

#6000 has since landed a different recovery that covers the same three cases without re-queueing. The plugin now skips the history replay on a resumed session and instead sends the diff between the chat context the resumption handle stands for and the current one. An output that arrives through update_chat_ctx while disconnected waits for the reconnect; an output queued in _msg_ch and dropped by the restart, or dequeued and mid-send when the socket died, is still marked unsent, so the handle snapshot leaves it out and the reconnect diff re-sends it as a tool response to the session that still holds the call open.

@longcw longcw closed this Sep 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

4 participants