fix(google): skip context replay on resumed realtime sessions - #6000
Conversation
cbc3f01 to
d4cfc32
Compare
|
Rebased onto current main and kept the focused resume-context guard intact. Ruff and py_compile pass. The focused pytest collection is currently blocked locally by an older installed livekit.protocol package missing agent_simulation; I did not widen the branch to environment or dependency changes. |
|
The red The job then kept running until GitHub canceled the runner several hours later. All style/type checks and blockguard tests are green. I am leaving the branch unchanged unless you see a link between this IPC process-pool failure and the resumed realtime session context path. |
d4cfc32 to
bee2c4d
Compare
|
Updated The branch now invalidates the stored session resumption handle when Validation: I also rebased onto current |
|
Rebased onto current main. The resume-replay path is still there on main (realtime_api.py still seeds unconditionally at the if turns check), so this stays relevant. The conflict was the mixed-tools test class appended at the same spot, union-merged. tests/test_plugin_google_llm.py passes 26/26 locally. |
bee2c4d to
b85fff7
Compare
A chat context update that arrives while no session is active now waits in _pending_chat_ctx instead of clearing the resumption handle. On connect, a fresh session adopts it and replays the history as before; a resumed session queues only the diff against the known context through the same path the connected update uses, so a tool result produced during the restart reaches the session that still holds the call open. The resumable=false clause is dropped: the 2.5 native-audio model sends resumable=None with an empty handle after every connect and every client message, which left the session without a handle until the next model turn.
longcw
left a comment
There was a problem hiding this comment.
thanks for the pr! I merged the main and made some changes to it.
_sync_chat_ctx advances _chat_ctx when it queues an item, but the handle covering that item arrives only after the next model turn, so a restart in between left the item outside the server state with nothing to resend it. Snapshot _chat_ctx whenever a handle arrives and diff the resumed connect against that snapshot instead. A handle passed through RealtimeModel has no snapshot; adopt the first chat ctx as the server's rather than replaying it onto the resumed session.
_sync_chat_ctx advances _chat_ctx when it queues, so a handle that the receive task processes before the send task has drained the queue recorded those items as covered, and a restart in between dropped them with the channel. Build the chat ctx events as subclasses that carry their item ids, track the ids as unsent until the send task has sent the event, and leave still-unsent items out of the snapshot taken on a new handle so the resume diff re-sends them.
| self._resumption_chat_ctx = llm.ChatContext( | ||
| [ | ||
| item | ||
| for item in self._chat_ctx.items | ||
| if item.id not in self._unsent_item_ids | ||
| ] | ||
| ) |
There was a problem hiding this comment.
🔴 Queued updates replay twice after failures
When one context send fails with later events queued, _resumption_chat_ctx excludes every unsent item. The reconnect requeues them while _mark_restart_needed preserves their original events. Gemini receives duplicate messages or tool results after recovery.
Prompt for agents
Track failed in-flight context events separately from context events still waiting in `_msg_ch`. On an error restart, the reconnect diff must restore the failed event without duplicating queued events. One approach is to drain/reset tagged chat-context events before rebuilding the diff, while preserving the existing retry semantics for unrelated realtime input. Add a test with at least two queued chat-context updates where the first send raises and verify each message or tool result reaches the replacement socket exactly once.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
an error restart discards the queue as well, and the resume diff is the only thing that re-sends.
Re-vendors upstream livekit-plugins-google 1.8.1 (from 1.8.0) and re-applies the fork's patches on top. Upstream 1.8.1 is 1.8.0 plus exactly one change: livekit/agents#6000, "skip context replay on resumed realtime sessions" (realtime_api.py only, +95/-32). llm.py, stt.py, tts.py, utils.py and beta/ are byte-identical to 1.8.0. pyproject bumps its own floor to livekit-agents>=1.8.1, so the fork pin moves to >=1.8.1,<1.9. One fork patch is dropped because upstream fixed it: the stash-and-replay of in-flight tool results across an update_tools restart (_pending_tool_result). #6000 covers the same failure differently -- the plugin reconnects with the last session resumption handle, skips the history replay, and sends only the diff against the context that handle stands for. A result dropped by the restart is still marked unsent, so the handle snapshot leaves it out and the reconnect diff re-sends it to the session that still holds the call open. Keeping our stash on top would deliver the result twice. Patches carried forward: - preserve pre-filled tool params (keep JSON-schema `default`; send parameters_json_schema on the realtime path) - gemini-3.1 Live: allow generate_reply via LiveClientRealtimeInput (with the NOT_GIVEN -> "." fallback), and drop model_turn parts arriving after a turn is finalized - emit recoverable connection errors after the retry backoff, not before scripts/verify.sh: drops the _pending_tool_result check with the patch and greps the NOT_GIVEN fallback for the 3.1 check. Verified against a real install of livekit-agents 1.8.1: all static and runtime checks pass. Also adds a .gitignore and untracks a committed __pycache__ .pyc.
fix(google): skip the chat context replay on a resumed realtime session
Problem: Every reconnect replays the local chat context with
send_client_content, but the plugin always reconnects with the last session resumption handle, and the resumed session already holds the conversation. The replay duplicates the history on every restart, including theupdate_toolsrestart that follows every tool call.Fix: When a resumption handle is set, the connect path skips the replay and sends only the diff between the chat context the handle stands for and the current one, tool results included. A snapshot of the context is taken each time a handle arrives, and an update that arrives while disconnected waits in
_pending_chat_ctxuntil the next connect.Fixes #5985. Lands before #7153, which relies on the deferred diff for the tool result path.
Context for reviewing and coding agents