Skip to content

fix(google): skip context replay on resumed realtime sessions - #6000

Merged
longcw merged 6 commits into
livekit:mainfrom
he-yufeng:fix/google-realtime-resume-context
Sep 9, 2026
Merged

longcw merged 6 commits into
livekit:mainfrom
he-yufeng:fix/google-realtime-resume-context

Conversation

@he-yufeng

@he-yufeng he-yufeng commented Jun 7, 2026 •

Copy link
Copy Markdown
Contributor

fix(google): skip the chat context replay on a resumed realtime session

Problem: Every reconnect replays the local chat context with send_client_content, but the plugin always reconnects with the last session resumption handle, and the resumed session already holds the conversation. The replay duplicates the history on every restart, including the update_tools restart that follows every tool call.

Fix: When a resumption handle is set, the connect path skips the replay and sends only the diff between the chat context the handle stands for and the current one, tool results included. A snapshot of the context is taken each time a handle arrives, and an update that arrives while disconnected waits in _pending_chat_ctx until the next connect.

Fixes #5985. Lands before #7153, which relies on the deferred diff for the tool result path.

Context for reviewing and coding agents

What a resumed session restores, measured against the live API

A probe on gemini-2.5-flash-native-audio-preview-12-2025 and gemini-3.1-flash-live-preview sent two facts, closed the socket, and reconnected with the latest handle: the model reported both. Reconnecting with the handle issued before the second fact, on 3.1, reported only the first, so a handle is a snapshot of the server state at the time it was issued. Updates arrive right after each client message and each model turn. SessionResumptionConfig.transparent, which would add last_consumed_client_message_index, is rejected on the Gemini Developer API.

How to see it

Spy on google.genai.live.AsyncSession.send_client_content, run one turn through RealtimeSession, call update_tools([...]), and read what the second connection sends. On main it sends the user turn and the model turn again; with this change it sends nothing, and the model still answers a question about the first turn. tests/test_plugin_google_realtime.py::test_resumed_session_skips_chat_ctx_replay covers the decision with a fake AsyncLive.connect.

Why the disconnected update is a diff and not a fresh session

The earlier revision cleared the handle when update_chat_ctx ran while disconnected, so the next connect replayed everything as a fresh session. That path is the one every tool result takes during the update_tools restart, and the fresh replay uses exclude_function_call=True, so the call and its output never reached the model. Sending the diff through the same compute_chat_ctx_diff as the connected path delivers the function_call_output as a tool response to the resumed session, which still holds the call open.

Why the diff runs against a snapshot and not against _chat_ctx

_sync_chat_ctx advances _chat_ctx when it queues an item, and the handle covering that item arrives only after the next model turn on 2.5. A restart in between leaves the item outside the server state while _chat_ctx already holds it, so a diff against _chat_ctx would send nothing and the message or tool result would be lost. _resumption_chat_ctx is copied from _chat_ctx when each handle arrives, minus the items whose events are still queued in _msg_ch: _sync_chat_ctx builds its events as _ChatCtxContent and _ChatCtxToolResponse, subclasses that carry the item ids, and the send task subtracts those ids from _unsent_item_ids once the send returns, so a handle that lands before the send task has run does not claim them. The resumed connect diffs the pending or current context against that snapshot. A live run that sent a message and called update_tools at once showed the message dropped by the channel drain on the first socket and re-sent on the second.

Why a caller-provided handle adopts the first context without sending it

With RealtimeModel(session_resumption=SessionResumptionConfig(handle=...)) the baseline is unknown: _chat_ctx starts empty and a diff would replay the whole persisted history onto a session that already holds it. _resumption_chat_ctx stays None until the first handle arrives, and in that state the resumed connect adopts the pending context as the server's and sends nothing.

Why the handle is not cleared on resumable=false

The 2.5 native-audio model sends a session_resumption_update with resumable=None and an empty handle after every connect and every client message. Clearing on not resumable left the session without a handle until the next model turn, so a restart in that window became a fresh session.

What is not covered

The probe could not reproduce the 1007 from #5985: a four-turn replay on a resumed 3.1 session went through, so that error has another trigger.

devin-ai-integration[bot]

This comment was marked as resolved.

@he-yufeng
he-yufeng force-pushed the fix/google-realtime-resume-context branch from cbc3f01 to d4cfc32 Compare June 9, 2026 10:45
@he-yufeng

Copy link
Copy Markdown
Contributor Author

Rebased onto current main and kept the focused resume-context guard intact. Ruff and py_compile pass. The focused pytest collection is currently blocked locally by an older installed livekit.protocol package missing agent_simulation; I did not widen the branch to environment or dependency changes.

@he-yufeng

Copy link
Copy Markdown
Contributor Author

The red unit-tests job does not appear to be from this Google realtime resume-context change. The failure is in tests/test_ipc.py::test_slow_initialization, where the test leak detector finds a pending ProcJobExecutor._supervise_task after process initialization exits with code -10:

ERROR tests/test_ipc.py::test_slow_initialization - Failed: Test leaked tasks
Coroutine: ProcJobExecutor._supervise_task

The job then kept running until GitHub canceled the runner several hours later. All style/type checks and blockguard tests are green. I am leaving the branch unchanged unless you see a link between this IPC process-pool failure and the resumed realtime session context path.

@he-yufeng
he-yufeng force-pushed the fix/google-realtime-resume-context branch from d4cfc32 to bee2c4d Compare June 13, 2026 19:39
@he-yufeng
he-yufeng requested a review from a team as a code owner June 13, 2026 19:39
@he-yufeng

Copy link
Copy Markdown
Contributor Author

Updated bee2c4d9e to handle the disconnected-update case from the review.

The branch now invalidates the stored session resumption handle when update_chat_ctx() receives a local chat-context update while no realtime session is active. That makes the next reconnect start from the current local context instead of trying to resume from a server handle that points at an older transcript. It also clears the handle when the server reports a non-resumable session update.

Validation:

python -m py_compile livekit-plugins/livekit-plugins-google/livekit/plugins/google/realtime/realtime_api.py tests/test_plugin_google_llm.py
.\.venv\Scripts\python.exe -m pytest tests/test_plugin_google_llm.py -q
# 18 passed
python -m ruff check livekit-plugins/livekit-plugins-google/livekit/plugins/google/realtime/realtime_api.py tests/test_plugin_google_llm.py
python -m ruff format --check livekit-plugins/livekit-plugins-google/livekit/plugins/google/realtime/realtime_api.py tests/test_plugin_google_llm.py
git diff --check origin/main..HEAD

I also rebased onto current origin/main before pushing.

devin-ai-integration[bot]

This comment was marked as resolved.

@he-yufeng

Copy link
Copy Markdown
Contributor Author

Rebased onto current main. The resume-replay path is still there on main (realtime_api.py still seeds unconditionally at the if turns check), so this stays relevant. The conflict was the mixed-tools test class appended at the same spot, union-merged. tests/test_plugin_google_llm.py passes 26/26 locally.

@he-yufeng
he-yufeng force-pushed the fix/google-realtime-resume-context branch from bee2c4d to b85fff7 Compare August 19, 2026 00:38
devin-ai-integration[bot]

This comment was marked as resolved.

A chat context update that arrives while no session is active now
waits in _pending_chat_ctx instead of clearing the resumption handle.
On connect, a fresh session adopts it and replays the history as
before; a resumed session queues only the diff against the known
context through the same path the connected update uses, so a tool
result produced during the restart reaches the session that still
holds the call open.

The resumable=false clause is dropped: the 2.5 native-audio model
sends resumable=None with an empty handle after every connect and
every client message, which left the session without a handle until
the next model turn.
devin-ai-integration[bot]

This comment was marked as resolved.

@longcw longcw left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for the pr! I merged the main and made some changes to it.

_sync_chat_ctx advances _chat_ctx when it queues an item, but the
handle covering that item arrives only after the next model turn, so a
restart in between left the item outside the server state with nothing
to resend it. Snapshot _chat_ctx whenever a handle arrives and diff the
resumed connect against that snapshot instead.

A handle passed through RealtimeModel has no snapshot; adopt the first
chat ctx as the server's rather than replaying it onto the resumed
session.
devin-ai-integration[bot]

This comment was marked as resolved.

_sync_chat_ctx advances _chat_ctx when it queues, so a handle that the
receive task processes before the send task has drained the queue
recorded those items as covered, and a restart in between dropped them
with the channel. Build the chat ctx events as subclasses that carry
their item ids, track the ids as unsent until the send task has sent
the event, and leave still-unsent items out of the snapshot taken on a
new handle so the resume diff re-sends them.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 new potential issue.

Devin Review

Comment on lines +1213 to +1219
self._resumption_chat_ctx = llm.ChatContext(
[
item
for item in self._chat_ctx.items
if item.id not in self._unsent_item_ids
]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Queued updates replay twice after failures

When one context send fails with later events queued, _resumption_chat_ctx excludes every unsent item. The reconnect requeues them while _mark_restart_needed preserves their original events. Gemini receives duplicate messages or tool results after recovery.

Prompt for agents
Track failed in-flight context events separately from context events still waiting in `_msg_ch`. On an error restart, the reconnect diff must restore the failed event without duplicating queued events. One approach is to drain/reset tagged chat-context events before rebuilding the diff, while preserving the existing retry semantics for unrelated realtime input. Add a test with at least two queued chat-context updates where the first send raises and verify each message or tool result reaches the replacement socket exactly once.
Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

an error restart discards the queue as well, and the resume diff is the only thing that re-sends.

@longcw
longcw merged commit eb2e96f into livekit:main Sep 9, 2026
15 checks passed
Darshak03 pushed a commit to Darshak03/livekit-google-plugin that referenced this pull request Sep 14, 2026
Re-vendors upstream livekit-plugins-google 1.8.1 (from 1.8.0) and re-applies
the fork's patches on top.

Upstream 1.8.1 is 1.8.0 plus exactly one change: livekit/agents#6000, "skip
context replay on resumed realtime sessions" (realtime_api.py only, +95/-32).
llm.py, stt.py, tts.py, utils.py and beta/ are byte-identical to 1.8.0.
pyproject bumps its own floor to livekit-agents>=1.8.1, so the fork pin moves
to >=1.8.1,<1.9.

One fork patch is dropped because upstream fixed it: the stash-and-replay of
in-flight tool results across an update_tools restart (_pending_tool_result).
#6000 covers the same failure differently -- the plugin reconnects with the
last session resumption handle, skips the history replay, and sends only the
diff against the context that handle stands for. A result dropped by the
restart is still marked unsent, so the handle snapshot leaves it out and the
reconnect diff re-sends it to the session that still holds the call open.
Keeping our stash on top would deliver the result twice.

Patches carried forward:
- preserve pre-filled tool params (keep JSON-schema `default`; send
  parameters_json_schema on the realtime path)
- gemini-3.1 Live: allow generate_reply via LiveClientRealtimeInput (with the
  NOT_GIVEN -> "." fallback), and drop model_turn parts arriving after a turn
  is finalized
- emit recoverable connection errors after the retry backoff, not before

scripts/verify.sh: drops the _pending_tool_result check with the patch and
greps the NOT_GIVEN fallback for the 3.1 check. Verified against a real
install of livekit-agents 1.8.1: all static and runtime checks pass.

Also adds a .gitignore and untracks a committed __pycache__ .pyc.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants