Repository navigation
fix(voice): include tool execution context in replies - #1802
rosetta-livekit-bot[bot] wants to merge 20 commits into
Conversation
…1525) Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com> Co-authored-by: u9g <jason.lernerman@livekit.io>
Agent.llmNode now returns ReadableStream<ChatChunk | string | FlushSentinel>, but the agent_v2 hook overrides and AgentHookAdapter still declared the narrower ChatChunk | string union, so passing super.llmNode as the fallback failed to type-check. Widen the override return types and the adapter's fallback/return signatures to include FlushSentinel. Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Brian Yin <brian.yin@livekit.io> Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com> Co-authored-by: u9g <jason.lernerman@livekit.io>
Catch end-call close listener errors to avoid unhandled rejections during shutdown, and make public tool type guards return false for null inputs.
Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com>
Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com>
Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com>
…egment (#1760) Co-authored-by: Cursor <cursoragent@cursor.com>
…#1698) Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com> Co-authored-by: Toubat <brian.yin@livekit.io>
# Conflicts: # agents/src/voice/agent_activity.test.ts # agents/src/voice/generation.ts # agents/src/voice/generation_tts_timeout.test.ts
The test asserted an exact tick count ([0,1,2]) against real timers with a 5ms margin, which flakes on loaded CI runners. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com>
Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com>
Co-authored-by: rosetta-livekit-bot[bot] <282703043+rosetta-livekit-bot[bot]@users.noreply.github.com>
🦋 Changeset detectedLatest commit: 5c5a574 The changes in this PR will be included in the next version bump. This PR includes changesets to release 35 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
| chatCtx.insert(toolMessages); | ||
| // Refresh conversation items added during tool execution, such as an inline | ||
| // AgentTask sub-conversation merged into the agent chat context at handoff return. | ||
| chatCtx.merge(this.agent._chatCtx, { excludeInstructions: true }); |
There was a problem hiding this comment.
🚩 Realtime path does not have equivalent merge for tool execution context
The pipeline path now merges this.agent._chatCtx into the local chatCtx at agents/src/voice/agent_activity.ts:3020 to capture items added during tool execution (e.g., sub-agent handoff conversations). The realtime path at agents/src/voice/agent_activity.ts:3572 copies from realtimeSession.chatCtx and pushes tool outputs, but does NOT perform a similar merge from this.agent._chatCtx. This may be intentional since realtime models manage their own chat context via realtimeSession.updateChatCtx(), but if the same handoff-return scenario can occur in the realtime path, those sub-conversation items would be missing.
Was this helpful? React with 👍 or 👎 to provide feedback.
| const serverVad = this.#opts.serverVad; | ||
| const commitStrategy = serverVad === undefined || serverVad === null ? 'manual' : 'vad'; |
There was a problem hiding this comment.
🚩 ElevenLabs STT default commit_strategy changed from 'vad' to 'manual'
The old code at plugins/elevenlabs/src/stt.ts:653 used this.#opts.serverVad === null ? 'manual' : 'vad', meaning undefined (the default when no serverVad option is passed) resulted in 'vad' commit strategy. The new code uses serverVad === undefined || serverVad === null ? 'manual' : 'vad', so the default is now 'manual'. This is a behavioral change for users who don't explicitly set serverVad: their ElevenLabs STT streams will now use manual commit strategy instead of server-side VAD. The test updates at stt.test.ts:283,366,396-407 confirm this is intentional. The changeset description mentions this avoids persisting per-call overrides, but the commit_strategy default flip is the more impactful change for existing users.
Was this helpful? React with 👍 or 👎 to provide feedback.
| if (tools.length > 0 && providerTools.length > 0) { | ||
| throw new Error('Gemini does not support mixing function tools and provider tools'); | ||
| } | ||
|
|
||
| if (onlySingleType && tools.length > 0) { | ||
| return tools; | ||
| } | ||
|
|
||
| tools.push(...providerTools); | ||
|
|
||
| return tools.length > 0 ? tools : undefined; | ||
| } |
There was a problem hiding this comment.
🚩 Google LLM now enforces single-type tool constraint via toToolsConfig
The new toToolsConfig function at plugins/google/src/utils.ts:163-215 throws an error when both function tools and provider tools (GeminiTool subclasses or legacy geminiTools) are present: 'Gemini does not support mixing function tools and provider tools'. This is a new runtime validation that didn't exist before. The onlySingleType parameter (used by the non-realtime LLMStream.run() at plugins/google/src/llm.ts:363) returns early with just function tools when both types are present, effectively silently dropping provider tools. This inconsistency between the two code paths (throw vs. silently drop) could be confusing. The realtime path at plugins/google/src/realtime/realtime_api.ts:1445 does not pass onlySingleType, so it would throw if both types are present.
Was this helpful? React with 👍 or 👎 to provide feedback.
Summary
Testing
Ported from livekit/agents#6084
Original PR description
When a function tool awaits an inline
AgentTask, the sub-conversation it runs is merged into the agent'schat_ctxat handoff-return. But the reply generated after tool execution uses achat_ctxsnapshot taken before the tool ran, so it's blind to everything said inside the task — in practice the agent re-asks for fields the task already collected.This refreshes the snapshot with any chat items added during tool execution before generating the tool reply, mirroring the existing
update_instructions()refresh just below it.Found while benchmarking the hotel receptionist example, where the card-collection
AgentTaskwould complete and the agent would immediately ask for the card number again.