Skip to content

fix(realtime): separate server turn detection from server replies - #6963

Closed
longcw wants to merge 1 commit into
mainfrom
longc/rt-server-segmentation-caps
Closed

longcw wants to merge 1 commit into
mainfrom
longc/rt-server-segmentation-caps

Conversation

@longcw

@longcw longcw commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

Problem

capabilities.turn_detection tells the framework who owns the input audio buffer. The openai plugin sets it from create_response, which says who answers the turn.

With turn_detection=ServerVad(create_response=False) the server still segments and commits every turn, but the framework reads the capability as client-side turn taking and commits as well. Each turn logs input_audio_buffer_commit_empty, which #6642 suppressed rather than removed.

Seven reads of this capability want the "the server segments" meaning, so barge-in gatekeeping and the default VAD were wrong for this configuration too.

Fix

turn_detection goes back to its first meaning: the server detects turns, and it segments and commits the input audio itself.

The half it absorbed becomes auto_turn_reply_generation, which is True by default. Only the openai plugin needs a change, and a plugin that does not know the field behaves as before.

Five of the seven wrong reads become correct with no edit. AgentActivity gains _rt_server_reply_enabled for the six places that want "the server answers", and only the two commit_audio() calls needed an explicit gate.

Note: with an explicit create_response=False, endpointing now comes from the server's own speech events, so a client turn_detection="vad" is ignored with a warning. Pass turn_detection=None to keep full client control. This ordering is what closes the gap where response.create could otherwise overtake the server's commit and answer a turn the conversation does not hold yet.

capabilities.turn_detection tells the framework who owns the input
audio buffer, but the openai plugin set it from create_response, which
says who answers the turn. With ServerVad(create_response=False) the
server keeps segmenting and committing every turn while the framework
commits as well, so each turn logged input_audio_buffer_commit_empty.

turn_detection goes back to meaning "the server segments and commits
the input audio". The half it absorbed becomes
auto_turn_reply_generation, which defaults to True, so only the openai
plugin needs a change. Five of the seven reads that wanted the
segmenting meaning become correct with no edit, and AgentActivity
reads the new flag at the six places that want "the server answers".

Endpointing for an explicit create_response=False now comes from the
server, so a client turn_detection="vad" is ignored with a warning.
Pass turn_detection=None to keep full client control.
@longcw
longcw requested a review from a team as a code owner August 24, 2026 06:22
@longcw
longcw marked this pull request as draft August 24, 2026 06:24

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Devin Review: No Issues Found

Devin Review analyzed this PR and found no potential bugs to report.

View in Devin Review to see 1 additional finding.

Open in Devin Review

@longcw

longcw commented Aug 24, 2026

Copy link
Copy Markdown
Contributor Author

closing and create a fix in openai realtime instead #6965

@longcw longcw closed this Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants