Conversation
A realtime provider may finalize the same input-audio transcription item more than once, each time with a longer revision of the transcript. RoomIO ignored ev.item_id when forwarding user transcripts, so every final opened a fresh segment id and clients rendered N stacked segments, each containing all of its predecessors β while the chat context correctly collapsed the revisions into one message via _upsert_item. Derive the segment id from the item id when present, so a repeated final for the same item reuses the segment id and clients revise the segment in place. Events without an item_id (the STT path) keep the previous behavior. An open segment with no visible text yet (e.g. the empty interim emitted on input_speech_stopped) is re-keyed in place instead of flushed, so no empty orphan segment is published. Fixes livekit#6710 Generated with AI Co-Authored-By: AI <ai@example.com>
|
jingyan seems not to be a GitHub user. You need a GitHub account to be able to sign the CLA. If you have already a GitHub account, please add the email address used for this commit to your account. You have signed the CLA already but the status is still pending? Let us recheck it. |
| else: | ||
| # nothing visible was published yet: re-key the open segment | ||
| self._current_id = segment_id | ||
| return |
There was a problem hiding this comment.
π‘ An empty user caption can be left permanently unfinished on clients using the older captions channel
The still-open caption is silently given a new identity (self._current_id = segment_id at livekit-agents/livekit/agents/voice/room_io/_output.py:272) even though an empty version of it was already sent to clients, so that first empty caption is never closed out and lingers as an unfinished entry.
Impact: Clients using the legacy captions channel can accumulate a stale, never-finalized empty caption for every user turn in realtime sessions.
Why the legacy sink differs from the stream sink: it publishes empty interims
The realtime path emits UserInputTranscribedEvent(transcript="", is_final=False) on input_speech_stopped with no item_id (livekit-agents/livekit/agents/voice/agent_activity.py:1922). _forward_user_transcript forwards it to capture_text("").
_ParticipantStreamTranscriptionOutput.capture_textreturns early when the cleaned text is empty (livekit-agents/livekit/agents/voice/room_io/_output.py:520-521), so nothing was ever written and re-keying the open segment is safe β that's what the new test asserts._ParticipantLegacyTranscriptionOutput.capture_texthas no such guard: it always calls_publish_transcription(self._current_id, "", final=False)(livekit-agents/livekit/agents/voice/room_io/_output.py:301). So by the time the next event carries anitem_id, an interim segment with the randomSG_...id has already been published, yet_pushed_textis""soset_segment_idtakes the "nothing visible was published yet" branch and swaps_current_id. The originally published segment is never sent again withfinal=True.
Before this PR the same _current_id was kept for the whole turn, so the empty interim was always finalized.
Prompt for agents
In _ParticipantLegacyTranscriptionOutput (livekit-agents/livekit/agents/voice/room_io/_output.py), set_segment_id uses an empty _pushed_text to decide that "nothing visible was published yet" and re-keys the open segment in place. That inference holds for _ParticipantStreamTranscriptionOutput (its capture_text returns before creating a writer when the cleaned text is empty) but not for the legacy sink: legacy capture_text unconditionally publishes an interim rtc.TranscriptionSegment, including for the empty interim the realtime path emits on input_speech_stopped. Re-keying then abandons that already-published segment, which never receives a final=True update. Consider tracking whether an interim was actually published for the current segment (a flag set in _publish_transcription/capture_text and cleared in _reset_state) and using that flag instead of _pushed_text to choose between re-keying and flushing, or skip publishing legacy interims whose visible text is empty.
Was this helpful? React with π or π to provide feedback.
|
closing in favor of #6729 |
A realtime provider may finalize the same input-audio transcription item more than once, each time with a longer revision of the transcript. RoomIO ignored ev.item_id when forwarding user transcripts, so every final opened a fresh segment id and clients rendered N stacked segments, each containing all of its predecessors β while the chat context correctly collapsed the revisions into one message via _upsert_item.
Derive the segment id from the item id when present, so a repeated final for the same item reuses the segment id and clients revise the segment in place. Events without an item_id (the STT path) keep the previous behavior. An open segment with no visible text yet (e.g. the empty interim emitted on input_speech_stopped) is re-keyed in place instead of flushed, so no empty orphan segment is published.
Fixes #6710