Conversation
a realtime provider can finalize the same input-audio item more than once, each time with a longer transcript. The chat context upserts those revisions onto a single message, but the rendered transcript opened a new segment for each final, so clients stacked N segments where each contained all of its predecessors and the two views of the conversation diverged. The segment a provider item was published under is now reused when that item is finalized again. Text with no item id (the stt paths) still opens a fresh segment per final.
the legacy output also resets after flushing, which pointed the item at a segment that was never published - so the legacy channel kept stacking the revisions while the stream channel replaced them.
capture_text returns early until a participant is known, so a reservation made for an item can go unconsumed. Left behind, the next item published into it and replaced the previous utterance on screen instead of appearing below. The reservation is now derived fresh each time rather than persisting.
the realtime session opens the segment with an empty, item-less update when the user stops speaking, so the transcript carrying the item id arrives into an already open capture. Binding only when a segment opened therefore never recorded the item and the first revision still opened a second line. Text with no item no longer keeps the previous item association either: the empty update would otherwise bind the old item to the new utterance segment, and a late correction of that item would land on top of the newer line.
how could this happen? it feels like we should fix that in the realtime model plugin instead of in output. a final user transcript shouldn't carry content that already sent in previous final transcript. |
|
Fair question — I went and checked rather than assuming, and I think the cumulative final is the contract rather than the bug, but I'd rather you decide. How it happens: the provider re-finalizes an item. The xAI plugin already separates the two cases — it emits Why the transcript is cumulative by design: msg = llm.ChatMessage(role="user", content=[ev.transcript], id=ev.item_id)
self._agent._chat_ctx._upsert_item(msg)— it replaces the message content for that id. So a later final has to carry the item's full text; if it carried only what's new, the upsert would leave the chat context holding just the tail. The openai realtime plugin builds the same shape deliberately, accumulating deltas in That's also why replacement rather than append: a revision isn't always an extension. A provider can correct earlier words, and a delta can't express that while a full-text replace can. So the plugins and the chat context already agree that repeated finals are revisions — If you'd still rather it live in the plugin, the two options I can see both lose something, so it's worth being explicit about which you want:
Happy to close this and do it in the xAI plugin if you prefer one of those, or narrow this to the stream output only. Just say which. |
|
@longcw whenever you get a chance — the question above is still yours to call. Short version: the framework already treats a repeated final as a revision ( Happy to close this and do it plugin-side instead if you'd rather — just say which. |
|
I am going to fix it in xai plugin, closing this one for now |
|
the root cause is the plugin sends |
Fixes #6710.
Problem
RoomIO._forward_user_transcriptnever readsev.item_id:and
flush()ends with_reset_state(), which mints a freshSG_segment id. So when a realtime provider finalizes the same input-audio item more than once — each final carrying a longer transcript — every final opens a new segment, and the client stacks N segments where each contains all of its predecessors.Meanwhile
AgentActivity._on_input_audio_transcription_completeddoes key on the id and upserts the revisions onto one chat message, so the chat context and the rendered transcript end up disagreeing about the same conversation. @RogerHarr1's report has the measurements: one item finalized 4 times over 28.8s, rendering as four stacked segments.Fix
The segment an item was published under is remembered and reused when that item is finalized again, so a revision replaces what the client is already showing instead of appending beside it.
_capture_item()is applied to both children of_ParticipantTranscriptionOutput— the legacyrtc.Transcriptionoutput and the text-stream output — since both mint their own ids and would otherwise disagree with each other.Text that arrives with no item id (the STT paths, where
item_idisNone) is untouched: each final still opens its own segment.Verification
tests/test_room_io.py:item_atwice yields the same segment id, and a followingitem_bgets a different one — this fails onmainitem_id=None, consecutive finals still get distinct segmentsruff format --check,ruff check,mypy -p livekit.agentsand the room_io/session suites pass locally.