Skip to content

fix(room_io): keep a re-finalized user transcript in its own segment - #6711

Closed
biztex wants to merge 4 commits into
livekit:mainfrom
biztex:fix/user-transcript-item-id
Closed

biztex wants to merge 4 commits into
livekit:mainfrom
biztex:fix/user-transcript-item-id

Conversation

@biztex

@biztex biztex commented Aug 5, 2026

Copy link
Copy Markdown
Contributor

Fixes #6710.

Problem

RoomIO._forward_user_transcript never reads ev.item_id:

await self._user_tr_output.capture_text(ev.transcript)
if ev.is_final:
    self._user_tr_output.flush()

and flush() ends with _reset_state(), which mints a fresh SG_ segment id. So when a realtime provider finalizes the same input-audio item more than once — each final carrying a longer transcript — every final opens a new segment, and the client stacks N segments where each contains all of its predecessors.

Meanwhile AgentActivity._on_input_audio_transcription_completed does key on the id and upserts the revisions onto one chat message, so the chat context and the rendered transcript end up disagreeing about the same conversation. @RogerHarr1's report has the measurements: one item finalized 4 times over 28.8s, rendering as four stacked segments.

Fix

The segment an item was published under is remembered and reused when that item is finalized again, so a revision replaces what the client is already showing instead of appending beside it.

_capture_item() is applied to both children of _ParticipantTranscriptionOutput — the legacy rtc.Transcription output and the text-stream output — since both mint their own ids and would otherwise disagree with each other.

Text that arrives with no item id (the STT paths, where item_id is None) is untouched: each final still opens its own segment.

Verification

tests/test_room_io.py:

  • finalizing item_a twice yields the same segment id, and a following item_b gets a different one — this fails on main
  • with item_id=None, consecutive finals still get distinct segments

ruff format --check, ruff check, mypy -p livekit.agents and the room_io/session suites pass locally.

a realtime provider can finalize the same input-audio item more than once,
each time with a longer transcript. The chat context upserts those revisions
onto a single message, but the rendered transcript opened a new segment for
each final, so clients stacked N segments where each contained all of its
predecessors and the two views of the conversation diverged.

The segment a provider item was published under is now reused when that item
is finalized again. Text with no item id (the stt paths) still opens a fresh
segment per final.
@biztex
biztex requested a review from a team as a code owner August 5, 2026 16:07
devin-ai-integration[bot]

This comment was marked as resolved.

the legacy output also resets after flushing, which pointed the item at a
segment that was never published - so the legacy channel kept stacking the
revisions while the stream channel replaced them.
devin-ai-integration[bot]

This comment was marked as resolved.

capture_text returns early until a participant is known, so a reservation
made for an item can go unconsumed. Left behind, the next item published into
it and replaced the previous utterance on screen instead of appearing below.
The reservation is now derived fresh each time rather than persisting.
devin-ai-integration[bot]

This comment was marked as resolved.

the realtime session opens the segment with an empty, item-less update when
the user stops speaking, so the transcript carrying the item id arrives into
an already open capture. Binding only when a segment opened therefore never
recorded the item and the first revision still opened a second line.

Text with no item no longer keeps the previous item association either: the
empty update would otherwise bind the old item to the new utterance segment,
and a late correction of that item would land on top of the newer line.
@longcw

longcw commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

a realtime provider finalizes the same input-audio item more than once — each final carrying a longer transcript — every final opens a new segment

how could this happen? it feels like we should fix that in the realtime model plugin instead of in output. a final user transcript shouldn't carry content that already sent in previous final transcript.

@biztex

biztex commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Fair question — I went and checked rather than assuming, and I think the cumulative final is the contract rather than the bug, but I'd rather you decide.

How it happens: the provider re-finalizes an item. The xAI plugin already separates the two cases — it emits is_final=False for status="in_progress" and only delegates the "completed" one — so these aren't partials leaking through. @RogerHarr1's instrumentation logged four is_final=True events for a single item_id over 28.8s, with the transcript growing 30 → 78 → 132 → 177 chars.

Why the transcript is cumulative by design: _on_input_audio_transcription_completed does

msg = llm.ChatMessage(role="user", content=[ev.transcript], id=ev.item_id)
self._agent._chat_ctx._upsert_item(msg)

— it replaces the message content for that id. So a later final has to carry the item's full text; if it carried only what's new, the upsert would leave the chat context holding just the tail. The openai realtime plugin builds the same shape deliberately, accumulating deltas in _input_transcript_accumulators and emitting the running total rather than each delta.

That's also why replacement rather than append: a revision isn't always an extension. A provider can correct earlier words, and a delta can't express that while a full-text replace can.

So the plugins and the chat context already agree that repeated finals are revisions — RoomIO is the one place that treats each as a new utterance, which is how the two views ended up disagreeing for the same conversation.

If you'd still rather it live in the plugin, the two options I can see both lose something, so it's worth being explicit about which you want:

  1. Suppress later finals — the corrected text is dropped entirely, and the transcript keeps whatever the first (shortest) final said.
  2. Emit only new content — needs _upsert_item to append instead of replace, across every realtime plugin, and still can't represent a correction to earlier words.

Happy to close this and do it in the xAI plugin if you prefer one of those, or narrow this to the stream output only. Just say which.

@biztex

biztex commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

@longcw whenever you get a chance — the question above is still yours to call.

Short version: the framework already treats a repeated final as a revision (_upsert_item replaces the message by item_id, and the openai plugin emits a running total rather than deltas), so RoomIO is the only layer that doesn't. Moving it into the plugin means either dropping the corrected text or switching that upsert to append, which can't express a provider correcting earlier words.

Happy to close this and do it plugin-side instead if you'd rather — just say which.

@longcw

longcw commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

I am going to fix it in xai plugin, closing this one for now

@longcw longcw closed this Aug 6, 2026
@longcw

longcw commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

the root cause is the plugin sends final transcripts with accumulated content because xAI realtime emits transcription event differently. fixing it in room io is in a wrong layer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

room_io: user transcript segments ignore item_id, so an in-place item revision renders as N separate segments

2 participants