Skip to content

Getting a "WARNING livekit.agents - The realtime API returned a text content part, which is not supported" warning during conversations #1143

Description

@zacharyw

Hello - I'm not sure if this is a bug or something I'm doing wrong.

Using RealtimeModel`MultimodalAgent` I am attempting to start a conversation and seed it with some conversation history so the user can pick up where they left off when the conversation.

I am setting the modality to "audio" to try and ensure text responses aren't being used, but I sometimes get this warning/error, and no audio output is produced. It seems to happen somewhat randomly: the more convo history there is, the more likely it seems to happen.

Some code for how I've set things up:

model = openai.realtime.RealtimeModel(
            instructions=data['globalPrompt'],
            voice='shimmer',
            temperature=0.8,
            # max_response_output_tokens=float('inf'),
            modalities=['audio'],
            turn_detection=openai.realtime.ServerVadOptions(
                threshold=0.9, prefix_padding_ms=200, silence_duration_ms=500
            ),
        )

agent = MultimodalAgent(model=model)

agent.start(ctx.room)

logger.info("starting agent")
        
session = model.sessions[0]
# Add messages to conversation history if needed
for message in data.get('messages', []):
    logger.info(f"role: {message['role']}, content: {message['content']}")
    session.conversation.item.create(
        llm.ChatMessage(role='assistant' if message['role'] == 'system' else 'user', content=message['content'])
    )
        
session.response.create()

messages is an array of messages returned from my API that just contains content (a string) and role (a string). It seems that setting role to "assistant" instead of "system" seems to reduce the frequency of this issue, but it could be a placebo effect.

This code is based on the example code from the integration guide: https://docs.livekit.io/agents/openai/multimodalagent/

Which is interesting that this example code sets audio and text modalities, when text doesn't really seem to work at all.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    questionFurther information is requested

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions