Skip to content

Extending VoicePipelineAgent with a generative video avatar (we have the model, we need to figure out how to plug it in) #1424

Description

@jmizgajski

Dear Livekit team,

my company has a it's own real-time video avatar model and would like to use the abstraction of VoicePipelineAgent to run it with different LLMs and your implementation of function calls. (the current demo on your repo only shows how to send video directly to room, not how to use an agent that produces video)

A natural place to plug something like this in would be to extend the tts component with video-frame generation or subclass VoicePipelineAgent so that it requires an additional avatar service that is fed with the chunks produced by tts. Alternatively the tts service would be replaced with avatar service that produces both audio and video.

Would you be willing to provide an example/stub that does that or guide me in how to best implement this?

Activity

  1. MikeKras commented on Jan 29, 2025

    @MikeKras

    +1 - would be helpful to know this

  2. Maelstro commented on Jan 29, 2025

    @Maelstro
    Contributor

    +1, video streaming capabilities for agent are much needed

  3. davidzhao commented on Jan 29, 2025

    @davidzhao
    Member

    we'll be sharing an integration guide on this very soon

  4. jmizgajski commented on Jan 29, 2025

    @jmizgajski
    Author

    Really glad to hear that!

    @davidzhao can you provide a ballpark estimate of how soon is very soon? days? weeks? months?

  5. davidzhao commented on Jan 30, 2025

    @davidzhao
    Member

    within a week :)

  6. joeylin commented on Feb 2, 2025

    @joeylin

    +1 would be helpfull

  7. fa9r commented on Feb 4, 2025

    @fa9r
    Contributor

    @jmizgajski @davidzhao We have already built a heavily latency-optimized AvatarPipelineAgent abstraction at Beyond Presence and might be open to contributing that to LiveKit core if there is enough interest from the community :)

    We’ve been optimizing our implementation for months and have it running in HD video resolution at <1s end-to-end latency (from EU) now, you can check it out here: https://bey.chat/

    Due to a lot of inbound interest, we’re already building a developer API for our avatar models right now, and we could also contribute a corresponding LiveKit integration that would enable anyone to build demos like the one I shared above in a matter of hours with LiveKit

    Added latency for the audio-to-video avatar call is minimal with our implementation (~0.1s) and the main latency driver of our demo is VAD/STT actually, so even 700ms end-to-end latency or less might be achievable for people who already have hyper-optimized LiveKit audio agents today

    Upvote if you’re interested and want us to contribute this and @davidzhao happy to discuss a closer collaboration in more detail if you want :)

  8. jmizgajski commented on Feb 7, 2025

    @jmizgajski
    Author

    Dear @davidzhao,

    It's been a week so I am checking in for your updated estimate on this. How is it looking right now? What is the updated ETA?

    Also @fa9r offered donating their code with some really impressive RT stats. (kudos @fa9r, I'd love to see it as part of livekit or even as PR for now, so I could test it). Is that impacting your timelines in anyway? (@davidzhao)

    Is anything I could do here to help? like testing or reviewing the PR with this, adding some performance tests?

  9. longcw commented on Feb 11, 2025

    @longcw
    Contributor

    Here is an example for the avatar integration #1391

    It's based on the new version wip that provides much more flexible input and output definitions.

  10. fa9r commented on Feb 11, 2025

    @fa9r
    Contributor

    @longcw @davidzhao Amazing, looks great overall, we can definitely work with that 🙌

    We had also already discussed internally how the overall system interaction with a LiveKit agent might work and came up with two main options:

    • Option 1: Shared LK room between avatar / agent / end user -> less overall latency, but the setup is somewhat unintuitive for developers; we were also not sure whether there is a way to prevent the agent backend to show up in the call as a 3rd ghost participant to the user in this case
    • Option 2: Separate LK rooms between avatar / agent and agent / end user -> cleaner from an architecture PoV, but video needs to be transmitted twice, and receiving/forwarding streams adds complexity on agent side

    Looking at #1391, it seems like you went with Option 1, using data channel for "hidden" agent-avatar communication and muting incoming avatar streams in agent based on participant metadata. Could you perhaps share some more info on why you went with Option 1 over 2 (just latency?), and how LK handles the ghost participant issue I mentioned above?

  11. longcw commented on Feb 13, 2025

    @longcw
    Contributor

    @longcw @davidzhao Amazing, looks great overall, we can definitely work with that 🙌

    We had also already discussed internally how the overall system interaction with a LiveKit agent might work and came up with two main options:

    • Option 1: Shared LK room between avatar / agent / end user -> less overall latency, but the setup is somewhat unintuitive for developers; we were also not sure whether there is a way to prevent the agent backend to show up in the call as a 3rd ghost participant to the user in this case
    • Option 2: Separate LK rooms between avatar / agent and agent / end user -> cleaner from an architecture PoV, but video needs to be transmitted twice, and receiving/forwarding streams adds complexity on agent side

    Looking at #1391, it seems like you went with Option 1, using data channel for "hidden" agent-avatar communication and muting incoming avatar streams in agent based on participant metadata. Could you perhaps share some more info on why you went with Option 1 over 2 (just latency?), and how LK handles the ghost participant issue I mentioned above?

    It should be straightforward to hide the agent in the room from the client side based on the participant kind or id . The option 2 that has two rooms and makes the avatar join the two rooms to pass through the user audio seems increasing the complexity a lot (maybe also latency).

  12. fa9r commented on Feb 13, 2025

    @fa9r
    Contributor

    It should be straightforward to hide the agent in the room from the client side based on the participant kind or id

    @longcw hiding a specific participant kind would seem most intuitive to me, but I think this wouldn't work currently since both avatar and agent have kind "agent", right?

    Related to this comment on the PR I guess:

    we might want a new kind to distinguish [agents and avatars]

  13. davidzhao commented on Feb 13, 2025

    @davidzhao
    Member

    our current plan is providing sufficient hooks on the client-side to select the avatar participant. I don't think we should be hiding the controlling agent at all. There will be use cases where the frontend needs to send data to the agent (either via RPC calls or data streams)

  14. fa9r commented on Feb 13, 2025

    @fa9r
    Contributor

    Agreed that the agent should not be hidden by default, but IMO it would be helpful to point out in the avatar example README what the canonical way to do hide the agent is (since I imagine for 90% of avatar use cases you would prefer to simulate a 1:1 conversation & not have the agent show up as 3rd participant)

    As a workaround, we are simply hiding all agents with camera disabled atm. Does the job, but I imagine there's probably a cleaner way.

  15. Sanjeevkumar-woks commented on Feb 16, 2025

    @Sanjeevkumar-woks

    @fa9r the demo you shared was absalutly stunning, can you tell me the basic tech involved behind this ? is it using livekit behide the seen or it is independent webrtc, ?

  16. ChenghaoMou commented on Feb 28, 2025

    @ChenghaoMou
    Contributor

    @longcw I have managed to integrate an open-source lip sync model with the dev branch. Here is a recording of an interaction on the playground:

    output.mp4

    Noticeably, it is not as fast as the one from @fa9r. This is still the old STT+LLM+TTS stack. I had previously a working version as a Speech to Video plugin (similar to #806) into the pipeline itself with WebSocket. But the dev branch made it much simpler.

    Some issues I had working with the dev branch on this feature:

    1. I had to put video generation in a separate thread so as not to compete with the main thread event loop. Otherwise, the video becomes choppy even though it has many frames lining up already in the queue.
    2. The dev branch has some significant changes (good ones) but still took some time to fully integrate.

    I am not at liberty to share the code, but here are the resources used, in case you want to give it try:

    1. an A10G machine with 25FPS + added <1s time to first frame.
    2. https://github.com/TMElyralab/MuseTalk (MIT)
  17. fa9r commented on Mar 18, 2025

    @fa9r
    Contributor

    @fa9r the demo you shared was absalutly stunning, can you tell me the basic tech involved behind this ? is it using livekit behide the seen or it is independent webrtc, ?

    @Sanjeevkumar-woks yeah it's based on LiveKit behind the scenes, you can check out our plugin code here: #1669 :)

  18. akshitdudeja125 commented on Mar 18, 2025

    @akshitdudeja125

    Hi @fa9r, we were interested in using your plugin.Is there any implementation for the Voice Pipeline agent to use for now?
    And any expected time for the API to go live?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    questionFurther information is requested

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions