Repository navigation
Extending VoicePipelineAgent with a generative video avatar (we have the model, we need to figure out how to plug it in) #1424
Description
Activity
- addedquestionFurther information is requestedFurther information is requested
on Jan 29, 2025 +1 - would be helpful to know this
+1, video streaming capabilities for agent are much needed
we'll be sharing an integration guide on this very soon
Reacted by JM and Xianqi LIUReally glad to hear that!
@davidzhao can you provide a ballpark estimate of how soon is
very soon? days? weeks? months?within a week :)
Reacted by Zhicai Guo and peng2219Reacted by JM and Gustav GrimbergReacted by JM, Sanjeevkumar Managutti, ozitrance, Dang Khoa and Felix AltenbergerReacted by Long Chen and JM+1 would be helpfull
@jmizgajski @davidzhao We have already built a heavily latency-optimized
AvatarPipelineAgentabstraction at Beyond Presence and might be open to contributing that to LiveKit core if there is enough interest from the community :)We’ve been optimizing our implementation for months and have it running in HD video resolution at <1s end-to-end latency (from EU) now, you can check it out here: https://bey.chat/
Due to a lot of inbound interest, we’re already building a developer API for our avatar models right now, and we could also contribute a corresponding LiveKit integration that would enable anyone to build demos like the one I shared above in a matter of hours with LiveKit
Added latency for the audio-to-video avatar call is minimal with our implementation (~0.1s) and the main latency driver of our demo is VAD/STT actually, so even 700ms end-to-end latency or less might be achievable for people who already have hyper-optimized LiveKit audio agents today
Upvote if you’re interested and want us to contribute this and @davidzhao happy to discuss a closer collaboration in more detail if you want :)
Reacted by JM, lucjac, xcodejunkiex, Dang Khoa, Sujith Thirumalaisamy, BK and Alex RodDear @davidzhao,
It's been a week so I am checking in for your updated estimate on this. How is it looking right now? What is the updated ETA?
Also @fa9r offered donating their code with some really impressive RT stats. (kudos @fa9r, I'd love to see it as part of livekit or even as PR for now, so I could test it). Is that impacting your timelines in anyway? (@davidzhao)
Is anything I could do here to help? like testing or reviewing the PR with this, adding some performance tests?
Reacted by Dang Khoa and Felix AltenbergerHere is an example for the avatar integration #1391
It's based on the new version wip that provides much more flexible input and output definitions.
Reacted by Felix Altenberger, JM, lucjac and David Zhao@longcw @davidzhao Amazing, looks great overall, we can definitely work with that 🙌
We had also already discussed internally how the overall system interaction with a LiveKit agent might work and came up with two main options:
- Option 1: Shared LK room between avatar / agent / end user -> less overall latency, but the setup is somewhat unintuitive for developers; we were also not sure whether there is a way to prevent the agent backend to show up in the call as a 3rd ghost participant to the user in this case
- Option 2: Separate LK rooms between avatar / agent and agent / end user -> cleaner from an architecture PoV, but video needs to be transmitted twice, and receiving/forwarding streams adds complexity on agent side
Looking at #1391, it seems like you went with Option 1, using data channel for "hidden" agent-avatar communication and muting incoming avatar streams in agent based on participant metadata. Could you perhaps share some more info on why you went with Option 1 over 2 (just latency?), and how LK handles the ghost participant issue I mentioned above?
@longcw @davidzhao Amazing, looks great overall, we can definitely work with that 🙌
We had also already discussed internally how the overall system interaction with a LiveKit agent might work and came up with two main options:
- Option 1: Shared LK room between avatar / agent / end user -> less overall latency, but the setup is somewhat unintuitive for developers; we were also not sure whether there is a way to prevent the agent backend to show up in the call as a 3rd ghost participant to the user in this case
- Option 2: Separate LK rooms between avatar / agent and agent / end user -> cleaner from an architecture PoV, but video needs to be transmitted twice, and receiving/forwarding streams adds complexity on agent side
Looking at #1391, it seems like you went with Option 1, using data channel for "hidden" agent-avatar communication and muting incoming avatar streams in agent based on participant metadata. Could you perhaps share some more info on why you went with Option 1 over 2 (just latency?), and how LK handles the ghost participant issue I mentioned above?
It should be straightforward to hide the agent in the room from the client side based on the participant kind or id . The option 2 that has two rooms and makes the avatar join the two rooms to pass through the user audio seems increasing the complexity a lot (maybe also latency).
Reacted by Felix AltenbergerIt should be straightforward to hide the agent in the room from the client side based on the participant kind or id
@longcw hiding a specific participant kind would seem most intuitive to me, but I think this wouldn't work currently since both avatar and agent have kind "agent", right?
Related to this comment on the PR I guess:
we might want a new kind to distinguish [agents and avatars]
our current plan is providing sufficient hooks on the client-side to select the avatar participant. I don't think we should be hiding the controlling agent at all. There will be use cases where the frontend needs to send data to the agent (either via RPC calls or data streams)
Reacted by Felix AltenbergerAgreed that the agent should not be hidden by default, but IMO it would be helpful to point out in the avatar example README what the canonical way to do hide the agent is (since I imagine for 90% of avatar use cases you would prefer to simulate a 1:1 conversation & not have the agent show up as 3rd participant)
As a workaround, we are simply hiding all agents with camera disabled atm. Does the job, but I imagine there's probably a cleaner way.
Reacted by David Zhao@fa9r the demo you shared was absalutly stunning, can you tell me the basic tech involved behind this ? is it using livekit behide the seen or it is independent webrtc, ?
@longcw I have managed to integrate an open-source lip sync model with the dev branch. Here is a recording of an interaction on the playground:
output.mp4
Noticeably, it is not as fast as the one from @fa9r. This is still the old STT+LLM+TTS stack. I had previously a working version as a Speech to Video plugin (similar to #806) into the pipeline itself with WebSocket. But the dev branch made it much simpler.
Some issues I had working with the dev branch on this feature:
- I had to put video generation in a separate thread so as not to compete with the main thread event loop. Otherwise, the video becomes choppy even though it has many frames lining up already in the queue.
- The dev branch has some significant changes (good ones) but still took some time to fully integrate.
I am not at liberty to share the code, but here are the resources used, in case you want to give it try:
- an A10G machine with 25FPS + added <1s time to first frame.
- https://github.com/TMElyralab/MuseTalk (MIT)
Reacted by Long Chen, Xianqi LIU, Chris Wilson and Abdul Aziz Naufal Ramadhan@fa9r the demo you shared was absalutly stunning, can you tell me the basic tech involved behind this ? is it using livekit behide the seen or it is independent webrtc, ?
@Sanjeevkumar-woks yeah it's based on LiveKit behind the scenes, you can check out our plugin code here: #1669 :)
Hi @fa9r, we were interested in using your plugin.Is there any implementation for the Voice Pipeline agent to use for now?
And any expected time for the API to go live?
Dear Livekit team,
my company has a it's own real-time video avatar model and would like to use the abstraction of
VoicePipelineAgentto run it with different LLMs and your implementation of function calls. (the current demo on your repo only shows how to send video directly to room, not how to use an agent that produces video)A natural place to plug something like this in would be to extend the
ttscomponent with video-frame generation or subclassVoicePipelineAgentso that it requires an additionalavatarservice that is fed with the chunks produced bytts. Alternatively thettsservice would be replaced withavatarservice that produces both audio and video.Would you be willing to provide an example/stub that does that or guide me in how to best implement this?