Intro0:00
So, uh, my name's Kwindla. I work at a company called Daily. We make real-time audio and video infrastructure, and when GPT-4 was released last year, we started building lots and lots of experiments. We thought GPT-4 was so good at conversational exchange in text that it really should be able to talk to you, literally talk to you, with your voice.
So we built a bunch of things like this. This is probably a year ago, and in GenAI terms it's ancient, ancient history, but probably worth looking at as context.
I'm your friendly chatbot, here to assist you with any questions or tasks you have in mind. How can I help you today?
Uh, can you tell me a joke?
Absolutely. Why don't skeletons fight each other?
Oh, you always start with that one. Give me a different joke.
Alright, I'll give it another go. What do you call fake spaghetti? An imp.
Oh, I've heard that one too. Come on, maybe just—
So we used "Tell me a joke" from, you know, very early on all the way through to today, and I think we know all the jokes that all the LLMs we use regularly, uh, tell, which is funny, but not necessarily funny in the same way they were the first time we heard them.
So this is sort of a really high-level schematic of what we're trying to do here,right? We've got a user on a phone or a laptop, they want to talk to their device, and then somewhere in the cloud we've got a bunch of GPUs that are doing a whole lot of heavy-duty compute, and we need to talk to those, that cloud computing resource, somehow.
So as soon as you build stuff like the video we just saw, a couple of things come very much top of mind. One is speed really matters, and the other is architectural flexibility is really, really important. I'll talk about both of those things today.
Architecture2:03
Let's start with architectural flexibility. So that really nice clean diagram gets really messy really fast. This is not unusual in an engineering software development problem domain, but I thought I'd kind of make a slide based on, like, it looked at a bunch of source code, I thought about all the conversations I've had with, you know, my colleagues and our customers and friends who are building this stuff, and it turns out you have to, at some level, kind of be aware of a bunch of these things if you want to build real-time robust voice AI stuff, deploy it, scale it to production.
It's a little bit of an intimidating map. We're definitely putting the multi- and multimodal AI here, all the way from audio processing, things like echo cancellation, and CPU management when you're encoding and decoding audio and video, through networking issues like firewall traversal, all the way through to building things like retrieval augmented generation and tool calling so that your real-world applications are really, really useful.
We can collect that kind of messy map into a few, a little bit more kind of high-level categories. It's worth going through these just really quickly because I think they give you a sense of what that map is.
So you need really robust and low-latency media processing and transport. You've got to encode the media, you've got to send it over the network, that's got to work really well, it's got to work really fast. You need really good and fast transcription, at least until the future of truly multimodal audio-native models comes, which will happen at some point.
And even after that, you probably are going to need to go from audio to text for lots of kinds of AI use cases. You have to do lots of real-time data pipeline and buffer management. So I think in Discord I've probably maybe 20 or 30 times answered the question, "Why is my audio stream not working?"
when I did local development on my Mac and then I pushed it to an Intel box in the cloud. It's because, you know, Indian issues always get you if you're writing low-level audio stuff. But there's a lot of sort of pipeline and buffer management that comes into play here.
You want to be able to swap between models for a whole bunch of reasons, or use multiple models together. You generally need to call out to external systems. You have to do things like phrase endpointing, which is the fancy academic term for when is the person done talking and when do they expect the AI to talk.
You need to handle interruptions really gracefully, and that actually is a whole rabbit hole of its own. People will interrupt the bot, and you need to figure out what you're going to do and how you're going to handle that and how you're going to maintain state when people interrupt.
You have to do echo cancellation unless you can convince everybody in the world who's using your thing to wear headphones 100% of the time. And I've been doing audio-video development on the Internet for a really long time. Like, our core product at Daily is a WebRTC tech stack, so I can tell you that you will not convince everybody in the world to wear headphones all the time.
Maybe if you're only deploying to people in, like, a call center professional context, but even then it's pretty tough. And you need good fast text-to-speech or voice generation on the other end.
The good thing is, it's a lot, but when it all comes together, it really is amazing to build with these new GenAI tools. You can build stuff that feels magical and new to people, but that is genuinely useful, that sort of just feels like how things should be.
Healthcare5:25
We built this is an example of something we built for a healthcare customer that solved a real pain point. I think if you've been to the doctor, which I'm sure all of us have, you've had to fill out way more forms than you thought you should at subsequent doctor visits.
And maybe you filled those out on paper, maybe you, you know, got a link to a form to fill out before the visit, which is like the current state of the art. What if an LLM could just call you on your phone, or you could click a link in an app, and you can have a conversation and you never had to fill out a form again?
Hello, chat. My name is Jessica, and I am with Tri-County Health Services. To ensure your privacy and to confirm your identity, could you please provide me with your birthday, including the year?
Uh, yes, January 1st.
I appreciate that, Chad. Can you also tell me the.
Oh, sorry, yeah, 1983.
Thank you for providing your birthday, Chad. Now, could you please list your current prescriptions, including the medication name and the dosage for each?
Uh, yeah, I take lisinopril and also gabapentin.
Great. Could you please provide the dosage?
Oh, yeah, sorry, 10 milligrams and 100 milligrams respectively.
Thank you for that information, Chad. Additionally, do you have any allergies that we should be aware of? If so, please list them.
Uh, yeah, I'm allergic to latex, and I'm also allergic to penicillin.
Thank you for letting us know about your allergies, Chad. Next.
So I think you can see how all those pieces we just talked about are actually necessary to pull together something like that, but when you do, it feels simple, it feels like it's supposed to. That's a real video, nothing in there is edited, the callouts to the external system are mocked, they're not, it's not actually a real electronic health record systems on the back end, but it's, it's, they're real tool calling.
And that stuff is now starting to move into production in, in the real world. So I think in a lot of cases the GenAI frontier is moving very quickly, and people are really sort of pulling it out of our hands, which is super exciting.
So let's switch gears a little bit and talk about fast, because that's the other really big piece here. How fast do we need? How fast is fast? And also just a quick side note, so we're talking, there are always two kinds of fast when you're talking about engineering things.
Latency7:43
There's throughput and there's latency. These days, for conversational interactions, throughput is pretty okay for all the tools we all use today. Like, LLMs and other tools can generate content as fast as people can read it or listen to it.
But what's hard is latency. And latency is that sort of time to first byte, time to first token. In lots and lots of engineering contexts, there's trade-offs between throughput and latency, complicated relationships between throughput and latency. One of the graphs that I sometimes show in these talks is that throughput tends to improve by an order of magnitude every couple of years in lots of domains.
Latency improvements tend to be linear, like way behind throughput improvements. So latency is hard, and latency is mostly what bites us here. Human conversational latency, like if I am talking to another person, it feels weird to me if that person doesn't respond in about half a second.
Sometimes people respond actually a lot faster. We seem as humans to be doing, like, speculative decoding, next token prediction, just like natively. Like, that's what we do. I know what you're going to say, four or five words before you finish saying it, I'm queuing up my response, I'm sort of doing my inference in this, like, greedy fashion.
If you say something I didn't expect, well, I can, like, reroute. But most of the time I'mright. And most of the time, if you actually record people in conversation, they'll respond in, like, 200 or 300 milliseconds, commonly. And if they don't, they'll give you some kind of cue.
So the, the sort of 500 millisecond target is, is pretty important, because we hit that uncanny valley pretty quickly when we're above it. In fact, I think that video of my colleague Chad that you just saw, if you watched it with a critical eye, what I hope you saw was pretty cool orchestration of, like, state-of-the-art GenAI stuff, and probably slower response times than really should be there.
So we've spent the last couple months, I've spent the last couple months really sort of thinking a lot about how to improve these response times. And just as a kind of benchmark, like, relative, like, another number that shows how hard this is, like, Gemini Pro's time to first token, it's like 900 milliseconds.
So if you're aiming for 500 milliseconds, you're already almost double, even before you do anything else, even before you send stuff, you know, over the network for other services or anything. So what models and tools you choose are constrained, they matter a lot, it matters a lot how you string them together.
So just to pop up a level again, this is what we're trying to achieve. And the most powerful tool we have today for making everything run fast in this domain is actually putting as much together into one compute container as we possibly can.
So if the, if the really, really big things we're trying to do are natural language, speech-to-text, and then phrase endpointing, so when should the bot do processing or talk, and then LLM inference, and then voice output, if we can put all those things together and run them locally and colocated, we're, like, way ahead of where we are if we can't do that.
And this is worth emphasizing because I think, I'm sure, like, 95, 98, 99 percent of stuff we're all building today with GenAI, we're calling out to hosted services. There's a lot of really, really good reasons for that. But that's tough in this domain if latency is what you're prioritizing.
And latency might not be what you're prioritizing, and that's okay, like, there's lots of different trade-offs you can make, but if you're trying to make things really, really, really fast, you need to figure out how to host stuff yourself and how to host stuff in a way where you can tune and control and combine everything.
So this is the part of the talk where I, like, look at the clock and I look out at all of you and I try to figure out how much tolerance you have for me talking about latency. Because I will, maybe ironically, we'll talk about latency for hours and hours, it's what I'm obsessed with as an engineer.
I do think it's worth just quickly kind of going over this list of, like, kind of the best you can hope for latency numbers for a typical voice AI context, because some of them are non-obvious. So first, what are we actually measuring?
We're measuring the time, like, what do we really, really care about? We measure the time I stop talking. So if there's, like, a green waveform on one side and a purple waveform on the other side of this, like, you know, audio editor, the time I stop talking, and then there's some kind of gap, usually silence, we could play hold music if it's too long a gap.
And then there's another waveform on the other side when I first start to hear the LLM talking to me. That's the gap we care about, the voice-to-voice latency. And that has to include everything, it has to include audio encoding, sending stuff over the network, all the processing, sending stuff back, playing it out to speakers.
The very first number here is actually kind of shockingly high. If you're using the laptop mic on macOS, it's like 30, 40 milliseconds of delay. There are some good reasons for that. There's a lot of audio processing, again, echo cancellation, background noise removal, gain control, all that stuff plays in, plays a role here.
As these kind of use cases get more and more important, maybe hardware and operating system manufacturers will push this stuff down. But today, you're sort of 40 milliseconds into your 500 millisecond target time before the bytes ever even get into any of what we think of as software we control.
And then you're encoding the audio, you're sending it over the network, you're doing some packet handling, you're reassembling those packets on the other side in what we call a jitter buffer, you're decoding, and then you're finally at the, like, transcription stage.
If you are really careful about transcription and endpointing, you can probably get down to a couple hundred milliseconds. Then you're doing LLM inference. In this context, you're really hoping you can get your time to first token from the LLM in maybe another 100 milliseconds.
Then you're doing some speech, sorry, text-to-speech, and then you're going over the network and out the other side. These are real measured numbers, sort of medians. There's obviously some hand-waving because it's hard to measure everything perfectly, and there's huge amounts of variation.
But you're, you're at like 600 to 700 milliseconds, optimistically but realistically, in this world. So that gives you a sort of sense of where we are.
Yesterday, we posted a demo on Twitter of really trying to drive all this latency down. And we have, like, a container that's open source that you can kind of look at and play with and hack on that gets you to, like, the 500 to 700 milliseconds of latency.
FastBot14:22
Hello, I'm FastBot. How can I help you today?
Let's play a geography game. Do you want to give me some geography questions to answer?
That's the spirit. I'll give you a question. What is the world's largest desert?
Oh, I think it's the Sahara.
That's correct. The Sahara Desert is indeed the largest hot desert in the world, covering most of North Africa. Well done. Next question. Which river is the longest in South America?
It's got to be the, not the Nile, sorry, the Amazon.
No need to apologize. You're absolutelyright again. The Amazon River is indeed the longest river in South America and one of the longest rivers in the world. Good job. Next question. What is the capital city of Australia?
Canberra.
Correct again. Canberra is the capital city of Australia. You're on a roll. Next question. Which mountain range runs along the border between France and Spain?
Pyrenees.
Exactlyright. The Pyrenees mountain range does indeed run along the border between France and Spain. You're really showing off your geography skills here. Next question. What is the largest city in Scandinavia?
Visit Stockholm.
I think there might be a slight mistake. Visit.com is not a city, but rather a travel website.
I was trying to say Stockholm.
No worries. I knew what you meant. And you're absolutelyright again. Stockholm is indeed the largest city in Scandinavia and the capital of Sweden. You're doing great. Next question. Which city is the largest in Poland?
Krakow.
Close but not quite. Krakow is a beautiful city in Poland, but it's not the largest. The correct answer is Warsaw, which is the capital and largest city of Poland.
Oh, I feel like I should have known that. If we're going to talk about capitals, maybe I'll ask you some questions. What's the capital of France?
I'm ready for a challenge. The capital of France is Paris.
Germany?
Not quite. Well, Germany is a country with a great—
No, I meant, what's the capital of Germany?
Good one. The capital of Germany is indeed Berlin.
And what's the capital of Stockholm? I mean, sorry, what's the capital of Sweden?
No worries. The capital of Sweden is Stockholm.
What a—
So the, the, the idea there was not to do prompt engineering, because clearly I could have better done a better job on prompt engineering, but to kind of really test latency and show all the latency numbers in one place, because it's useful to visualize stuff.
And again, that UI and everything is open source. So the, the big number in the middle
is the, the voice-to-voice latency measured on the client side, and then the numbers on the edge are metrics that are being sent from the server that's doing all that work, sort of on the individual pieces, the measurements sent back, so you can sort of see what the variation is and correlate and kind of get good intuitions about this stuff.
The architecture here is two models by a company called Deepgram for the transcription and the voice generation that are really good compromises between how good they are and how fast they are. And Deepgram has a hosted service, but they also let you run those models on premises in little Docker containers.
And that's Llama 3 8B, I think, because I couldn't quite get 70B to run as fast as I wanted it to, although in theory that's possible. And I'll post some links to this if you want to look at it more.
PipeCat18:11
So because we solved so many problems over and over, we thought it would be great to have an open source framework for this stuff. I think we've seen this in other parts of AI landscape. Things like LangChain and LlamaIndex are really valuable.
This is sort of that for real-time and multimodal AI. And this slide probably looks familiar because I stole the list of hard problems from this slide and made a slide that I moved higher up in the talk here for today.
But this is an open source framework called PipeCat. It's gotten a bunch of traction recently. It's vendor neutral, even though it came out of work that we've done at Daily early on on this. And we're just really excited about this.
It's super fun to be getting lots of community contributions now. And if you are trying to build really fast multimodal AI stuff, I think it's at least worth taking a look at. You can build things like conversational bots and speech-to-speech, language translation apps, and voice-controlled agents of various kinds, like, that control your software user interfaces, and real-time vision model stuff, like the awesome last presentation is also, like, baked into PipeCat services.
Now, here's all the stuff that's supported in PipeCat today. We're adding stuff all the time. You can add stuff. So if you're interested in building, please hang out with us in the PipeCat Discord. If you want to contribute a service plugin, please do that.
If you want to be a maintainer for an open source project that's a lot of fun, ping me. Maintainers are, you know, gold in the open source world. We're all always trying to recruit great maintainers. And just last slide about the context here.
So this is the PipeCat star rating list, and the day that it went vertical was the GPT-4.0 announcement. We're going to get great multimodal models, and they'll be incredibly useful. They'll make building super fast stuff easier and easier.
But A, they're not here yet, and B, we're still going to need orchestration layers for all this stuff. Also, the, the, the demo that I showed just a minute ago that I posted yesterday is now at 175,000 views on Twitter.
Outro20:01
So there's more and more and more interest in voice AI, and we'd love to have people come build with us.





