AIAI EngineerJul 31, 2025· 16:30

Your realtime AI is ngmi — Sean DuBois (OpenAI), Kwindla Kramer (Daily)

Sean DuBois (OpenAI, Pion) and Kwindla Kramer (Daily, Pipecat) argue that realtime AI systems must be built from the network layer up, with WebRTC essential for voice AI over WebSockets. WebSockets rely on TCP, causing glitches in 10–15% of connections, while WebRTC uses UDP for sub-second voice-to-voice latency (they demonstrate ~500ms target). Kwindla shows a latency breakdown from a Pipecat app, and Sean notes WebRTC handles packet loss, jitter, and bandwidth estimation automatically. They demo Squabbert, a Raspberry Pi stuffed animal using peer-to-peer WebRTC with MLX Whisper and Gemma 3 for local AI. Yashin, a non-technical mom, shares her bilingual learning app built with Sean's guidance, highlighting WebRTC's accessibility.

Transcript

Opening Skit0:00

Alright, Squabbert, you ready to get packed up?

Squabbert0:18

I don't know, I'm pretty nervous.

Oh, relax. You got nothing to worry about.

Squabbert0:22

But this is like the worst idea ever for a live demo. A live, unscripted conversation with a non-deterministic LLM? On conference Wi-Fi?

Come on. Your prompt is great, and I mean, it's gotta be tough.

Squabbert0:35

With on-stage audio and a room full of echoes? Sometimes my text-to-speech even mispronounces my own name. Why did you even name me, Squabbert? What if I say "Squimbly" or something again?

Okay, take a breath—well, you don't really do that. Let's just take it one step at a time. Just start with the intro.

Squabbert0:53

I guess you'reright. Okay, here I go. Hi everybody, I'm Squabbert, here to take us on a whirlwind tour of the wonderful world of WebRTC. Please welcome... Sean and Kwin.

Sean DuBois1:08

Hey, I'm Sean. I work on WebRTC at OpenAI. Um, some of the things you might be familiar with are the Realtime API, our 1-800-CHATGPT. You can call itright from your phone. Um, before I worked at OpenAI, I worked on the GO implementation of WebRTC called Pion.

Sean & Kwin1:08

Kwindla Hultman Kramer1:23

And I'm Kwin. I work at Daily on real-time audio and video infrastructure, and on an open-source voice agent framework called Pipecat. Uh, today we're going to talk about how to build natural, fast, human-like voice experiences. Uh, we're going to give you a crash course on low-latency audio and video, and I hope we'll show you a couple of things you might not have thought of around voice AI before.

Latency1:47

Sean DuBois1:47

Um, if you want to build a conversational voice experience that people really love, you're going to stress a lot about latency. Nothing else matters if your AI responds too slowly.

Kwindla Hultman Kramer1:57

Building voice AI experiences is similar to other kinds of AI engineering in most ways. If you've built multi-turn agents, a lot of that will port over to building voice agents. But the big difference is latency.

Sean DuBois2:11

Everything in a voice AI app needs to be built from the ground up for fast response times. If you're talking to a person, around 500 milliseconds sounds natural. When talking to an AI system, people bring those same expectations: response latencies much above a second.

In general, doom your voice agent to very low completion rates and low NPS scores and hangups.

Kwindla Hultman Kramer2:33

And we're talking here about voice-to-voice latency. So this is the time, uh, between when I, the human, stop talking and the time I hear the first audio byte come back from the LLM.

Let's take a look at how latency adds up in a typical voice-to-voice AI application. So this is a breakdown from a real voice AI app running in a web browser on macOS, talking over the internet to a voice agent running in the cloud, running on Pipecat.

Sean DuBois3:03

A couple of things to note: our voice latency is just under a second. That's good, but not great. We can make things a little faster, but that comes with trade-offs: lower quality or cost. Second, it's frustratingly easy to do even worse than this.

Your LLM might be slower, other things get in the way, or worst of all, Bluetooth.

Kwindla Hultman Kramer3:22

Don't get me started on Bluetooth.

Sean DuBois3:25

But the single biggest mistake we see people making is using the wrong approach to sending and receiving audio over the network. It's time to talk about WebRTC and WebSockets.

Kwindla Hultman Kramer3:36

If you're new to building voice applications, you probably think, "Hey, I need a, like, a long-lived, uh, connection. I'm going to send audio and video over this long-lived connection." I've used WebSockets for long-lived data connections before. I'm just going to write some WebSockets code.

WebSockets vs WebRTC3:36

Kwindla Hultman Kramer3:51

That's great if you're doing long-lived, short, small amounts of data. It doesn't work for real-time audio. In fact, WebSockets are almost the opposite of what you want from a network engineering perspective for real-time audio and video.

Sean DuBois4:05

So let's do a compare and contrast on this.

WebSockets are great if you're trying to deliver audio, and you want something really easy that can target all platforms. If you're trying to build a prototype, lots of different platforms, WebSockets is the way to go. On the other hand, WebRTC solves a bunch of things around handling, giving you high-quality audio, high bandwidth, low latency.

But the catch is, it can be a lot more complicated to implement. And, um, it specializes, but it can be frustrating. Lots of applications use both, but for different things. So here's the TL;DR, if you only remember one thing from this talk: use WebSockets for those server-to-server use cases and small amounts of structured data and places that you want to prototype.

Use WebRTC if you're sending audio and video streams over the internet from your web app, your native, um, that's where it excels.

Kwindla Hultman Kramer4:58

So why is it so important to use WebRTC for real-time edge-to-cloud audio? A WebSocket is a TCP connection. TCP guarantees in-order delivery of network packets. If you send some data, that data is going to arrive exactly as you send it, or it's not going to arrive at all.

You put packets in your operating system's send queue, that OS queue is going to keep trying to send them until they either get hacked by the other side or your connection completely times out. And this is, in general, what you want if you're doing most network programming.

If you're making a web request, for example, this is perfect. It's not what you want if you're aiming for conversational latency.

Sean DuBois5:35

Remember that we're trying to hit a voice-to-voice latency of under one second, and ideally even better. What we want to ignore is things like the occasional packet loss. So imagine if a packet is dropped. I don't really care about what happened a second ago.

So WebRTC does clever math and buffer management that we're going to talk about more to hide where that happens.

Kwindla Hultman Kramer5:54

So this is the first and most important thing WebRTC does for you. It's all that machinery that sends packets as fast as possible, ignores packets that don't arrive inside that very tight latency budget we're operating within. You can think of it as, like, super-fast, best-effort networking.

If this were all WebRTC could do for you compared to WebSockets, it would still be worth using WebRTC just for this. Because you literally can't implement this on top of a TCP stack or on top of WebSockets. Again, the operating system is just going to try to keep sending whatever you tell it to send.

It's going to block everything if you have any packet loss or significant sort of jitter or delay in the network. And in real world, we have lots and lots of real world data on this. In the real world, this means you will get audio glitchiness or high latency or unexpected socket disconnections in 10 to 15 percent of your network connections.

Sean DuBois6:43

But WebRTC does a lot more than that. Um, if you go and you try to build the same application in WebSockets, you have to handle resampling, you have to handle packetization and doing all that. Bandwidth estimation. Networks are constantly changing and fluctuating, so you can't just send one bitrate.

Um, and you also get standard APIs for getting the stats and observability. This is all just built into WebRTC, but if you decide to do WebSockets, you have to build it yourself. So if you look at this code up on the screen, on theright side is an example of WebRTC sending one bidirectional stream of audio.

On the left side is WebSockets. And you want to spend more time building your application and less time worrying about things like sample rates. That's why you pick WebRTC.

Kwindla Hultman Kramer7:25

And this is real code using your OpenAI Realtime API, which you offer both, both options for developers.

Sean DuBois7:30

Yes.

Kwindla Hultman Kramer7:31

So I hope we've convinced you that you should use WebRTC if you're doing edge-to-cloud audio, and especially audio and video. Um, we love talking about this stuff. If you come find us later, we will talk your ear off about jitter buffers and packet management and bandwidth shaping and all that stuff.

But what we want to do now is move on and talk about a whole other category of fun stuff, which is what you can actually do with WebRTC. Um, I'll start by saying you can embed real-time audio in any app you write.

Real Uses7:46

Kwindla Hultman Kramer7:56

Any website, any iOS app, any Android app. Lots of fun embedded stuff. And the network connections will just work. You will get good audio on any device, any platform, almost any real world network connection.

Sean DuBois8:09

And I bet you use WebRTC today already. If you used Facebook Messenger, WhatsApp, Zoom, Discord, you know, any of these applications, they're using WebRTC. Um, but you didn't know that there's even more cool things happening with WebRTC. Um, I worked with a company that was doing surgery over the internet.

Um, people will telerop, um, vehicles in the field. It's super cool. Um, WebRTC is kind of the standard language of the real-time world. Um, and that's what makes it so easy that we can go build conversational intelligence on top of it.

Um, all this stuff has already been solved.

Kwindla Hultman Kramer8:41

I mean, in the new LLM era, we know lots of people who spend hours talking to computers, uh, driving their development environments with voice, doing brainstorming, treating the computer as a personal assistant, a coach, a therapist, a researcher.

Voice Future8:41

Kwindla Hultman Kramer8:54

I'm convinced voice is going to be the core building block of the next generation of UIs, of the UIs for the generative AI era. We have to do a little iteration before we figure out really what those UIs look and sound like.

Uh, but that future seems very clear to me. One of the things I say to people to kind of try to communicate how excited I am about building all this stuff is, we all lived through the last platform shift.

It was recent enough that the move for most computing from desktop to mobile happened within our memories. And so by analogy, we are in late 2007 now. We have the first iPhones, but we haven't yet invented pull to refresh.

Sean DuBois9:35

What keeps me so excited and motivated is I feel like voice is like the next, um, bicycle for the mind. You know, today we've only been able to use computers with our eyes and our hands, but now I can talk to it as well.

Um, think of all those situations where, like, you have your voice available, but you don't have your hands. Um, and the other great thing is you can keep all this computing power remotely and access it and have small devices near you.

Kwindla Hultman Kramer10:00

So on the small devices front, we brought a friend with us today. Uh, he's a little shy sometimes, so give him just a second.

Squabbert Demo10:00

Kwindla Hultman Kramer10:16

So, uh, Squabbert, can you tell people a little bit about your tech stack?

Hang on, I'm going to refresh it.

Kwindla Hultman Kramer10:23

This is Squabbert, who Chad built and Chad's daughter Ella named Squabbert.

Squabbert, can you tell the people a little bit about your tech stack?

Squabbert10:44

I'm a friendly stuffed animal running a web interface locally on a Raspberry Pi. I'm connected directly to a Python process running on a laptop, using MLX Whisper, Gemma 3, and a mildly buggy custom logit sampler that Kwin wrote.

Kwindla Hultman Kramer10:58

I mean, come on now. Like, I had a good reason that I wanted to write a logit sampler, and I, I wanted you to be able to do something that even the really amazing big LLMs running in the cloud can't do.

Do you know what I'm talking about?

Squabbert11:14

You're talking about my syllable counting talent. It's pretty neat, even if I do say so myself.

Kwindla Hultman Kramer11:19

I mean, counting syllables is like the 2025 version of how many R's are in strawberry. And, you know, Squabbert, why don't we demonstrate that a little bit? Can you make up a poem, like a four-line poem about computer programming, and use only two syllable words?

Squabbert11:41

Two. Okay, listen. Logic, coding. People, knowing. System, working. Future, growing.

Kwindla Hultman Kramer11:51

One more maybe? Uh, that was good, but I've heard you do better.

Squabbert11:58

Okay, here's another one. Pixels dancing. Colors bright. Program running, feels soright. Coding faster. Day and night, making magic. Pure delight.

Kwindla Hultman Kramer12:06

You forgot about the two syllable words thing, but we'll, we'll, we'll let it go. Definitely the first one had two syllable words. I think, I think we should say goodbye just in the interest of time, Squabbert. But thank you so much.

Yashin12:17

Woo!

Kwindla Hultman Kramer12:23

I think Squabbert did really, really well. Um, so that, that is a Raspberry Pi connected with a peer-to-peer WebRTC connection directly to my laptop over the same local area network.

Sean DuBois12:37

So that was the serverless WebRTC connection. Squabbert's talking directly to the laptop. But what's super cool about WebRTC is you have all these different choices to how you want to connect to things. So you could do this local connection, or something like Squabbert could connect to a server, um, up running on another in the cloud and do all the AI stuff.

Connection Paths12:37

Sean DuBois12:55

And then the third option is you can go and connect up to something like Pipecat and make it multi-party. So bring LLMs into meetings or other places like that. Um, super awesome how, like, you get all this flexibility to build these things the way that, like, matches your application.

So we want to close the video with actually a builder who's in the audience. Um, super excited. And, um, what's the best part about this is that, um, when she built this, she had never written any code before this.

Yashin's App13:10

Sean DuBois13:21

And so what makes me so excited about the future of voice is if we can make this easy enough, people that have really innovative inspirational ideas can go and do stuff themselves.

Yashin13:33

Hi. I'm Yashin. I'm a mom of two bilingual kids. I'm raising my kids bilingual because I want them to connect with my cultural roots. And that's a wish shared by many parents of bilingual children. But raising bilingual kids is hard.

It's expensive, it's time-consuming, and it often feels like a big chore on both the kids and parents. I believe we can do a lot better. With today's technology, I really believe we can make language education feel more natural and even fun for the kids.

Here's a quick clip of what I've been working on.

Hi there, buddy. Ready to have some fun today? 你好呀 , 小朋友 。 今天我们一起玩 , 好不好 ?

Yashin14:19

好 。

太好了 。 我们可以学一些简单的词汇 。 How about we start with something fun? Let's say hello. In Mandarin, we say 你好 . Can you say 你好 ?

Yashin14:34

你好 . Okay. It's still early. But I'm excited about what's possible. I'm not technical, but with just a little bit of guidance from a kind member of this community, I was able to bring this first version to life.

I already have a group of eager testers, mostly parents like me, who are excited to try this. If this also sounds exciting to you, I would love to connect. Thanks so much for watching, and let's build this future together.

Closing15:13

Kwindla Hultman Kramer15:13

So we put the QR code up here for Yashin's project, which I absolutely love. She's here, uh, if you're interested. If you're here in the audience and, uh, you're interested in multilingual stuff or building these kind of things, please find her.

It's such a great, great, great thing. Um, if you're watching on YouTube, here's the QR code. Sean and I are super excited about these kind of projects. And, I mean, the, the person Yashin shouted out in the video is, of course, Sean, who has done more than anybody I know to make WebRTC accessible to everyone.

Sean DuBois15:43

Um, the idea that we want to leave you with is if you have an idea in voice AI and WebRTC, whether you've been a programmer for years or you're just getting started, like, we are here to support you.

We're so excited. And we believe that things like Pipecat and LibPeer, like, if we can make this easier, we're going to see the next generation of really exciting, innovative projects.

Kwindla Hultman Kramer16:03

So come find us in the hallway or online. We hang out on Discord and Twitter and LinkedIn.

Sean DuBois16:08

And, uh, here are the resources that you can scan and, uh, hopefully they find helpful. So Kwin wrote an amazing book that I believe is in your bag. Um, and then, uh, yeah. So thank you so much.

Kwindla Hultman Kramer16:19

We can't wait to see what you build.

Yashin16:21

Woo!