AIAI EngineerJun 27, 2025· 21:43

Milliseconds to Magic: Real‑Time Workflows using the Gemini Live API and Pipecat

Kwindla Kramer (Daily, Pipecat) and Shrestha Basu Mallick (Google DeepMind, Gemini API) argue that voice is the most natural interface and that the Gemini Live API combined with the Pipecat framework enables developers to build real-time multimodal voice agents covering the full stack from models to application code. They demo a voice-driven task management app that handles grocery, reading, and work lists, showing impressive context-aware tool use (e.g., consolidating lists, web searching for 'Dream Count' author) but also jagged edge limitations like turn detection errors and persistent misspelling of 'Kwin.' They discuss how capabilities like turn detection migrate down the stack over time. The episode also covers proactivity, multilinguality, telephony integration, and experimental native audio models for emotive, steerable dialogue.

  1. 0:00Voice Magic
  2. 3:08The Stack
  3. 6:57Demo Prep
  4. 8:50Live Demo
  5. 17:16Demo Debrief
  6. 19:22Grandmother's Knot

Powered by PodHood

Transcript

Voice Magic0:00

Shrestha Basu Mallick0:16

So voice is the most natural of interfaces. Humans are storytellers, talkers, listeners, conversationalists. We think out aloud. We learn to talk before we learn to read. And most of us talk faster than we type. We express emotion through our voices, and we use sound to understand the world around us.

Kwindla Hultman Kramer0:46

We've been working together for the past few months. Shrestha, from the angle of models and APIs, and me from the application layer and agent framework direction. And I think we both believe that voice is a critical and universal building block for the whole next generation of Gen AI, especially at the UI level, but more generally as well.

I mean, those of us who are early adopters of personal voice AI talk to our computers all the time. We think of the LLMs we talk to as sounding boards and coaches and interfaces to everything that lives on our devices and in the cloud.

Shrestha Basu Mallick1:19

And this is not just an early adopter phenomenon,right? Like, we already have voice agents deployed at scale. Language translation apps that translate between a patient and a doctor. Directed learning apps that a fourth grader can use to learn a topic they want to.

Speech therapy apps. And copilots that help people navigate complex enterprise software.

Kwindla Hultman Kramer1:44

One of the things we see in our work with customers at Daily is it's pretty common for people not to realize that they're talking to a voice agent on a phone call, even when you tell them at the beginning of the phone call that they're talking to an AI.

Shrestha Basu Mallick1:57

Yeah, and kids born today will probably take all of this for granted. But those of us who are living through this evolution of talking computers, this can sometimes feel like magic. But of course, anybody who's seen a really great magician prepare a magic trick knows that the magic is just the interface.

There's a lot of hard work that goes into creating that magic trick.

Kwindla Hultman Kramer2:25

So here's a partial list of the hard things that, doneright, collectively add up to that magic. So real-time responsiveness, which we've all, in this whole track all day, talked about as the foundation thing you have to getright or voice AI is unworkable, through the things that we're just starting to experiment with, like generating dynamic user interface elements for every conversational turn.

These are the things we've been hacking on and thinking about together for the past few months. And we're not going to go over all of these today, although we did have a little extra time in the session,right, Thor?

Thor said we could talk for, like, a couple hours, maybe.

Shrestha Basu Mallick2:57

Yeah.

Kwindla Hultman Kramer2:57

But we do have a framework that we thought would be useful to share with you. A framework that sort of maps onto how we've worked together from the model layer all the way up.

Shrestha Basu Mallick3:08

Yeah, and this barely scratches the surface, but here are the layers of the voice AI stack. So at the bottom, underpinning everything, you have the large language models that frontier labs like DeepMind work on. Then above that, you have carefully designed, but at this stage, constantly evolving, real-time APIs.

The Stack3:08

Shrestha Basu Mallick3:29

Google's version is called the Gemini Live API.

Above the APIs are the orchestration libraries and frameworks like Pipecat that help to manage and abstract away the complexity of building these real-time multimodal applications. And then, of course, at the top of the stack, you have the application code.

Kwindla Hultman Kramer3:52

For each of the hard things we listed on the previous slide, the code that implements that hard thing lives somewhere in that stack. So one of the ways we think about this is that there's a map, and you can sort of think about it two-dimensionally, maybe.

There's the where does the code live that kind of solves the hard problem that you're thinking about as a voice agent developer, where in the stack, and then how mature is our solution to thatright now.

Shrestha Basu Mallick4:17

Yeah, basically, how solved is this thing? And what we've tried to do here is map all of these various things that you need to getright on aright-to-left axis of maturity.

Kwindla Hultman Kramer4:29

And there are a couple of things that are kind of top of mind for me about this mapping. One is that I don't think of any of these things as more than about 50% solved. Totally arbitrary, like, personal thing.

Shrestha and I argued about it a little bit, like, what's theright way to represent that on this slide. But what we're trying to say is basically it's early. It's early for voice AI, and there's a lot of work to do at every part of the stack to get to that universal voice UI we're imagining.

Shrestha Basu Mallick4:55

Yeah, and secondly, as this technology matures, and we've already seen some of this happening, the capabilities tend to move down the stack. So what might happen is, in your one-off individual applications, you might write some code to solve a specifically difficult challenge.

Now, if enough people experience that challenge, then that tends to get built into the orchestration libraries and frameworks, and then eventually make its way into the APIs. But independently of all of that, the models themselves are getting more and more generally capable.

I mean, we just talked about semantic voice activity detection in the previous talk, so.

Kwindla Hultman Kramer5:39

Yeah, this is, like, a great follow-on to Tom's talk about turn detection, because I think turn detection is a perfect example of this. So, like, I built some of the first talk-to-an-LLM voice AI applications a little over two years ago now, and I tried to solve turn detectionright there in the application code, because there weren't any tools yet for it.

A few months later, we built what we thought were pretty generalized at the time state-of-the-art turn detection implementations into Pipecat, so moved down the layer into the framework. Now Shrestha has turn detection in the multimodal live APIs sort of inside the surface area of those same APIs that are doing inference and other things for you.

And I think all of us, as Tom said, expect the models over time to just do turn detection for us. And all those hard things, it varies depending on exactly what you're talking about of that long list we put together on that slide.

But in general, I think everything is moving down the stack, and then more and more interesting use cases are creating more things to put sort of at the top of the stack.

Shrestha Basu Mallick6:41

Yeah, I will say we have server-side turn detection built in, but we also allow you to turn off turn detection and use models like Daily and LiveKit. So should we start with the demo?

Kwindla Hultman Kramer6:57

Yeah, we can. We do have a demo to show you. And it's sort of a demo of some stuff I've been using in my own life every day for the last year or so. I've been experimenting with talking to my computer and my phone as much as I can to do various things, as you can imagine, because I post about it probably too often on social media.

Demo Prep6:57

Kwindla Hultman Kramer7:17

And one of the things that's become really sticky for me is kind of managing my priorities in, like, a very loose way with voice AI. These days, the conversations I have with my phone and my computer often tend to be in this sort of funny and interesting and new gray area for me between, like, a to-do list and brainstorming or kind of trying to work stuff out.

And because I write this code for myself and am always messing with it, this is definitely a cobbler's children never wear shoes situation. Like, I don't have unit tests. I don't have evals for this code. I'm always using, like, Shrestha's latest models.

I'm always using some bleeding-edge branch of Pipecat. So what I thought would be fun to do is fire up a version of this code, have Shrestha talk to it, and see what breaks. And then maybe we can talk about, like, where those bugs are and, like, how I could have done better next time.

Shrestha Basu Mallick8:11

Yeah, I will say, having played with this, there's quite a lot of variation in how the model behaves. So we'll see where we get today.

Kwindla Hultman Kramer8:19

One thing that I think is interesting is, as I write more and more of this code, I keep relearning the lesson that the way the models drive the whole application cycle is very different from my traditional programmer's view of things.

So half the things I'm sure you encountered when you were, like, playing with this are the model did something that my code actually didn't expect. That might have been bad, but it might actually have been good, and that's a whole sort of adjustment to how we think about these applications.

Shrestha Basu Mallick8:46

Yeah, there's definitely pleasant surprises also.

Live Demo8:50

Speaker 38:50

I am ready.

Shrestha Basu Mallick8:52

OK. Do I talk?

Kwindla Hultman Kramer8:55

I think you can just talk. And I don't think I can mute it, because I don't know how to use the mouse backwards very well.

Shrestha Basu Mallick9:01

OK.

Speaker 39:02

Yeah.

Shrestha Basu Mallick9:02

I just need to.

Speaker 39:03

You can talk whenever you're ready. No problem. I'm here to help whenever you need. Yep, I can hear you loud and clear.

Shrestha Basu Mallick9:11

Allright. OK, so.

Speaker 39:15

What can I help you with today?

Shrestha Basu Mallick9:16

Yes. Can you help me create a grocery list for asparagus pizza?

Speaker 39:24

Got it. Pizza ingredients are added to your girthy list. Anything else?

Shrestha Basu Mallick9:28

No. Can you list out the pizza ingredients for asparagus pizza, though?

Speaker 39:35

Sure. I've updated the grocery list to include pizza crust, mozzarella cheese, tomato sauce, and asparagus.

Shrestha Basu Mallick9:43

Anything else?

Would you add anything else to the list?

Speaker 39:52

I've added garlic and olive oil to the list. Does that soundright?

Shrestha Basu Mallick9:56

That sounds perfect. Allright. Now let's create my reading list. So.

Speaker 310:01

Great.

Shrestha Basu Mallick10:04

Can you create.

Speaker 310:05

I'm just going to start adding to your reading list.

Shrestha Basu Mallick10:06

OK, great.

Speaker 310:07

Absolutely. What books would you like to have?

Shrestha Basu Mallick10:09

Turn detection, yeah. OK, reading list.

Can you add the please add the first book, "Dream Count," and can you look up the author?

Speaker 310:20

OK, I've added "Quick" to your reading list.

Shrestha Basu Mallick10:22

No, that was not the book I wanted to add. Please add the first book, "Dream Count," and look up the author for me.

Hey, are you there?

Speaker 310:40

Got

it. "Segmentation Faults" is added to your reading list. What's next?

Shrestha Basu Mallick10:49

That is I guess I'll have to

it clearly couldn't find the author for that one. But let's try one last time. OK, one more book, "Dream Count," and can you look up the author?

Hey, are you there?

Speaker 311:18

OK, "Segmentation" has been added to your reading list.

Shrestha Basu Mallick11:22

It just likes your voice better than mine. Allright. No, "Segmentation Fault" is great, but I would really like to read "Dream Count." Can you add "Dream Count"?

Allright. Hey, are you still there?

Speaker 311:44

Yes, I'm still here. How can I help you?

Shrestha Basu Mallick11:46

OK, can you add the second and third books of the three-body problem trilogy?

Speaker 311:54

Sure. The second and third. Both "Segmentation Fault" and. While "Segmentation" and.

Shrestha Basu Mallick11:59

Can you add the second and third books of the three-body.

Speaker 312:03

Concept. Is there anything else I can help you with?

Shrestha Basu Mallick12:07

I swear it worked much better 10 times last night.

Speaker 312:13

I apologize for my performance is enough to part today. Is there anything specific you'd like me to try again?

Shrestha Basu Mallick12:20

Well, maybe you can try allright, but let's recover here. Please add the second and third books of the three-body problem trilogy.

Speaker 312:36

You want to add the second and third books of the three-body problem trilogy.

Shrestha Basu Mallick12:40

I just want to point out that it did look up the names and find it. Let's see if it'll cooperate. Yeah. This one might be in its training data, because it's a bit older. But the "Dream Count" one, which it was spectacular on last night, it did a Google search, because it's a new book.

Allright, let's move on. Can you create a new list for me called "Work Tasks"? So first item, I would like to say, create H2 Roadmap by end of day Friday.

Speaker 313:18

Got it. "Create H2 Roadmap by end of day Friday" is added to your work list.

Shrestha Basu Mallick13:22

I don't see it. Can you show me on the screen?

Speaker 313:27

I've displayed the work list for you.

Shrestha Basu Mallick13:29

Allright. Let's add another item.

Speaker 313:33

OK, what's the next item?

Shrestha Basu Mallick13:35

Finish writing podcast script by end of day Thursday, and can you add the exact date?

Speaker 313:42

OK, I've added "Finish writing podcast script by end of day Thursday" to your work list.

Shrestha Basu Mallick13:47

What's the exact date for end of day Thursday?

Speaker 313:52

Today is Wednesday, June 4, 2025. So end of day Thursday is June 5, 2025.

Shrestha Basu Mallick13:59

Can you add that to the list?

Speaker 314:02

Yes, I've added the exact date to the work list.

Shrestha Basu Mallick14:06

Allright.

Speaker 314:06

It now says "Finish writing podcast script by end of day Thursday, June 5, 2025."

Shrestha Basu Mallick14:11

Perfect.

Speaker 314:11

Is there anything else you'd like to add to the list?

Shrestha Basu Mallick14:14

Yes, actually.

Speaker 314:14

Or anything else?

Shrestha Basu Mallick14:15

Yes, I would. Can you combine my grocery list, my reading list, and my work list and split them into two lists, personal tasks and work tasks?

Speaker 314:29

I've combined your grocery list, reading list, and work list. They are now displayed on the screen as separate lists.

Shrestha Basu Mallick14:37

Why are garlic oil garlic and olive oil separate from pizza ingredients? Can you flesh out the grocery list, please?

Speaker 314:47

You'reright. I've updated the grocery list to include garlic and olive oil as part of the pizza ingredients.

Shrestha Basu Mallick14:53

Allright, and can you consolidate the work list? I see two on the screen.

Speaker 314:59

You'reright. I apologize for the duplication. I've consolidated the work list to a single entry. You should now see the updated list on the screen.

Shrestha Basu Mallick15:08

OK, and then finally, can you do the following? Can you assign grocery list to Kwin, reading list to Shrestha, with the exception of "Segmentation Fault," which Kwin will read, and work list to Shrestha?

Speaker 315:25

I understand you want to assign the grocery list to Kwin, the reading list to Shrestha, except for "Segmentation Fault," which Kwin will read, and the work list to Shrestha as well.

Shrestha Basu Mallick15:35

Can you show that on the screen?

Speaker 315:39

I've updated the lists to show the assignments he specified. You should now see the updated lists on the screen.

Shrestha Basu Mallick15:45

And you've spelled my name correctly, which was the harder name, but you spelled Kwin incorrectly. But OK, we're out of time here. OK, so if we want the final thing here is we would like to have a nice background for this app.

So can you.

Speaker 316:04

I apologize for spelling your name.

Shrestha Basu Mallick16:07

OK, could you help us generate actually, I'll be more assertive. Please generate an app that has the words "Hello World" jumping up and down on the screen with the colors of Google. And then we also want two neon green ASCII cats which are animated.

Speaker 316:32

I'm now generating an app with "Hello World" jumping up and down in Google colors and two animated neon green ASCII cats. It may take a moment for the app to appear on your screen.

Kwindla Hultman Kramer16:52

Detention.

Shrestha Basu Mallick17:08

So yeah, there you have the Google colors and the cats from Pipecat. So with that, I'll hand it over to Kwin.

Demo Debrief17:16

Kwindla Hultman Kramer17:16

So first, thank you or you should go back to your pride of place,right in the middle. Thank you for being such a good sport. Very messy code on my part, including things like basically no instructions to the LLM about how to display text on the screen and just telling it it has a function that can display text on the screen, and it sort of guesses and learns in context, as you can tell from Shrestha, about when it should clear the screen, because there's an optional clear argument to the add text to the screen function.

And it's super impressive, but also super jagged frontier about whether it kind of can intuit what you want to do in those contexts. So thank you for, like, doing this, because this is what I do all the time with this code, trying to figure out, like, what these models can do and what kind of code you have to write and what you don't have to scaffold for them to do well.

Shrestha Basu Mallick18:03

Yeah, and it's been, you know, playing with this. Every turn is different. And it's interesting to see the things that it struggles with, like your name. Even if I spell out the exact letters, it somehow really wants to spell Kwin the way it spells.

I think it also I mean, turn detection, as we saw, there's a lot of work that can be done, of course, there. And I'm trying to remember. There's and there's, of course, a lot of variation in sometimes here.

Like, there are times when it gets the grocery list perfect and, you know, combines the list perfectly, and sometimes it's a bit in the middle, like here.

Kwindla Hultman Kramer18:39

And the way this code works is it just for a given, like, session, it loads lots and lots and lots of previous conversational sessions in user assistant, user assistant sort of messages. Does sometimes, depending on the version of the code I've, like, got running, it summarizes a little bit.

Sometimes it doesn't. So we really are leaning on the intelligence of the LLM to do all of the sort of contextual understanding about what we mean by a list, what we mean by the context in which we are talking about that list.

It is super amazing that it works at all, basically, in my mind. And it's all voice-driven, and it's all multimodal from the ground up. We have a whole nother video we can show, but I definitely think we're out of time.

So we will.

Guest19:19

You have a final talk, and everyone seems to be excited, so.

Grandmother's Knot19:22

Kwindla Hultman Kramer19:22

Can't you one minute of.

Shrestha Basu Mallick19:24

Maybe we should talk about our grandmothers?

Kwindla Hultman Kramer19:25

Oh, yes. I totally forgot that part. Sorry. Let's skip past the demo where it gets the grocery list perfect.

Shrestha Basu Mallick19:31

I think maybe this crowd would like to see that demo.

Kwindla Hultman Kramer19:34

That was great. So this has been fun for me to work on, because, like, it's so relevant to my everyday life. But in Shrestha, and we're talking about it, and I think there's actually something else that really kind of hooked me that she said.

Shrestha Basu Mallick19:46

Yeah, so, you know, my grandmother was Indian, of course, and she used to wear this cloth garment called a sari. And her way of reminding herself when she had to do things was tying knots on the sari. Of course, and then I was chatting with Kwin, and what was incredible is apparently his grandmother in North Carolina, so very different from Calcutta in India, used to tie strings around her fingers.

And firstly, you know, this is kind of incredible. You know, no matter how many continents separate us, like, smart people come up with the same generally intelligent patterns. But it's also incredible how technology allows humans to evolve. Now, the one problem with either the knots or the strings is you knew you had to remember something, but you didn't know what it was.

So you still relied on your memory. And, you know, ultimately, that's why I do the work I do at Google, because I want to build the technologies that enable, you know, an infinite world of creative possibilities tomorrow or even today across continents.

And I just want to say that we believe that voice is the most natural of interfaces, and there will come a world where most of the interaction with language models will happen via voice. And the Gemini models are trained to be multimodal from the ground up.

So of course, they ingest text, voice, but also images and video. So if you have any questions about Gemini, please reach out to me on X, on LinkedIn, email, wherever. Happy to work with builders like yourself.

Kwindla Hultman Kramer21:28

Yeah, thanks for coming to the talk, and we would love to see what you build with these models and APIs.