# Milliseconds to Magic: Real‑Time Workflows using the Gemini Live API and Pipecat

AI Engineer · 2025-06-27

<https://aie.addtry.com/1a94625b-6062-4b31-a4d2-620f2b696a71>

Kwindla Kramer (Daily, Pipecat) and Shrestha Basu Mallick (Google DeepMind, Gemini API) argue that voice is the most natural interface and that the Gemini Live API combined with the Pipecat framework enables developers to build real-time multimodal voice agents covering the full stack from models to application code. They demo a voice-driven task management app that handles grocery, reading, and work lists, showing impressive context-aware tool use (e.g., consolidating lists, web searching for 'Dream Count' author) but also jagged edge limitations like turn detection errors and persistent misspelling of 'Kwin.' They discuss how capabilities like turn detection migrate down the stack over time. The episode also covers proactivity, multilinguality, telephony integration, and experimental native audio models for emotive, steerable dialogue.

## Questions this episode answers

### How do voice AI capabilities mature and shift down the technology stack?

Kwindla Kramer and Shrestha Basu Mallick present a maturity mapping where capabilities like turn detection, real-time responsiveness, and dynamic UI generation are at best 50% solved. They explain a pattern: these capabilities first appear in custom application code, then get incorporated into orchestration frameworks like Pipecat, and later become features of real-time APIs like Gemini Live. The ultimate goal is for models to handle them natively. Turn detection exemplifies this trajectory, having moved from hand-coded solutions to API-level integration.

[4:04](https://aie.addtry.com/1a94625b-6062-4b31-a4d2-620f2b696a71?t=244000)

### How does the Gemini Live API handle turn detection, and can I use my own?

Shrestha Basu Mallick states that the Gemini Live API has server-side turn detection built in, but developers can disable it and use external systems like Daily or LiveKit if they prefer. In the demo, the built-in turn detection sometimes cut off responses or misheard names, while Kwindla Kramer noted that Pipecat originally implemented state-of-the-art turn detection before such features moved into the API.

[5:39](https://aie.addtry.com/1a94625b-6062-4b31-a4d2-620f2b696a71?t=339000)

### What happened when Shrestha Basu Mallick tested a personal task management voice agent built with Gemini Live and Pipecat?

Shrestha Basu Mallick tested a voice agent that created grocery, reading, and work lists, searched for book authors, added tasks with specific dates, and combined lists. It even generated a 'Hello World' app with animated ASCII cats on command. Turn detection misfired multiple times and the agent misspelled 'Kwin' despite correction, highlighting its impressive context awareness but also the jagged frontier of current voice AI.

[8:50](https://aie.addtry.com/1a94625b-6062-4b31-a4d2-620f2b696a71?t=530000)

## Key moments

- **[0:00] Voice Magic**
  - [1:19] Voice agents are already deployed at scale in language translation, learning apps, speech therapy, and enterprise copilots.
- **[3:08] The Stack**
  - [4:29] Kwindla Kramer: 'I don't think of any of these things as more than about 50% solved.'
  - [4:55] Voice AI capabilities move from application code to frameworks, APIs, and eventually into models themselves.
  - [5:39] Turn detection for voice AI has moved from application code to Pipecat framework to Gemini Live API, and will eventually be handled by models.
- **[6:57] Demo Prep**
  - [6:57] Kwindla Kramer has been using a self-built voice AI daily for a year to manage priorities and brainstorming.
- **[8:50] Live Demo**
  - [10:44] Gemini Live repeatedly mishears 'Dream Count' as 'Segmentation Faults' during the demo, highlighting voice recognition challenges.
  - [12:56] Gemini Live successfully looks up 'Three-Body Problem' from training data but struggles with new book 'Dream Count' that requires a Google search.
  - [13:52] Gemini Live accurately infers that 'end of day Thursday' means June 5, 2025, when context says today is June 4.
  - [16:04] Demo: Gemini Live generates an animated app with 'Hello World' in Google colors and ASCII cats from a voice command.
- **[17:16] Demo Debrief**
  - [17:16] Kwindla Kramer explains his voice agent code had no display instructions; the LLM learned to show text in context.
- **[19:22] Grandmother's Knot**
  - [19:46] Shrestha’s Indian grandmother tied knots on a sari for reminders; Kwin’s North Carolina grandmother tied finger strings — same pattern across continents.
  - [21:12] Gemini models are trained to be multimodal from the ground up, ingesting text, voice, images, and video.

## Speakers

- **Kwindla Hultman Kramer** (guest)
- **Shrestha Basu Mallick** (guest)

## Topics

Voice Agents, Modular AI Orchestration

## Mentioned

Daily (company), DeepMind (company), Google (company), Gemini Live API (product), LiveKit (product), Pipecat (product)

## Transcript

### Voice Magic

**Shrestha Basu Mallick** [0:16]
So voice is the most natural of interfaces. Humans are storytellers, talkers, listeners, conversationalists. We think out aloud. We learn to talk before we learn to read. And most of us talk faster than we type. We express emotion through our voices, and we use sound to understand the world around us.

**Kwindla Hultman Kramer** [0:46]
We've been working together for the past few months. Shrestha, from the angle of models and APIs, and me from the application layer and agent framework direction. And I think we both believe that voice is a critical and universal building block for the whole next generation of Gen AI, especially at the UI level, but more generally as well.

I mean, those of us who are early adopters of personal voice AI talk to our computers all the time. We think of the LLMs we talk to as sounding boards and coaches and interfaces to everything that lives on our devices and in the cloud.

**Shrestha Basu Mallick** [1:19]
And this is not just an early adopter phenomenon,right? Like, we already have voice agents deployed at scale. Language translation apps that translate between a patient and a doctor. Directed learning apps that a fourth grader can use to learn a topic they want to.

Speech therapy apps. And copilots that help people navigate complex enterprise software.

**Kwindla Hultman Kramer** [1:44]
One of the things we see in our work with customers at Daily is it's pretty common for people not to realize that they're talking to a voice agent on a phone call, even when you tell them at the beginning of the phone call that they're talking to an AI.

**Shrestha Basu Mallick** [1:57]
Yeah, and kids born today will probably take all of this for granted. But those of us who are living through this evolution of talking computers, this can sometimes feel like magic. But of course, anybody who's seen a really great magician prepare a magic trick knows that the magic is just the interface.

There's a lot of hard work that goes into creating that magic trick.

**Kwindla Hultman Kramer** [2:25]
So here's a partial list of the hard things that, doneright, collectively add up to that magic. So real-time responsiveness, which we've all, in this whole track all day, talked about as the foundation thing you have to getright or voice AI is unworkable, through the things that we're just starting to experiment with, like generating dynamic user interface elements for every conversational turn.

These are the things we've been hacking on and thinking about together for the past few months. And we're not going to go over all of these today, although we did have a little extra time in the session,right, Thor?

Thor said we could talk for, like, a couple hours, maybe.

**Shrestha Basu Mallick** [2:57]
Yeah.

**Kwindla Hultman Kramer** [2:57]
But we do have a framework that we thought would be useful to share with you. A framework that sort of maps onto how we've worked together from the model layer all the way up.

**Shrestha Basu Mallick** [3:08]
Yeah, and this barely scratches the surface, but here are the layers of the voice AI stack. So at the bottom, underpinning everything, you have the large language models that frontier labs like DeepMind work on. Then above that, you have carefully designed, but at this stage, constantly evolving, real-time APIs.

### The Stack

**Shrestha Basu Mallick** [3:29]
Google's version is called the Gemini Live API.

Above the APIs are the orchestration libraries and frameworks like Pipecat that help to manage and abstract away the complexity of building these real-time multimodal applications. And then, of course, at the top of the stack, you have the application code.

**Kwindla Hultman Kramer** [3:52]
For each of the hard things we listed on the previous slide, the code that implements that hard thing lives somewhere in that stack. So one of the ways we think about this is that there's a map, and you can sort of think about it two-dimensionally, maybe.

There's the where does the code live that kind of solves the hard problem that you're thinking about as a voice agent developer, where in the stack, and then how mature is our solution to thatright now.

**Shrestha Basu Mallick** [4:17]
Yeah, basically, how solved is this thing? And what we've tried to do here is map all of these various things that you need to getright on aright-to-left axis of maturity.

**Kwindla Hultman Kramer** [4:29]
And there are a couple of things that are kind of top of mind for me about this mapping. One is that I don't think of any of these things as more than about 50% solved. Totally arbitrary, like, personal thing.

Shrestha and I argued about it a little bit, like, what's theright way to represent that on this slide. But what we're trying to say is basically it's early. It's early for voice AI, and there's a lot of work to do at every part of the stack to get to that universal voice UI we're imagining.

**Shrestha Basu Mallick** [4:55]
Yeah, and secondly, as this technology matures, and we've already seen some of this happening, the capabilities tend to move down the stack. So what might happen is, in your one-off individual applications, you might write some code to solve a specifically difficult challenge.

Now, if enough people experience that challenge, then that tends to get built into the orchestration libraries and frameworks, and then eventually make its way into the APIs. But independently of all of that, the models themselves are getting more and more generally capable.

I mean, we just talked about semantic voice activity detection in the previous talk, so.

**Kwindla Hultman Kramer** [5:39]
Yeah, this is, like, a great follow-on to Tom's talk about turn detection, because I think turn detection is a perfect example of this. So, like, I built some of the first talk-to-an-LLM voice AI applications a little over two years ago now, and I tried to solve turn detectionright there in the application code, because there weren't any tools yet for it.

A few months later, we built what we thought were pretty generalized at the time state-of-the-art turn detection implementations into Pipecat, so moved down the layer into the framework. Now Shrestha has turn detection in the multimodal live APIs sort of inside the surface area of those same APIs that are doing inference and other things for you.

And I think all of us, as Tom said, expect the models over time to just do turn detection for us. And all those hard things, it varies depending on exactly what you're talking about of that long list we put together on that slide.

But in general, I think everything is moving down the stack, and then more and more interesting use cases are creating more things to put sort of at the top of the stack.

**Shrestha Basu Mallick** [6:41]
Yeah, I will say we have server-side turn detection built in, but we also allow you to turn off turn detection and use models like Daily and LiveKit. So should we start with the demo?

**Kwindla Hultman Kramer** [6:57]
Yeah, we can. We do have a demo to show you. And it's sort of a demo of some stuff I've been using in my own life every day for the last year or so. I've been experimenting with talking to my computer and my phone as much as I can to do various things, as you can imagine, because I post about it probably too often on social media.

### Demo Prep

**Kwindla Hultman Kramer** [7:17]
And one of the things that's become really sticky for me is kind of managing my priorities in, like, a very loose way with voice AI. These days, the conversations I have with my phone and my computer often tend to be in this sort of funny and interesting and new gray area for me between, like, a to-do list and brainstorming or kind of trying to work stuff out.

And because I write this code for myself and am always messing with it, this is definitely a cobbler's children never wear shoes situation. Like, I don't have unit tests. I don't have evals for this code. I'm always using, like, Shrestha's latest models.

I'm always using some bleeding-edge branch of Pipecat. So what I thought would be fun to do is fire up a version of this code, have Shrestha talk to it, and see what breaks. And then maybe we can talk about, like, where those bugs are and, like, how I could have done better next time.

**Shrestha Basu Mallick** [8:11]
Yeah, I will say, having played with this, there's quite a lot of variation in how the model behaves. So we'll see where we get today.

**Kwindla Hultman Kramer** [8:19]
One thing that I think is interesting is, as I write more and more of this code, I keep relearning the lesson that the way the models drive the whole application cycle is very different from my traditional programmer's view of things.

So half the things I'm sure you encountered when you were, like, playing with this are the model did something that my code actually didn't expect. That might have been bad, but it might actually have been good, and that's a whole sort of adjustment to how we think about these applications.

**Shrestha Basu Mallick** [8:46]
Yeah, there's definitely pleasant surprises also.

### Live Demo

**Speaker 3** [8:50]
I am ready.

**Shrestha Basu Mallick** [8:52]
OK. Do I talk?

**Kwindla Hultman Kramer** [8:55]
I think you can just talk. And I don't think I can mute it, because I don't know how to use the mouse backwards very well.

**Shrestha Basu Mallick** [9:01]
OK.

**Speaker 3** [9:02]
Yeah.

**Shrestha Basu Mallick** [9:02]
I just need to.

**Speaker 3** [9:03]
You can talk whenever you're ready. No problem. I'm here to help whenever you need. Yep, I can hear you loud and clear.

**Shrestha Basu Mallick** [9:11]
Allright. OK, so.

**Speaker 3** [9:15]
What can I help you with today?

**Shrestha Basu Mallick** [9:16]
Yes. Can you help me create a grocery list for asparagus pizza?

**Speaker 3** [9:24]
Got it. Pizza ingredients are added to your girthy list. Anything else?

**Shrestha Basu Mallick** [9:28]
No. Can you list out the pizza ingredients for asparagus pizza, though?

**Speaker 3** [9:35]
Sure. I've updated the grocery list to include pizza crust, mozzarella cheese, tomato sauce, and asparagus.

**Shrestha Basu Mallick** [9:43]
Anything else?

Would you add anything else to the list?

**Speaker 3** [9:52]
I've added garlic and olive oil to the list. Does that soundright?

**Shrestha Basu Mallick** [9:56]
That sounds perfect. Allright. Now let's create my reading list. So.

**Speaker 3** [10:01]
Great.

**Shrestha Basu Mallick** [10:04]
Can you create.

**Speaker 3** [10:05]
I'm just going to start adding to your reading list.

**Shrestha Basu Mallick** [10:06]
OK, great.

**Speaker 3** [10:07]
Absolutely. What books would you like to have?

**Shrestha Basu Mallick** [10:09]
Turn detection, yeah. OK, reading list.

Can you add the please add the first book, "Dream Count," and can you look up the author?

**Speaker 3** [10:20]
OK, I've added "Quick" to your reading list.

**Shrestha Basu Mallick** [10:22]
No, that was not the book I wanted to add. Please add the first book, "Dream Count," and look up the author for me.

Hey, are you there?

**Speaker 3** [10:40]
Got

it. "Segmentation Faults" is added to your reading list. What's next?

**Shrestha Basu Mallick** [10:49]
That is I guess I'll have to

it clearly couldn't find the author for that one. But let's try one last time. OK, one more book, "Dream Count," and can you look up the author?

Hey, are you there?

**Speaker 3** [11:18]
OK, "Segmentation" has been added to your reading list.

**Shrestha Basu Mallick** [11:22]
It just likes your voice better than mine. Allright. No, "Segmentation Fault" is great, but I would really like to read "Dream Count." Can you add "Dream Count"?

Allright. Hey, are you still there?

**Speaker 3** [11:44]
Yes, I'm still here. How can I help you?

**Shrestha Basu Mallick** [11:46]
OK, can you add the second and third books of the three-body problem trilogy?

**Speaker 3** [11:54]
Sure. The second and third. Both "Segmentation Fault" and. While "Segmentation" and.

**Shrestha Basu Mallick** [11:59]
Can you add the second and third books of the three-body.

**Speaker 3** [12:03]
Concept. Is there anything else I can help you with?

**Shrestha Basu Mallick** [12:07]
I swear it worked much better 10 times last night.

**Speaker 3** [12:13]
I apologize for my performance is enough to part today. Is there anything specific you'd like me to try again?

**Shrestha Basu Mallick** [12:20]
Well, maybe you can try allright, but let's recover here. Please add the second and third books of the three-body problem trilogy.

**Speaker 3** [12:36]
You want to add the second and third books of the three-body problem trilogy.

**Shrestha Basu Mallick** [12:40]
I just want to point out that it did look up the names and find it. Let's see if it'll cooperate. Yeah. This one might be in its training data, because it's a bit older. But the "Dream Count" one, which it was spectacular on last night, it did a Google search, because it's a new book.

Allright, let's move on. Can you create a new list for me called "Work Tasks"? So first item, I would like to say, create H2 Roadmap by end of day Friday.

**Speaker 3** [13:18]
Got it. "Create H2 Roadmap by end of day Friday" is added to your work list.

**Shrestha Basu Mallick** [13:22]
I don't see it. Can you show me on the screen?

**Speaker 3** [13:27]
I've displayed the work list for you.

**Shrestha Basu Mallick** [13:29]
Allright. Let's add another item.

**Speaker 3** [13:33]
OK, what's the next item?

**Shrestha Basu Mallick** [13:35]
Finish writing podcast script by end of day Thursday, and can you add the exact date?

**Speaker 3** [13:42]
OK, I've added "Finish writing podcast script by end of day Thursday" to your work list.

**Shrestha Basu Mallick** [13:47]
What's the exact date for end of day Thursday?

**Speaker 3** [13:52]
Today is Wednesday, June 4, 2025. So end of day Thursday is June 5, 2025.

**Shrestha Basu Mallick** [13:59]
Can you add that to the list?

**Speaker 3** [14:02]
Yes, I've added the exact date to the work list.

**Shrestha Basu Mallick** [14:06]
Allright.

**Speaker 3** [14:06]
It now says "Finish writing podcast script by end of day Thursday, June 5, 2025."

**Shrestha Basu Mallick** [14:11]
Perfect.

**Speaker 3** [14:11]
Is there anything else you'd like to add to the list?

**Shrestha Basu Mallick** [14:14]
Yes, actually.

**Speaker 3** [14:14]
Or anything else?

**Shrestha Basu Mallick** [14:15]
Yes, I would. Can you combine my grocery list, my reading list, and my work list and split them into two lists, personal tasks and work tasks?

**Speaker 3** [14:29]
I've combined your grocery list, reading list, and work list. They are now displayed on the screen as separate lists.

**Shrestha Basu Mallick** [14:37]
Why are garlic oil garlic and olive oil separate from pizza ingredients? Can you flesh out the grocery list, please?

**Speaker 3** [14:47]
You'reright. I've updated the grocery list to include garlic and olive oil as part of the pizza ingredients.

**Shrestha Basu Mallick** [14:53]
Allright, and can you consolidate the work list? I see two on the screen.

**Speaker 3** [14:59]
You'reright. I apologize for the duplication. I've consolidated the work list to a single entry. You should now see the updated list on the screen.

**Shrestha Basu Mallick** [15:08]
OK, and then finally, can you do the following? Can you assign grocery list to Kwin, reading list to Shrestha, with the exception of "Segmentation Fault," which Kwin will read, and work list to Shrestha?

**Speaker 3** [15:25]
I understand you want to assign the grocery list to Kwin, the reading list to Shrestha, except for "Segmentation Fault," which Kwin will read, and the work list to Shrestha as well.

**Shrestha Basu Mallick** [15:35]
Can you show that on the screen?

**Speaker 3** [15:39]
I've updated the lists to show the assignments he specified. You should now see the updated lists on the screen.

**Shrestha Basu Mallick** [15:45]
And you've spelled my name correctly, which was the harder name, but you spelled Kwin incorrectly. But OK, we're out of time here. OK, so if we want the final thing here is we would like to have a nice background for this app.

So can you.

**Speaker 3** [16:04]
I apologize for spelling your name.

**Shrestha Basu Mallick** [16:07]
OK, could you help us generate actually, I'll be more assertive. Please generate an app that has the words "Hello World" jumping up and down on the screen with the colors of Google. And then we also want two neon green ASCII cats which are animated.

**Speaker 3** [16:32]
I'm now generating an app with "Hello World" jumping up and down in Google colors and two animated neon green ASCII cats. It may take a moment for the app to appear on your screen.

**Kwindla Hultman Kramer** [16:52]
Detention.

**Shrestha Basu Mallick** [17:08]
So yeah, there you have the Google colors and the cats from Pipecat. So with that, I'll hand it over to Kwin.

### Demo Debrief

**Kwindla Hultman Kramer** [17:16]
So first, thank you or you should go back to your pride of place,right in the middle. Thank you for being such a good sport. Very messy code on my part, including things like basically no instructions to the LLM about how to display text on the screen and just telling it it has a function that can display text on the screen, and it sort of guesses and learns in context, as you can tell from Shrestha, about when it should clear the screen, because there's an optional clear argument to the add text to the screen function.

And it's super impressive, but also super jagged frontier about whether it kind of can intuit what you want to do in those contexts. So thank you for, like, doing this, because this is what I do all the time with this code, trying to figure out, like, what these models can do and what kind of code you have to write and what you don't have to scaffold for them to do well.

**Shrestha Basu Mallick** [18:03]
Yeah, and it's been, you know, playing with this. Every turn is different. And it's interesting to see the things that it struggles with, like your name. Even if I spell out the exact letters, it somehow really wants to spell Kwin the way it spells.

I think it also I mean, turn detection, as we saw, there's a lot of work that can be done, of course, there. And I'm trying to remember. There's and there's, of course, a lot of variation in sometimes here.

Like, there are times when it gets the grocery list perfect and, you know, combines the list perfectly, and sometimes it's a bit in the middle, like here.

**Kwindla Hultman Kramer** [18:39]
And the way this code works is it just for a given, like, session, it loads lots and lots and lots of previous conversational sessions in user assistant, user assistant sort of messages. Does sometimes, depending on the version of the code I've, like, got running, it summarizes a little bit.

Sometimes it doesn't. So we really are leaning on the intelligence of the LLM to do all of the sort of contextual understanding about what we mean by a list, what we mean by the context in which we are talking about that list.

It is super amazing that it works at all, basically, in my mind. And it's all voice-driven, and it's all multimodal from the ground up. We have a whole nother video we can show, but I definitely think we're out of time.

So we will.

**Guest** [19:19]
You have a final talk, and everyone seems to be excited, so.

### Grandmother's Knot

**Kwindla Hultman Kramer** [19:22]
Can't you one minute of.

**Shrestha Basu Mallick** [19:24]
Maybe we should talk about our grandmothers?

**Kwindla Hultman Kramer** [19:25]
Oh, yes. I totally forgot that part. Sorry. Let's skip past the demo where it gets the grocery list perfect.

**Shrestha Basu Mallick** [19:31]
I think maybe this crowd would like to see that demo.

**Kwindla Hultman Kramer** [19:34]
That was great. So this has been fun for me to work on, because, like, it's so relevant to my everyday life. But in Shrestha, and we're talking about it, and I think there's actually something else that really kind of hooked me that she said.

**Shrestha Basu Mallick** [19:46]
Yeah, so, you know, my grandmother was Indian, of course, and she used to wear this cloth garment called a sari. And her way of reminding herself when she had to do things was tying knots on the sari. Of course, and then I was chatting with Kwin, and what was incredible is apparently his grandmother in North Carolina, so very different from Calcutta in India, used to tie strings around her fingers.

And firstly, you know, this is kind of incredible. You know, no matter how many continents separate us, like, smart people come up with the same generally intelligent patterns. But it's also incredible how technology allows humans to evolve. Now, the one problem with either the knots or the strings is you knew you had to remember something, but you didn't know what it was.

So you still relied on your memory. And, you know, ultimately, that's why I do the work I do at Google, because I want to build the technologies that enable, you know, an infinite world of creative possibilities tomorrow or even today across continents.

And I just want to say that we believe that voice is the most natural of interfaces, and there will come a world where most of the interaction with language models will happen via voice. And the Gemini models are trained to be multimodal from the ground up.

So of course, they ingest text, voice, but also images and video. So if you have any questions about Gemini, please reach out to me on X, on LinkedIn, email, wherever. Happy to work with builders like yourself.

**Kwindla Hultman Kramer** [21:28]
Yeah, thanks for coming to the talk, and we would love to see what you build with these models and APIs.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
