AIAI EngineerMay 9, 2026· 19:27

Voice AI: when is the "Her" moment? — Neil Zeghidour, CEO, Gradium AI

Neil Zeghidour, CEO of Gradium AI, argues that voice AI remains far from the 'Her' ideal because cascaded systems (speech-to-text, LLM, text-to-speech) suffer from high latency—tool calls alone add 500ms to 4 seconds—while human response time is ~200ms. Speech-to-speech models reduce latency but are half-duplex, meaning they cannot handle overlapping speech or backchanneling, unlike Moshi, Gradium's full-duplex model. However, Moshi lacked intelligence, tool calls, and paralinguistic understanding—the ability to infer tone, hesitation, or discomfort from voice, which is stripped away in text. Cost is another barrier: TTS bills burn through fundraising before user bases grow. Gradium's solution is Phonon, an on-device TTS model running on smartphone CPUs, offering privacy and eliminating API fees. The path forward requires combining full-duplex natural conversation with the reliability and smarts of cascaded systems.

Transcript

Intro0:00

Neil Zeghidour0:15

Hi everyone. Uh, thanks for being here today. We're really happy to talk with you. It's theright time for this talk because we had a lot of great presentations around voice this morning, and I wanted to take a bit, you know, time to reflect where we are in the—modestly, where we areright now in the voice AI domain and what is left to do as main challenges.

Quick intro about Gradium: our mission is to unlock the unrealized potential of voice AI. So basically, we train voice models, speech-to-text, text-to-speech, speech-to-speech, whether it's transformation, translation, speech-to-speech dialogue, and so on and so forth. We want to be a main model provider for voice for everyone, building voice agents and voice solutions.

So we are not working on orchestration, we are not working on specific verticals, we just make building blocks for people who want to build voice AI.

Voice Cloning1:15

Neil Zeghidour1:15

I mean, I can let it be said by a famous podcaster that I cloned this morning:

there should be audio that is output.

Speaker 21:32

Man, have you seen what Gradium is doing with voice cloning? It's kind of crazy. Like, seriously.

Neil Zeghidour1:37

Can we have—

Speaker 21:38

You have 10 seconds of your voice. That's it.

Neil Zeghidour1:40

Can we have much more volume, please?

Speaker 21:45

Man, have you seen what Gradium is doing with voice cloning? It's kind of crazy. Like, seriously. You record, like, 10 seconds of your voice, that's it, and the system analyzes the tone, the pitch, the accent, all those little quirks that make your voice yours.

Then boom, you type text and it talks back.

Neil Zeghidour2:02

Okay, I guess you recognize maybe Joe Rogan. I hope you did. So basically, this is a spin-off from a nonprofit lab we created 2 years ago with funding from philanthropists, including Eric Schmidt, Rudolf Seidel, and Xavier Niel. The main idea was to create a lab that does open research, and we focused mostly on speech.

So we developed Moshi, which was the first speech-to-speech model for conversation, speech-to-speech translation, pocket TTS, most recently a CPU model. And we decided to also create this for-profit structure to make products that can be used in production beyond open source.

So the goal of the talk is basically based on the "Her" movie. So this has been the most overused, most annoying analogy, I think, in the field. At the same time, it's extremely relevant because it's 13 years old, and if we look at one of the introduction scenes—so that's when the main character meets his AI voice, Samantha, for the first time—it still sounds like it was, you know, anticipating, I think, what interaction could look like.

Her Moment2:39

Speaker 33:11

Oh, what—what do I call you? Do you have a name?

Speaker 43:14

Um, yes. Samantha.

Speaker 33:18

Wait, where'd you get that name from?

Speaker 43:20

I gave it to myself, actually.

Speaker 33:23

How come?

Speaker 43:24

Because I like the sound of it. Samantha.

Neil Zeghidour3:29

And, you know, then came the, like, the trend of the "Her" moment. And we got so many "Her" moments on Twitter. And in real life, again, I don't want to be mean to anyone, so I will also make fun of myself.

It's more like, you know, trying to be pragmatic about what was the promise and where we areright now. So this is a very recent demo from, for me, what is the best voice AI company in the world, which is 11 Labs.

Speaker 54:00

Hello, you're speaking to the government's AI helper. How can I help you today?

Speaker 64:05

I would like to start a new business.

Speaker 54:09

To start a new business, you'll need to choose a business structure and name, complete the registration, and upload required documents.

Neil Zeghidour4:17

And this one is a demo that I did a few months ago with a Richie Minnie, so our presenter this morning. You know, but it was very risky, I recognize. I was very scared when I did it. But basically, putting our streaming voice models into a Richie Minnie.

Speaker 24:33

Action. Just say the word.

Neil Zeghidour4:35

Okay, maybe you can do something a bit fun. I want to get, you know, to improve my health overall, and so I'm looking for a bro to go to the gym. Can you take the personality and the voice of a gym addict?

Speaker 74:50

Hey, I'm Logan, your gym bro. Let's crush those gains together. You ready to lift, sweat, and feel awesome? I've got your back, no excuses, just results.

Neil Zeghidour4:58

So, okay, this—it's fine. It sounds more natural than it used to. But in both cases, you know, I mean, we're still not there,right? The latency is still quite high. The ability to handle simultaneous speaking between the user and the system is not there.

Intelligence starts to become much better, and I think that's also why we see all this traction around voice agents, because there are agents, actually, to whom we can give a voice and they can be useful. And also, this is mostly just glorified a text model with a voice around it, and so, you know, anything that is not in the text will not be able to be leveraged.

And maybe I could ask one last time to increase a bit the volume. I think that could be even better. And so what does it take to get there? So basically, we had a very nice presentation this morning by Samuel about cascaded systems, so I will go quickly.

Speech-to-text, LLM text-to-speech, that's the classic cascade. In our case, we do streaming speech-to-text, streaming text-to-speech, with voice cloning, semantic VAD, the classic stuff. So latency, you know, getting to a fast conversation. So we have a very fast TTS, so that's the latency for our TTS compared to a few other models.

Cascaded Limits5:48

Neil Zeghidour6:11

But, you know, that means that just the TTS is still more than 200 milliseconds, while in a human conversation you need the entire stack of understanding, producing an answer, and pronouncing it to be around 200 milliseconds. So that will none of this will allow to have a conversation that sounds human.

And this is just latency for text conversation. There is no tool call, no actual task that is performed. Now, that's another scene from the movie.

Speaker 46:41

Okay, let's start with your emails. You have several thousand emails regarding LA Weekly, but it looks like you haven't worked there in many years.

Speaker 36:48

Oh, yeah.

Neil Zeghidour6:49

Okay, so here it just went instantly into all the emails and gathered information and so on and so forth. Obviously, we are not there at all. Today, if you have a voice agent that is supposed to use a tool, you're going to wait.

Tool Call Latency7:03

Neil Zeghidour7:03

The tool call, you know, we are fighting for latency of the TTS, trying to grab 10 milliseconds, 20 milliseconds, and then you have a tool call or open router that is going to have a latency between 500 milliseconds and 4 seconds.

So in a way, sometimes you feel like we're fighting for something that is not the most relevant anymore. I think now the main bottleneck is becoming the tool call, which is very unpredictable, and we should have models that are resilient to very complex tool calls.

So what we are, you know, one solution to do that is to have fillers. So basically, you know, your LLM, it splits into two things. So it sends a tool call, and while it waits for getting the result back, it can keep the conversation going in a natural way, and then it retrieves the result and it tries to insert it back naturally in the conversation.

So just because I always do live demos, I will do a very short improvised live demo with something that was vibe-coded very quickly, so it maybe go very wrong. But the idea here is to have, like, just a vibe-coded travel agent, and that can, when I ask it to retrieve places, it's going to try to find a nice thing to say about the location I'm going to.

Speaker 28:11

Hello, this is Colin from Wanderlust Travel. I see you're looking to book a trip for two people from April 10 to the 13th, 2026. Where would you like to travel to for your getaway?

Neil Zeghidour8:24

I want to go to Tokyo.

Speaker 28:26

Tokyo is such an incredible choice. It's a fascinating mix ofultra-modern skyscrapers and beautiful, peaceful shrines. I found some fantastic options for you. You could stay at the Fairmont for a serene.

Neil Zeghidour8:41

So this needs polishing. The main idea, you know, is that while it was trying to gather things and doesn't really know how long it will take to get the data, it tries to say a nice thing about the place you're going to.

So that's one way to get latency that is going to be more controlled and more reliable, even despite the complexity of the thing. But then we know that latency of cascaded system is inherently high,right? And so then what I hear a lot from customers, from investors, and so on, is what about speech-to-speech?

Speech-to-Speech8:51

Neil Zeghidour9:13

So I think a big confusion is what is speech-to-speech. So speech-to-speech is the idea that now, instead of having the three blocks, you only have one that does everything together,right? So instead of having audio that goes in speech-to-text and then LLM text-to-speech, you have a model that takes speech as input and is going to output speech.

And that reduces latency a lot, but that's still not a human conversation. In particular, because every single speech-to-speech model except Moshi is half-duplex. What that means is even the best speech-to-speech model, we could argue, I don't know, maybe that's the advanced voice model of OpenAI, or Sesame, I'm a huge fan of the Sesame

voice model, is still half-duplex. What it means is that the model is either listening or it's speaking, and it cannot handle the ambiguity of a human conversation where you can have overlap between people speaking on one another, a lot of things happening, you cough, you know, you just do hm-hm, which is called backchanneling, and then it breaks completely.

Full-Duplex10:09

Neil Zeghidour10:09

Full duplex, that's a human conversation. It varies between cultures and languages, but for example, in Japanese, it's a sign of politeness and that you are actively listening to do a lot of backchanneling. So you say hm-hm, ah-ah-ah constantly when the other person is speaking, and you get up to 20% of the time that is overlapped between the people,right?

So that's what makes a conversation human. And so I shot this video in my hotel room this morning just to, you know, that's a speech-to-speech model, but that's not a full duplex model, and you see how it can become annoying.

Speaker 810:40

Hey, how's it going?

Speaker 410:42

Hey, I'm doing really well. Thanks for asking. How about you? How's your day going so far?

Speaker 810:47

I'm great. You know, I'm preparing a talk about voice AI and how far we are from, you know, from the "Her" movie. And, yeah, I'd just like to brainstorm a bit about it with you.

Speaker 411:01

Oh, that's a great topic. Yeah, I'd love to help you brainstorm. Are you thinking.

Speaker 811:04

Yep.

Speaker 411:06

Exactly. Yeah, I was thinking.

Speaker 811:07

No, no, I didn't mean to interrupt. You know, I was just saying, yeah, like that. You can just keep going. You know, don't mind me. That's something I typically do.

Speaker 411:15

Ah, no worries at all. Yeah, I was just going to say we could break it down into a few aspects.

Speaker 811:19

Sure.

Speaker 411:19

Like. Yeah, exactly.

Speaker 811:21

No, no, please stop interrupting. You know, it's called backchanneling. Humans do it all the time. It shows that you're just following the conversation, that you don't, you know, like, interrupt you in your flow. Just keep going.

Speaker 411:31

Ah, got it. Thanks for letting me know.

Speaker 811:33

No problem.

Speaker 411:34

Yeah.

Speaker 811:35

Ah, come on.

Neil Zeghidour11:36

Okay, so we are basically, I was a bit mean,right? That's, you know, that's my point. It's not an actual conversation,right? And it can become very annoying. In particular, that's also why, you know, most of the voice AI demos, they are shot in an empty, like, quiet room next to the phone and so on.

A lot of things can break. So what we did instead, in our case, was. Sorry, where am I in my presentation? I'mright here with Moshi, was the first full duplex system. So here you'll see my co-founder, Alex, talking to it.

It's almost two years old now. I think it's still kind of age well because what you'll see is, you know, they are going to talk on one another constantly, and it's just fine.

Speaker 912:17

So the planet is Sirius 22. Can you plot a trajectory course to it, please?

Speaker 1012:22

Yes, sir.

Speaker 912:23

Okay. How long is it going to take us to get there?

Speaker 1012:25

I've mapped it out. It's approximately five months to get there.

Speaker 912:28

Okay, that's not too bad. Do you think we have all we need on board the ship to start the mission?

Speaker 1012:33

Yes, sir. We have everything we need.

Speaker 912:35

Okay.

Neil Zeghidour12:36

So, you know, even when the model has guessed what you're going to say, it starts answering before you're done. At the same time, it can talk over it, and it's not ignoring what you're saying. It will, you know, consider it afterwards and so on.

So you have what is the most robust, to this day, conversational experience, robust to noise, to a lot of people speaking, and so on and so forth. But now, if we compare it

Paralinguistics12:59

Neil Zeghidour12:59

to the "Her" movie.

Speaker 313:00

So do you know what I'm thinkingright now?

Speaker 413:02

Well, I take it from your tone that you're challenging me, maybe because you're curious how I work. Do you want to know how I work?

Speaker 313:09

Yeah, actually.

Neil Zeghidour13:10

So maybe this was not very clear, but this snippet here, it's the AI understanding that the character is a bit uncomfortable. So that's paralinguistic understanding. You know, it's understanding all the cues that come from the way people speak.

Technically, that is in Moshi, that is in any speech-to-speech model, because this information is not lost. However, if you don't exploit this information

to make your model say relevant things, it's never going to exploit it. If you train it on the audio version of an instructor dataset and it's just factual question answering, why would it even try to capture this information?

So basically, Moshi, I think we saw, it's still the only full duplex model. Recently, NVIDIA published the personal plex model based on it. What was great is still the flow of it is just, honestly, impossible to match, I guess.

It's conversational and very robust. At the same time, the model was very stupid. So basically, it was, you know, just useless. You could talk for a few minutes, and then it was a bit pointless, you know? Because it was not an agent.

It had no tool call, no ability to do anything. It's impossible to use in production something that has no observability. You don't know if the, it's very hard to detect if someone said something that should be not accepted, and so on and so forth.

And there was no real paralinguistic understanding. So

the main takeaway for me is

we know that this nature of interactive models that are going to be full duplex, that are going to be really the, that will be the way to get an interaction that is as natural as you would have with a human.

But as long as we don't, we are not able to give to this kind of very natural-sounding models the same level of reliability, intelligence, and personalization as cascaded systems, I don't see a path towards

them, you know, replacing cascaded systems. So I used to be really at war against cascaded systems. I think they are so practical and so convenient that the main challenge, honestly, I think this, we solved that with Moshi. Anybody who implements it and trains it on better data will have something that sounds just indistinguishable from humans.

But all of this is going, all the challenge is going to happen here. A last point is the scalability. So now, let's say you have the best speech-to-speech model, okay? It solves all the stuff that I've talked about.

Scaling Walls15:34

Neil Zeghidour15:34

If we take, again, the analogy from the movie, you will talk to it maybe several hours a day, or it will always be on because, you know, it's on your computer when you work and you asynchronously ask things to it, and so on and so forth.

I'm not going to mention the cost of the API of our competitors, but voice is very expensive. Everybody in this room probably knows it. The voice model of most hyperscalers is run at a loss. It's a gigantic multimodal model, and they lose money every time you use it.

But it's, you know, it's kind of a marketing thing, and so on. It's fine. But now, if we want to make it an actual profitable product that people are using at a massive scale, just not going to work.

And in particular, anyone who tried developing a consumer app with voice realized that LLM now is almost nothing in terms of cost. All the bill is TTS. Speech-to-text is very cheap as well. Diarization is affordable. TTS is really what is going to consume most of the, and I saw people burning their fundraising in TTS bills, and they don't even get the opportunity to get their user base to grow.

So another aspect is privacy. The more you're going to open to your AI, the more, you know, you'll want it to be more controlled and private and feel more comfortable that things are not shared publicly. In particular, we see now people with MITOS being afraid that any single database in the world is going to be hacked in a few months.

And so, you know, you will be more comfortable if all your private data is local. And so to solve that, our first step is Gradium Phonon. It's on-device TTS. So on-device means a lot of things. For some people, on-device means it runs on a gamer GPU.

On-Device TTS16:57

Neil Zeghidour17:12

For us, on-device means it runs on a smartphone CPU. And so it's a very small model, less than 100 million parameters. And for its size, it works quite well. It's better than all the existing on-device models. Kokoro is a good one, but it doesn't have voice cloning.

And I'm out of time, so I will just play a short demo if I can. Yes.

Speaker 217:38

Geez Morty, stop looking for a signal. Gradium Phonon runs locally on the CPU, which means high-fidelity neural speechright here without those intergalactic cloud government hacks.

Speaker 1117:47

It is about simplicity with no servers and no waiting, just a smooth, quiet performance that stays on the device with local processing and total privacy. What else?

Speaker 217:58

Local processing, that sounds like the machine is making its own donut. Wait, if the CPU is doing all the work, can it sprinkle.

Neil Zeghidour18:05

Yeah, so this runs on a smartphone CPU, which means that you can use that to power, you know, any kind of voice application without paying a single cent of API fee. So we opened a private beta for this model.

Conclusion18:18

Neil Zeghidour18:18

The goal for us is to allow people to create consumer apps with voice, basically, and that they are able to scale the usage without having to lose money on the API. So the conclusion, the path forward is, for us, I have a strict opposition to some of our competitors who say that voice is a commodity now.

I think it's completely false. Voice is very challenging. The last nine is going to be the most difficult to solve. And for us, it's really about science and engineering, and it will get "Her" to us to the "Her" movie.

Sorry. So you can use us on gradium.ai, and if you want to join us to bring "Her" to life, yes, now I'm using this analogy. No. Just if you want to work on very exciting voice models, you can apply and join us.

Thanks a lot for your attention.