Intro0:00
Hi everyone. Uh, thanks for being here today. We're really happy to talk with you. It's theright time for this talk because we had a lot of great presentations around voice this morning, and I wanted to take a bit, you know, time to reflect where we are in the—modestly, where we areright now in the voice AI domain and what is left to do as main challenges.
Quick intro about Gradium: our mission is to unlock the unrealized potential of voice AI. So basically, we train voice models, speech-to-text, text-to-speech, speech-to-speech, whether it's transformation, translation, speech-to-speech dialogue, and so on and so forth. We want to be a main model provider for voice for everyone, building voice agents and voice solutions.
So we are not working on orchestration, we are not working on specific verticals, we just make building blocks for people who want to build voice AI.
Voice Cloning1:15
I mean, I can let it be said by a famous podcaster that I cloned this morning:
there should be audio that is output.
Man, have you seen what Gradium is doing with voice cloning? It's kind of crazy. Like, seriously.
Can we have—
You have 10 seconds of your voice. That's it.
Can we have much more volume, please?
Man, have you seen what Gradium is doing with voice cloning? It's kind of crazy. Like, seriously. You record, like, 10 seconds of your voice, that's it, and the system analyzes the tone, the pitch, the accent, all those little quirks that make your voice yours.
Then boom, you type text and it talks back.
Okay, I guess you recognize maybe Joe Rogan. I hope you did. So basically, this is a spin-off from a nonprofit lab we created 2 years ago with funding from philanthropists, including Eric Schmidt, Rudolf Seidel, and Xavier Niel. The main idea was to create a lab that does open research, and we focused mostly on speech.
So we developed Moshi, which was the first speech-to-speech model for conversation, speech-to-speech translation, pocket TTS, most recently a CPU model. And we decided to also create this for-profit structure to make products that can be used in production beyond open source.
So the goal of the talk is basically based on the "Her" movie. So this has been the most overused, most annoying analogy, I think, in the field. At the same time, it's extremely relevant because it's 13 years old, and if we look at one of the introduction scenes—so that's when the main character meets his AI voice, Samantha, for the first time—it still sounds like it was, you know, anticipating, I think, what interaction could look like.
Her Moment2:39
Oh, what—what do I call you? Do you have a name?
Um, yes. Samantha.
Wait, where'd you get that name from?
I gave it to myself, actually.
How come?
Because I like the sound of it. Samantha.
And, you know, then came the, like, the trend of the "Her" moment. And we got so many "Her" moments on Twitter. And in real life, again, I don't want to be mean to anyone, so I will also make fun of myself.
It's more like, you know, trying to be pragmatic about what was the promise and where we areright now. So this is a very recent demo from, for me, what is the best voice AI company in the world, which is 11 Labs.
Hello, you're speaking to the government's AI helper. How can I help you today?
I would like to start a new business.
To start a new business, you'll need to choose a business structure and name, complete the registration, and upload required documents.
And this one is a demo that I did a few months ago with a Richie Minnie, so our presenter this morning. You know, but it was very risky, I recognize. I was very scared when I did it. But basically, putting our streaming voice models into a Richie Minnie.
Action. Just say the word.
Okay, maybe you can do something a bit fun. I want to get, you know, to improve my health overall, and so I'm looking for a bro to go to the gym. Can you take the personality and the voice of a gym addict?
Hey, I'm Logan, your gym bro. Let's crush those gains together. You ready to lift, sweat, and feel awesome? I've got your back, no excuses, just results.
So, okay, this—it's fine. It sounds more natural than it used to. But in both cases, you know, I mean, we're still not there,right? The latency is still quite high. The ability to handle simultaneous speaking between the user and the system is not there.
Intelligence starts to become much better, and I think that's also why we see all this traction around voice agents, because there are agents, actually, to whom we can give a voice and they can be useful. And also, this is mostly just glorified a text model with a voice around it, and so, you know, anything that is not in the text will not be able to be leveraged.
And maybe I could ask one last time to increase a bit the volume. I think that could be even better. And so what does it take to get there? So basically, we had a very nice presentation this morning by Samuel about cascaded systems, so I will go quickly.
Speech-to-text, LLM text-to-speech, that's the classic cascade. In our case, we do streaming speech-to-text, streaming text-to-speech, with voice cloning, semantic VAD, the classic stuff. So latency, you know, getting to a fast conversation. So we have a very fast TTS, so that's the latency for our TTS compared to a few other models.
Cascaded Limits5:48
But, you know, that means that just the TTS is still more than 200 milliseconds, while in a human conversation you need the entire stack of understanding, producing an answer, and pronouncing it to be around 200 milliseconds. So that will none of this will allow to have a conversation that sounds human.
And this is just latency for text conversation. There is no tool call, no actual task that is performed. Now, that's another scene from the movie.
Okay, let's start with your emails. You have several thousand emails regarding LA Weekly, but it looks like you haven't worked there in many years.
Oh, yeah.
Okay, so here it just went instantly into all the emails and gathered information and so on and so forth. Obviously, we are not there at all. Today, if you have a voice agent that is supposed to use a tool, you're going to wait.
Tool Call Latency7:03
The tool call, you know, we are fighting for latency of the TTS, trying to grab 10 milliseconds, 20 milliseconds, and then you have a tool call or open router that is going to have a latency between 500 milliseconds and 4 seconds.
So in a way, sometimes you feel like we're fighting for something that is not the most relevant anymore. I think now the main bottleneck is becoming the tool call, which is very unpredictable, and we should have models that are resilient to very complex tool calls.
So what we are, you know, one solution to do that is to have fillers. So basically, you know, your LLM, it splits into two things. So it sends a tool call, and while it waits for getting the result back, it can keep the conversation going in a natural way, and then it retrieves the result and it tries to insert it back naturally in the conversation.
So just because I always do live demos, I will do a very short improvised live demo with something that was vibe-coded very quickly, so it maybe go very wrong. But the idea here is to have, like, just a vibe-coded travel agent, and that can, when I ask it to retrieve places, it's going to try to find a nice thing to say about the location I'm going to.
Hello, this is Colin from Wanderlust Travel. I see you're looking to book a trip for two people from April 10 to the 13th, 2026. Where would you like to travel to for your getaway?
I want to go to Tokyo.
Tokyo is such an incredible choice. It's a fascinating mix ofultra-modern skyscrapers and beautiful, peaceful shrines. I found some fantastic options for you. You could stay at the Fairmont for a serene.
So this needs polishing. The main idea, you know, is that while it was trying to gather things and doesn't really know how long it will take to get the data, it tries to say a nice thing about the place you're going to.
So that's one way to get latency that is going to be more controlled and more reliable, even despite the complexity of the thing. But then we know that latency of cascaded system is inherently high,right? And so then what I hear a lot from customers, from investors, and so on, is what about speech-to-speech?
Speech-to-Speech8:51
So I think a big confusion is what is speech-to-speech. So speech-to-speech is the idea that now, instead of having the three blocks, you only have one that does everything together,right? So instead of having audio that goes in speech-to-text and then LLM text-to-speech, you have a model that takes speech as input and is going to output speech.
And that reduces latency a lot, but that's still not a human conversation. In particular, because every single speech-to-speech model except Moshi is half-duplex. What that means is even the best speech-to-speech model, we could argue, I don't know, maybe that's the advanced voice model of OpenAI, or Sesame, I'm a huge fan of the Sesame
voice model, is still half-duplex. What it means is that the model is either listening or it's speaking, and it cannot handle the ambiguity of a human conversation where you can have overlap between people speaking on one another, a lot of things happening, you cough, you know, you just do hm-hm, which is called backchanneling, and then it breaks completely.
Full-Duplex10:09
Full duplex, that's a human conversation. It varies between cultures and languages, but for example, in Japanese, it's a sign of politeness and that you are actively listening to do a lot of backchanneling. So you say hm-hm, ah-ah-ah constantly when the other person is speaking, and you get up to 20% of the time that is overlapped between the people,right?
So that's what makes a conversation human. And so I shot this video in my hotel room this morning just to, you know, that's a speech-to-speech model, but that's not a full duplex model, and you see how it can become annoying.
Hey, how's it going?
Hey, I'm doing really well. Thanks for asking. How about you? How's your day going so far?
I'm great. You know, I'm preparing a talk about voice AI and how far we are from, you know, from the "Her" movie. And, yeah, I'd just like to brainstorm a bit about it with you.
Oh, that's a great topic. Yeah, I'd love to help you brainstorm. Are you thinking.
Yep.
Exactly. Yeah, I was thinking.
No, no, I didn't mean to interrupt. You know, I was just saying, yeah, like that. You can just keep going. You know, don't mind me. That's something I typically do.
Ah, no worries at all. Yeah, I was just going to say we could break it down into a few aspects.
Sure.
Like. Yeah, exactly.
No, no, please stop interrupting. You know, it's called backchanneling. Humans do it all the time. It shows that you're just following the conversation, that you don't, you know, like, interrupt you in your flow. Just keep going.
Ah, got it. Thanks for letting me know.
No problem.
Yeah.
Ah, come on.
Okay, so we are basically, I was a bit mean,right? That's, you know, that's my point. It's not an actual conversation,right? And it can become very annoying. In particular, that's also why, you know, most of the voice AI demos, they are shot in an empty, like, quiet room next to the phone and so on.
A lot of things can break. So what we did instead, in our case, was. Sorry, where am I in my presentation? I'mright here with Moshi, was the first full duplex system. So here you'll see my co-founder, Alex, talking to it.
It's almost two years old now. I think it's still kind of age well because what you'll see is, you know, they are going to talk on one another constantly, and it's just fine.
So the planet is Sirius 22. Can you plot a trajectory course to it, please?
Yes, sir.
Okay. How long is it going to take us to get there?
I've mapped it out. It's approximately five months to get there.
Okay, that's not too bad. Do you think we have all we need on board the ship to start the mission?
Yes, sir. We have everything we need.
Okay.
So, you know, even when the model has guessed what you're going to say, it starts answering before you're done. At the same time, it can talk over it, and it's not ignoring what you're saying. It will, you know, consider it afterwards and so on.
So you have what is the most robust, to this day, conversational experience, robust to noise, to a lot of people speaking, and so on and so forth. But now, if we compare it
Paralinguistics12:59
to the "Her" movie.
So do you know what I'm thinkingright now?
Well, I take it from your tone that you're challenging me, maybe because you're curious how I work. Do you want to know how I work?
Yeah, actually.
So maybe this was not very clear, but this snippet here, it's the AI understanding that the character is a bit uncomfortable. So that's paralinguistic understanding. You know, it's understanding all the cues that come from the way people speak.
Technically, that is in Moshi, that is in any speech-to-speech model, because this information is not lost. However, if you don't exploit this information
to make your model say relevant things, it's never going to exploit it. If you train it on the audio version of an instructor dataset and it's just factual question answering, why would it even try to capture this information?
So basically, Moshi, I think we saw, it's still the only full duplex model. Recently, NVIDIA published the personal plex model based on it. What was great is still the flow of it is just, honestly, impossible to match, I guess.
It's conversational and very robust. At the same time, the model was very stupid. So basically, it was, you know, just useless. You could talk for a few minutes, and then it was a bit pointless, you know? Because it was not an agent.
It had no tool call, no ability to do anything. It's impossible to use in production something that has no observability. You don't know if the, it's very hard to detect if someone said something that should be not accepted, and so on and so forth.
And there was no real paralinguistic understanding. So
the main takeaway for me is
we know that this nature of interactive models that are going to be full duplex, that are going to be really the, that will be the way to get an interaction that is as natural as you would have with a human.
But as long as we don't, we are not able to give to this kind of very natural-sounding models the same level of reliability, intelligence, and personalization as cascaded systems, I don't see a path towards
them, you know, replacing cascaded systems. So I used to be really at war against cascaded systems. I think they are so practical and so convenient that the main challenge, honestly, I think this, we solved that with Moshi. Anybody who implements it and trains it on better data will have something that sounds just indistinguishable from humans.
But all of this is going, all the challenge is going to happen here. A last point is the scalability. So now, let's say you have the best speech-to-speech model, okay? It solves all the stuff that I've talked about.
Scaling Walls15:34
If we take, again, the analogy from the movie, you will talk to it maybe several hours a day, or it will always be on because, you know, it's on your computer when you work and you asynchronously ask things to it, and so on and so forth.
I'm not going to mention the cost of the API of our competitors, but voice is very expensive. Everybody in this room probably knows it. The voice model of most hyperscalers is run at a loss. It's a gigantic multimodal model, and they lose money every time you use it.
But it's, you know, it's kind of a marketing thing, and so on. It's fine. But now, if we want to make it an actual profitable product that people are using at a massive scale, just not going to work.
And in particular, anyone who tried developing a consumer app with voice realized that LLM now is almost nothing in terms of cost. All the bill is TTS. Speech-to-text is very cheap as well. Diarization is affordable. TTS is really what is going to consume most of the, and I saw people burning their fundraising in TTS bills, and they don't even get the opportunity to get their user base to grow.
So another aspect is privacy. The more you're going to open to your AI, the more, you know, you'll want it to be more controlled and private and feel more comfortable that things are not shared publicly. In particular, we see now people with MITOS being afraid that any single database in the world is going to be hacked in a few months.
And so, you know, you will be more comfortable if all your private data is local. And so to solve that, our first step is Gradium Phonon. It's on-device TTS. So on-device means a lot of things. For some people, on-device means it runs on a gamer GPU.
On-Device TTS16:57
For us, on-device means it runs on a smartphone CPU. And so it's a very small model, less than 100 million parameters. And for its size, it works quite well. It's better than all the existing on-device models. Kokoro is a good one, but it doesn't have voice cloning.
And I'm out of time, so I will just play a short demo if I can. Yes.
Geez Morty, stop looking for a signal. Gradium Phonon runs locally on the CPU, which means high-fidelity neural speechright here without those intergalactic cloud government hacks.
It is about simplicity with no servers and no waiting, just a smooth, quiet performance that stays on the device with local processing and total privacy. What else?
Local processing, that sounds like the machine is making its own donut. Wait, if the CPU is doing all the work, can it sprinkle.
Yeah, so this runs on a smartphone CPU, which means that you can use that to power, you know, any kind of voice application without paying a single cent of API fee. So we opened a private beta for this model.
Conclusion18:18
The goal for us is to allow people to create consumer apps with voice, basically, and that they are able to scale the usage without having to lose money on the API. So the conclusion, the path forward is, for us, I have a strict opposition to some of our competitors who say that voice is a commodity now.
I think it's completely false. Voice is very challenging. The last nine is going to be the most difficult to solve. And for us, it's really about science and engineering, and it will get "Her" to us to the "Her" movie.
Sorry. So you can use us on gradium.ai, and if you want to join us to bring "Her" to life, yes, now I'm using this analogy. No. Just if you want to work on very exciting voice models, you can apply and join us.
Thanks a lot for your attention.





