Intro0:00
Alright. Hello everyone, I hope you have enough space.
Um, no, thanks so much for joining. We're building multilingual conversational AI agents. I know it's a bit of a mouthful, but yeah, hopefully we'll get something going.
Yeah, so different than on the poster, we're not from EvelynLabs, we're from ElevenLabs. Have— can I get a quick show of hands? Have you heard of ElevenLabs? Okay, everyone has heard of it, that's great. So we can— we can do some nice things today, maybe look at some new stuff, maybe you haven't played with yet.
So this is me. I'm Thor, here on theright side of the screen. I work on developer experience at ElevenLabs, so, you know, a lot of it is kind of the conversational AI, you know, agents platform. And I also have my colleague Paul here with me.
So if you have kind of any questions, you know, throughout the workshop, later on, we'll be floating around and, you know, happy to kind of, you know, answer your questions. So feel free to just put your hand up and we'll be floating around later.
Yeah, Paul works with me on developer experience, so if you have any feedback as well, you know, on documentation, examples, developer experience, or generally just the product, do let us know. Yeah, also this, you can scan the QR code to get the slides.
So there's a couple of resources that are linked in the slides, as well as there's a form where you can fill in your email address if you want to get some credits. So we can give you, you know, a couple of credits for the next three months to sort of, you know, play around with ElevenLabs and kind of try this out.
So yeah, feel free to just scan this and then, you know, kind of save it for later on. There's the resources linked in there as well. Cool. Yeah, also if you're building with ElevenLabs, I recommend you follow our ElevenLabs devs Twitter account.
This is specifically for, you know, updates in terms of API versions, client libraries, so anything, you know, if you're a developer building with ElevenLabs, that's a good place to follow and kind of, you know, be in the loop with what is happening.
And then, yeah, I mentioned, so if you scanned the QR code earlier, there's also the link so you can tap on the QR code as well to open the form. And if you just fill in your email address, we will send you after this workshop.
So in the workshop, you can get started with kind of the free account. That should be— should be plenty. But then, you know, we'll give you a coupon code or we'll send it to you via email for kind of the next three months to try it out.
Cool. And as I mentioned, yeah, the resources in the slides, so we'll, you know, kind of tap into these later on and then can kind of, you know, get started building. We can see as well kind of what folks are building.
You know, since we're looking at kind of multilingual conversational AI agents, do you just want to shout out kind of the languages that, you know, you're specifically looking to unlock with conversational AI? Anyone, you know, anything else than English?
Capabilities3:36
Any other languages?
Portuguese.
Portuguese, very nice. You're looking for like a Brazilian Portuguese accent or? Cool, yeah, we can look at that. So we got Portuguese. Any other languages?
Spanish.
Spanish, yes. We got some Spanish in there as well, that's good. So Portuguese, Spanish.
BP products,right, Hungary. Hungarian.
Hungarian, very good. I'm actually, do we have Hungarianright now? We don't yet, but I think hopefully soon. So we're working on the version three of the multilingual models, and I do think we'll need to double-check. Not that I promise something wrong, but we might be able, maybe not today, but maybe in a couple of weeks we can give you Hungarian, that's great.
Okay. Any other? Mandarin?
Hindi.
Hindi, yeah. Huge population that speaks Hindi. Then again, India has 50 plus languages, I believe. So we're working on some adding some additional, currently we have Hindi and Tamil, so we're working on some additional languages there. Okay, so we have a good mix that we can play around with.
Cool, yeah. If there's any other languages later, we can, you know, kind of explore those as well. Maybe one language that probably doesn't get spoken about often enough is barking, actually. If you want to build applications for the 900 million dogs that are out there, I actually looked this up, you can use our model.
So this is our most recent launch, two months ago now, I believe, the newest model we launched. So maybe we can have a little listen. The Dachshund is my favorite, actually.
The Golden Retriever, a little bit more.
And Joanne.
Yeah, that one's a bit feisty. But yeah, if you're building application for dogs, that might be a great use case for you. Did anyone see this launch like two months ago? No? Yeah, you saw it? Yeah, the unfortunate thing is we launched it first of April, actually, and so everyone thought it was an April Fool's.
So the timing was a bit unfortunate on that one. But as you can see, you know, text to bark, it's very real. No, in fact, it was an April Fool's. So be careful when you play this to your dog, they might get offended.
Because the context, we can't guarantee that it translates. But the actual sounds that you're hearing are generated by our sound effects model. So we do have a model. If you go to the app, there's this sound effects model here, and you can actually, this sound effect was also created.
But, so truck reversing, maybe that's a good one. You can click generate. And so basically what we do is we generate kind of four different samples for you that you can use. So, you know, not so interesting for your conversations per se, but like, you know, if you're creating video games, for example, I don't know if anyone is doing that.
We might be having some internet. I hope we have internet, no? Okay, the internet is looking okay. So, yeah, I'm not sure what's happening here. But so, for example, this was like a drum cowbell I generated recently. Actually, we do have a, if you Google ElevenLabs sound board, we built this recently, which was, which is pretty cool.
So you can kind of loop, it's basically like a drum machine, and you can get kind of the
sound effects. The drums are pretty nice as well. And, you know, they are mapped to like the keys on the keyboard. So you can, you can play like,
so that's pretty cool. And then you can add kind of new sound effects and stuff. You know, just to give you an idea of some of the things that we do. Now, obviously here, you know, we're talking about conversational AI agents.
And so specifically, if we kind of look at the different components that are involved in building conversational AI agents, we have the user who is speaking, you know, some language. And so we need to transcribe that speech into text.
Pipeline8:33
We then feed that into a large language model, which is kind of acting as the brain of, you know, our agent. So in our case, we currently don't build any intelligence models. So we partner with kind of the existing large language models providers.
So like your GPT-4.0, your Google Gemini, what have you. And then this large language model, which will act as kind of the brain of your agent, will generate a text output. And then we basically stream that text output back into speech.
So that is roughly kind of the pipeline that we have for building conversational AI agents. Now, with that, there is a bunch of system tools that are built in. So, for example, we have this language detection system tool, which facilitates the language switching that we're going to look at in a bit.
We also have, you know, function calling, tool calling. So, you know, if you need to give access to kind of specific functionality to your agent, you can do that here as well. And then, you know, there's different approaches.
So like if you've seen OpenAI Real-Time, for example, so this doesn't actually go through text,right? So it goes sound token to sound token, which has some benefits. But what we've seen, you know, for deploying conversational AI agents at scale and really understanding, you know, what is kind of happening.
So if you're going sound token to sound token, you're kind of flying blind a little bit. So, you know, you're trusting the model that sort of it actually replies intelligently. Whereas if you're going through text, you can have, you know, kind of better monitoring and sort of understanding what's going on in your conversation.
So while we're also exploring, you know, kind of sound to sound for kind of the conversational AI agents, for now what we found works best is sort of this pipeline that we've built. And we deploy kind of all these models very close to each other to kind of bring down the latency as much as possible.
Cool. And so now we can look at kind of the individual sort of components within that. So, for example, for speech to text, so this is actually the most recent model we actually launched. You know, that wasn't an April Fool's.
And so this is our speech to text model. So our automatic speech recognition model, so ASR. And this works, you know, kind of benchmark leading across 99 different languages at the moment. So what this does, you know, you can see sort of the functionality that's built in there.
ASR Demo11:08
There's, you know, word-level timestamps, there's speaker diarization, there's audio event tagging. So if you want in your transcript, you know, coughing, laughing, sort of some of these, you know, audio tags in there, you can enable that as well.
And it's all kind of within, you know, structured API responses, which is really nice. So what you can see here is, for example, you know, we have a conference call.
Quick check-in. Maple Street is a mess. Time to fix it.
Totally. Some of those potholes could swallow a small car.
Or a very brave skateboarder.
We start next week.
So what you can see here is, as we're kind of playing this audio, we see that the model recognizes the different speakers and kind of tags them as speaker one, speaker two. And then we have the word-level timestamps.
So you see as I play this.
Jonas, four-week timeline?
Yep, unless the concrete throws a.
So we can highlight the different words, kind of, you know, word-level here. So that's really, really useful. And, you know, obviously this is available through the API. So actually one thing I did, and you can try this out yourself, this is just available kind of for free to sort of demo it.
So if you're using Telegram, I've built kind of this little Telegram bot where you can forward voice messages or videos to, you know, the Telegram bot. And then you, it automatically identifies, okay, what language is that? And it gives you back the transcript of that message.
So, you know, you know this all too well. You're sitting in a meeting and your grandmother sends you a voice message and you don't know, oh, is thisurgent? Is this important? So what you can do is you can forward it to the bot and very quickly you will get the transcript back.
So if you want to try this out, you know, you can do that. So you can see it here. I was actually, I hope you recognized this if you were paying attention. Anyone recognize this?
You were just saying.
Yes, I was just saying that. So I was just recording a voice message on my phone. Thank you. Someone was paying attention, that's great. And so, you know, I was just recording a voice message and then I get the transcript back.
Now, the cool thing as well, so I actually live in Singapore. So if you spend time in Singapore, you might have heard kind of the Singlish,right? Which is sort of the Singapore English.
Full throttle hour this person. Bastard, you know. You, I asked for a plastic bag, you must put the thing inside the plastic bag for me,right? She never, she just put a plastic bag, throw the plastic bag on the table.
Anyone understand what's going on? Do we got any Singaporeans in the house? No? It's not that easy. Even after six years in Singapore, I still sometimes struggle with that. So what we can do is we can forward it to our transcription bot.
We can see, okay, it was received. It's transcribing it now. And yeah, go to hell, bastard. You know, I asked for a plastic bag, you must put the thing inside the plastic bag for me. So you can see these are kind of the problems we have in Singapore.
Now, another accent that might be a bit challenging,
Scottish.
Ole!
Some things I need to know about Scotland.
Eh, no, you need to know that, eh, that'sright, eh, split up the country. And all that. We don't want that. We want our addition council and all that, and that's basically that. That is pretty.
Anyone understand what's going on in Scotland? No? Okay, no Scottish people here. Again, we can forward it, we can see. And so this is really cool, actually, that, you know, even without fine-tuning the model, it can actually understand specific accents quite well.
So you need to know they're trying to split up the country and all that, and we don't want that. We want a decent council and all, and that's basically that. There you go. I actually now listen to this so often that I can actually hear it.
But yeah, you might not be familiar with it. Now, one example that's maybe a bit closer to where you're living, if you're here in the US.
And Mr. President, what you would like to say about the Bangladesh issue, because we saw and it is evident that how the deep state of the United States was involved to regime change during the Biden administration. And then Muhammad Yunus made Junior Soros also.
So what is your point of view about the Bangladesh?
What is the role that the deep state played in the situation in Bangladesh?
So what you're seeing here is the President has a translator for English to English translation. So we thought, you know, maybe he can just use this transcription bot. So if we forward that again. So even if the audio quality isn't that great, if there's like a lot of background noise, the model is actually very good at kind of identifying that.
Yeah, Bangladesh issue, United States, there you go. So yeah, so this is kind of the, you know, one component to it. So we need to, you know, understand what the user is saying and then feed that into our LLM to, you know, have kind of a meaningful conversation.
So for the other part, you know, we don't specifically provide the intelligence layer. So that's where we partner with kind of, you know, the leading model providers. You can also fine-tune your own model. You know, if you, for example, fine-tune and deploy it on, say, Google Vertex AI, you just need an open AI API compatible endpoint.
Voice & LLM17:05
And then you can plug in your custom LLM into this pipeline as well, which we can look at in a little bit. And so now we have the other component. So once the LLM starts streaming out the response, we then want to start streaming the speech as soon as possible.
So this pipeline is kind of streaming throughout. So we have, you know, the fastest kind of snappiest response possible. So for, you know, the actual text to speech, what's great with ElevenLabs, you know, we heard we want kind of Brazilian Portuguese accent,right?
So we have a huge library of voices that are available on the platform. So we actually have more than 5,000 different voices that you can choose from. And so if you go into your ElevenLabs account, you can go to voices and you can explore different voices.
Now, if you were sitting, you know, here in this workshop and you were like, oh, I really like this voice, you're in luck. So you can go to the voice library and you can type in German engineer. I'm originally from Germany.
And you find me.
True success is doing what you are born to do and doing it well.
Does that sound like me?
Ah, okay. It's trained on some of my YouTube videos. So I think maybe I talk a bit differently in the YouTube videos. But yeah, this is great. And the thing is, so I basically cloned my voice and I published it on the voice library.
And so anytime you use that voice, I get royalties. So this is a marketplace. We actually recently just surpassed the $5 million US dollar milestone that we paid out to, you know, our voice actors that are kind of publishing their voices on the platform.
And so, you know, in this workshop, if you use this voice, I'll be very grateful because then I can have a coffee later. That's great. No, but obviously, you know, you can use my voice if you want to.
You don't have to. The great thing is you can set really kind of narrow filters to find sort of the voice that, you know, you want. So, for example, if you want to choose Portuguese, you can just put the language filter to Portuguese and then you can choose kind of the accent.
So here, for example, we want a Brazilian Portuguese. We can then further kind of narrow this down in terms of, you know, sort of gender, age. So there's certain meta tags that we can apply. And then we can see maybe here.
Okay, I don't know. I haven't spent much time in Brazil.
Duas amigas partem busca de sossego em uma fazenda orgânica, longe do barulho da cidade.
Does that sound Brazilian?
No? Yeah? Okay, a little bit. So, I mean, there's a lot of voices available on the platform. So you can kind of see if you find one that sort of fits, you know, the local accent that you're looking for.
And then what you can do is you can go and kind of put, you know, all these different pieces together into your conversational AI agent. So here in the dashboard, we can go to conversational AI and we can actually configure a lot of our agent, you know,right there within the dashboard.
Agent Demo20:32
And then we can bring that into our application with the JavaScript SDKs, the Python SDKs, kind of depending on what applications you're building. So we can just do a quick demo maybe of an agent that I had built for a conference in Singapore.
So, you know, if you're familiar with Singapore, there are four official government languages. So if you're building applications, you know, for Singapore, you actually need to, you know, provide English, Mandarin Chinese, Malay, and Tamil. So these are kind of the official government languages in Singapore.
Now, obviously, there are other languages being spoken, you know, Hindi, Japanese, for example, you know, as well. But what you can do is you can, so we currently have within conversational AI, I think 31 different languages that the agent can speak and switch between and identify.
We're working on the next iteration, so the v3 of our multilingual models, which will bring, I believe, up to 99 different languages. So yeah, you know, no matter where you are in the world, you can kind of get there pretty good coverage.
So you can see here we can configure the LLM. So we can, you know, choose kind of any of the, you know, big LLM providers that are out there. We have directly integrated and kind of co-located with our platform.
Or you can put in your custom LLM, which is just an open AI API compatible endpoint. You can put in kind of a knowledge base. You can upload documents. You can, you know, reference your website. There's RAG built in as well.
Depending on kind of the model you're using, you might not want to put kind of the entire knowledge base into the context. So you can enable RAG there. And then the tools, you know, this is kind of your standard function calling, you know, that you might be familiar with when you're working with large language models.
So depending on the model you choose as well, there might be some nuances with how reliably the function calling works. But, you know, kind of the larger ones, GPT-4 plus Gemini 2.5, you know, these are generally very reliable with function calling.
And then we have the system tools that are kind of built in. So you can see here the language detection system tool. So we can just enable this. And we have this kind of, you know, configured by default to identify different languages and switch between them.
But you can also override kind of the LLM prompt to sort of customize that to your use case. And then maybe we can just give that a little test and see how that works.
Hi there. This is the API Day Singapore conversational AI agent. How can I help you today?
Oh, 不好意思 , 我不说英文 . 你可以说中文吗 ?
当然可以 . 请问有什么我可以帮您解答的问题吗 ?
Can you tell me what's the name of that ElevenLabs speaker?
The speaker from ElevenLabs at API Day Singapore is Thorsten Schaeff, who is listed as a developer advocate. Would you like to know more about his session or background?
Yeah, but what languages can you speak?
I can communicate in English, Chinese, Hindi, Malay, Japanese, and Tamil. If you prefer to continue in any of these languages, just let me know which one you'd like to use. Would you like to switch to another language now?
Yes. Can we switch to Hindi, please?
बिलकुल. अब हम हिंदी में बात कर सकते हैं. क्या आप थॉरस्टन शेफ या इलेवनलैब्स के सत्र के बारे में और जानकारी चाहते हैं?
Okay, unfortunately, I don't speak any Hindi. Can you repeat that in English?
Of course, we can continue in English. Would you like more information about Thorsten Schaeff or the ElevenLabs?
So I guess you get the gist. So with the language detection tool, there's kind of two, you know, different modes with this. So you can either, you know, as like in the first scenario with my broken Mandarin, I was basically just saying, oh, sorry, I don't speak any English.
Can we, like, speak Mandarin? And it would recognize, oh, okay, you know, even with my terrible Mandarin, it was like, oh, yeah, he's trying to speak Mandarin. Okay, so maybe let's reply in Mandarin. So it was doing that.
Or we can specifically ask, okay, what languages can you speak? You know, can you speak Hindi? Can we switch into Hindi, please? So this is kind of the built-in language detection system tool that we can use to facilitate kind of these multilingual conversations, which is really nice.
Cool. So this is kind of roughly what I wanted to show you sort of as a start. And now what we can do is kind of we have, you know, the next 30 minutes to play around with this yourself.
Hands-on25:51
And we're in the room and, you know, if you have any questions, we can answer them. So there are various different ways that you can configure your agents. So, you know, by default, you can get started in the dashboard and you can configure kind of a lot of the behavior and functionality in there.
And then, you know, if you go back to the resources, we have, you know, the documentation. We have different examples that you can use around conversational AI. So, for example, there's, you know, we have examples for Next.js to build that into your Next.js applications.
We have examples for Python. If you want to build it, you know, on like hardware devices somewhere, you might want to use Python on like your Raspberry Pi, for example. So we have the examples there and you can then, once you've configured that, you can bring that into, you know, your application.
Alternatively, you can also configure all nuances of your agents via the API. So actually, if you're building a marketplace where you are configuring agents on behalf of someone else, you know, you would generally do that through the API.
And we also have an MCP server. So if you're using, you know, Claude's desktop, you can bring in the MCP server and you can just tell a natural language, oh, please, you know, set up an ElevenLabs conversational AI agent, you know, with this voice and it'll go and it knows kind of what, you know, API calls to make to set up your agent.
But yes, so thanks so much for joining. If you have any questions, you know, we'll be floating around. You can also ask the questions now, you know, if you want to ask it kind of in the audience. But otherwise, you know, please just go to elevenlabs.io and create your account if you don't have one already.
And then you just go to app, go to conversational AI, and then you can create here in the agent's interface, you can create a new agent. And maybe we'll just start off with kind of the support agent here and then we can go through and kind of configure this agent with, you know, your voices, your languages.
And then, yeah, we'd love to hear a lot of different agents speak, you know, at the end of the next 30 minutes. Awesome. Thanks so much. Do let us know your questions and we'll be here for the next 30 minutes to help you set up your agent yourself.
Thank you.
Did anyone have questions that they wanted to ask in the room or? I think you're welcome. There's like microphones. Yeah, there. Do you want to just go up to the microphone and ask?
Technical Q&A28:38
Hello, everyone. First of all, great presentation. The question is related to you said that it can switch to different languages without fine-tuning it. Like, what is the background process of it? Can you explain it a bit more? Like, how well it shifts so perfectly that it can interact in the regional languages as well as the accent is also similar to that kind of thing?
Yeah. So here what you can see is in my agent configuration. So I can actually, within the voices tab, I can assign different voices to the different languages. So, for example,
in Singapore, the most commonly spoken Tamil accent in Singapore is the Chennai accent Tamil. And so basically I went to the voice library and I found a voice that is a Chennai accent Tamil. And I then basically added that to my voice library and just assigned that here.
So in the voices tab, you can configure the different voices for the different languages, which then means that, you know, the actual language detection part, that is the automatic speech recognition model. So the ASR model will actually identify which language, you know, with, so it basically assigns kind of a score, like a likelihood score that this is the language that is being spoken, as well as the transcript.
And so basically we use kind of with the system tool based on the confidence score of this language being spoken, we then automatically switch to their language in the background and basically use the voice that you configured to reply with for this language.
Does that roughly answer the question? Cool.
Yeah. Do you mind coming forward? So just because I think it's also being recorded, then we have it.
Great presentation, by the way. Quick question. Other than the
changing of the language or language detection or whatnot, does it have any other ability to make any other actions throughout the call? For example, use case appointment setting. Does it have actions to maybe call a webhook and to take a look to see if there's any appointments available, either in Make or in ALN and then report back?
Like, we can have that in the prompts.
Yeah, correct. So basically the configuration of that is a combination of your system prompt together with the tools. So you can configure custom tools. And so these can be server-side tools, which then would be a webhook call, you know, to your CRM, to your system.
So this is, you know, the standard kind of tool calling, function calling that the LLM supports. So, you know, like GPT-4 or like the more modern models generally support function calling and structured outputs. And so as long as the large language model that you're using to power your agent supports function calling, you can add your tools.
And this can be a combination of server-side tools. So, for example, you know, as you mentioned, like the scheduling, you can put in, so for example, we also have an example with cal.com. You can put in the API endpoints for cal.com and then the agent can actually look up, oh, okay, is there availability in the calendar?
And it can schedule. So it can ask for like the email address and then it can schedule, you know, the meeting for all the parties kind of through the conversational AI agent.
Okay, one more question. What would you suggest would be
a good structure for a conversational agent that has like super low latency? Obviously, the model plays a big role. So obviously, like price and model or like price and the latency are like the two biggest key factors when we want to do like an outbound or inbound dialing agent.
What would you suggest for, let's say, an outbound dialing agent or even an inbound to decrease the latency to make it seem more like conversational-like? Because obviously, if you use a really good model, it's really large, the latency is just super high and it just doesn't really, that's not a meaningful conversation to have.
Gotcha. I think that's a good question. I think we have some. So if you go to the documentation within the conversational AI, we have some kind of best practices. I'm not sure if we have like specific.
I do that by testing, but I want to skip the testing part.
Yeah. Do we have specific guidance on like the best model? It kind of depends on your use case,right?
Yeah.
Although I think it's like he's asking about like the best LLM to use,right, for.
Right.
Yeah. I think it depends on your use case. Like, you know, depending on kind of the function calling and, you know, how much of that you have, you can probably go down to like, yeah, a Gemini Flash or Flashlight to like reduce the latency in terms of that.
But I think on our end, in terms of like the voice models, so we use, where are the, yeah, we use by default kind of the Flash models for the speech generation. So depending on the languages that you want to support.
Yeah.
That's just a test. I think testing would, I think I'll be able to figure out with like proper testing with different models, Flash on, off, all that good stuff. Thank you so much, man.
Cool. Thanks.
I have a couple of questions. So number one, what's the total cost per minute coming down to?
You mean like by default? So it kind of depends what pricing tier you're on. So if you go to the pricing page, so here conversational AI. So depending on kind of the tier you're on, you know, there's a certain amount of minutes that are included.
So currently the pricing is based on call minutes. And then depending on which tier you're on, there are additional minutes that are charged at a specific price. So it does somewhat depend on
kind of the pricing tier that you're on.
If you want to have like an application where you need like really long interaction times, say you want to do a companion that will talk to you while you cook, is there something you can do to mitigate that cost?
The timing cost? Yeah, that's a good question. I think like for those use cases, it might not be great at the moment because like if they are just running this like 24/7 to like be able to talk to someone, the cost is pretty significant.
That's a good question. I think we, I mean, there's potentially like if you reach out to the sales team, there is some like custom pricing that we can do based on the use case. But I think for now this is charged based on minutes used like that the session is live.
Okay. And final question. I tried to do an agent in your dashboard and it had several tasks. So it was one onboarding task, then another follow-up task. And it kind of got confused. Like it was not like identifying when to do one task, when to do the other.
So is there a way to mitigate this or maybe have a multi-agent configuration?
Yeah. So there's, we have what we call agent
to agent, agent to agent transfers. So one way to do this is to, you know, basically set this up. And so this is a system tool as well where you basically set up different agents for different use cases.
So like certain use cases, you also might want to use a different LLM to power that use case, kind of, you know, depending on which LLM is sort of best for the task. And then you can configure different agents for different tasks.
And then you can configure kind of an orchestration setup that basically will then route kind of in the background to a different agent. If you keep the voice the same, this actually happens somewhat like silently without the user actually knowing that they're being transferred.
So it's not like it's an immediate transfer. It just means that you can, you know, sort of develop these agents potentially also across teams where you have, you know, one team that owns kind of this specific agent. And then in the background, you just kind of switch between the agents for like the different tasks.
Great. Thank you.
Thanks.
Cool. Any more questions?
Yeah. Latency. So in general, like simple examples usually, you know, great for demos, but when you have something more heavy, more enterprise grade and you have more data, more, you know, your RAGs taking time to come back, how do you not have a shitty experience?
Because, you know, think of it from the end user's perspective,right? They're like, oh, it was like a then blank for the next minute or two. Do you suggest fillers? Like, you know, does the conversational AI say, I'm thinking, let me think about it?
Or how do you kind of make it more natural? Because there is going to be like in an enterprise setup,right? Like if I have to then go look up a patient's claim, you know, that might go hit a database.
Once that comes back, it goes to other systems, does a whole bunch of things, but it takes time,right? Like how do you set it up so that, you know, latency that can't be avoided? How do we make the conversational experience better?
Yeah. So there's certain things that you can do. So, you know, generally kind of these are, you know, if you have like a big knowledge base, for example, like using RAG is kind of one way of kind of mitigating, you know, that's taking too much time.
And then also in terms of your tools, when you define your tools, you can configure the, where is it? Kind of the response time. So when you, you know, add parameters,
where was the configuration? I think there's a configuration for, yeah, like the timeout. So basically how long you want to wait for, you know, this tool to sort of come back. And also if you want to wait sort of for the response.
So I think the maximum timeout we allow is like 120 seconds. And then the agent will actually like say, oh, I'm currently looking that up in the system. Sorry, you know, we're still waiting kind of on the response.
So that is sort of built into the tooling here. Now, you know, depending on your, yeah, like use case, you probably want to put that kind of, you know, fairly low because like, yeah, if you're waiting on the call, I wonder if you can do something where it's like, oh, I'll call you back, you know, like I'll take that back and like take the action or, but I think for now it would just, yeah, depending on your timeout time, it basically will wait for the response and it will tell, like it will stay conversational to like, you know, talk the user through that, oh, we're still waiting on the tool response there.
And are all conversations linear or can they branch off, come back? Like while it's looking up something, can it come back in five minutes later? Oh, you know, I found, you know, in the meantime, they get other details from the customer or the patient or whoever.
No, that's a good question. I don't think so. I think, yeah, currently you would, I think in that case you would like orchestrate it in a way where you put it into a queue. The only thing then is like how do you update the agent?
I think like if you're using the, I will need to kind of look this up. There might be a way like with the web sockets where like as information comes back, you can inject that additional information into the conversation through a web socket notification.
Okay.
But yeah, I would need to look up kind of that specific use case because then what you could do is you put these tasks into a queue and you work through them in the background and like as kind of the responses come back, you can then inject them back into the context through kind of web socket events potentially, but would need to double check if that's feasible.
Okay. That's all I have. Thanks, man.
Cool. Yeah, thanks. Sweet. Is everyone, while we're like going through the questions, is everyone setting up their agents? Are we? No? Okay, good. Yeah, I see someone nodding there. Fantastic. Okay, cheers. Yeah.
Advanced Q&A43:56
Hello. Could you explain or demonstrate what happens if you have multiple languages within an input and output and how that might work if you have different accents for each language?
Yeah. So
you mean like if the person is kind of mixing different languages?
Yeah. Like for instance, if they say, explain the difference between Schadenfreude and 心在乐活 in Chinese.
Yeah. It's a good question. I haven't tested that specifically.
Maybe you can tryright now.
We can. Yeah. So what was it? Explain the difference between Schadenfreude and.
心在乐活 .
Ah, okay. 心在乐活 .
Great. Yeah.
Okay.
Hi there. This is the API Day Singapore Conversational AI Agent. How can I help you today?
Can you please explain the difference between Schadenfreude and 心在乐活 ?
Schadenfreude is a German term that means the pleasure or satisfaction one feels from witnessing another person's misfortune or failure. It is a well-known concept and is often used in English as well. However, 心在乐活 does not appear to be a recognized term in German or in psychology, and there is no widely accepted meaning for it.
It's possible there may be a spelling error or confusion with another term. Could you clarify what you mean by 心在乐活 or provide more context? That way I can give you a more accurate comparison.
It might be my pronunciation, but it's meant to be a Chinese term.
Thank you for clarifying that you meant a Chinese term. Unfortunately, 心在乐活 does not sound like a recognizable phrase in Mandarin or other major Chinese dialects, at least based on that pronunciation. If you can provide the Chinese characters, a different spelling, or more context about what the term describes or when it is used, I may be able to help you identify.
Yeah, it might be my pronunciation, but yeah, the problem is like if you mix
like, you know, well, English and German at least are both like Germanic languages,right? But then Mandarin is. So I think there might be challenges there. And like depending on the more languages you mix, it does get challenging. Yeah.
Do you have a recommendation then if you want to build like a language learning application, for example?
Yeah.
I wonder if there's certain things you can do with like the prompt, the system prompt in terms of like improving how it's being picked up. But
yeah, I think because we're like going through text here, like the language learning use case is a bit more challenging, especially if you're going, you know, Germanic languages versus,
yeah, it's a good question. I don't have an immediate answer for you there, but
yeah, actually you might want to try like a sound token to sound token like OpenAI Real-Time. I wonder if in that case it does better because it can, you know, it doesn't go through text.
Yeah. It does, but sometimes it like switches the accents too, which is kind of annoying because it will try to pronounce Chinese, for instance, in an English accent.
Ah, interesting. Yeah. So yeah, there's challenges with that. But yeah, that's a good one. I'll take that back and see, you know, kind of how we, so I know we have some, we have a customer in India, Supernova, that does, but it's specifically English learning for the Indian market.
So I think it's a bit of a different use case there.
Do you know if it produces English in a particular accent or is it like using the, you know, Indian phonetic sounds too?
I think there's like a case study, ElevenLabs, Supernova. So maybe you can look that up. There's a video. So maybe, yeah.
Yeah, I'll look it up. Thank you so much.
Maybe take a look at that and then we can, yeah, if you connect with us, we can also follow up on kind of some guidance on that use case specifically.
Perfect. Thank you.
Cool. Thanks. Allright.
Hey. I was curious if you're worried about scammers or fraudsters using these tools.
Yeah. So there's definitely, you know, a worry with that. Like obviously kind of all this technology. So one thing that is kind of very important for us, so you can, if you go to 11labs.io/safety, you can see kind of the safety tools that we're developing, you know, in parallel to our features.
So there's a bunch of things that we do, like specifically, you know, we do like live moderation for certain things. So actually when you publish your voice to the voice library, you can specify terms that you don't want your voice to say.
And then in this case, we actually have live moderation where we will make sure that your voice isn't used to generate kind of specific terms or sentences. We also kind of monitor in general what's being generated on the platform.
So kind of the moderation and sort of the other toolings that actually with any speech that is generated on our platform, we watermark it actually to the extent that we can trace back which account generated this specific speech.
So if we identify fraudulent activity, we can actually trace back which account generated kind of that and, you know, can kind of ban them or, you know, provide kind of information to the authorities as needed. Yeah. And so the other things is just kind of in terms of for we have like, where was the, we have like the voice capture that we developed when you are creating a professional voice clone.
We actually generate kind of a random sentence that you need to read out to verify that, you know, you have permission to clone this voice. So yeah, we do, you know, with kind of all the technology that we develop, we do put quite a large amount of, you know, focus and effort into safety tooling.
But yeah, there is obviously always a concern that your technology is being used for fraudulent activity. But I think so far, you know, we've been trying to mitigate that with like the safety tooling.
Definitely. Yeah. Looks like a lot of good guardrails in place. Thanks.
Thanks.
Yay, you're back.
Yeah, we were wanting to go. So just building on her one, like in our place, if a patient is asked, hey, you know, how are you feeling about this? And say they are, they try to speak in English, they might hold a part of the conversation in English, and then they might jump to Spanish, Spanish and Portuguese, and come back to English for, you know, like when they have to describe something that they can't in English, they kind of jump back to, you know, the language that they are most comfortable with.
Sometimes they jump around different languages.
How would, because like the way you explained it, it kind of, you have some kind of a router that checks what kind of language it is and then shoots it off. But within a conversation, they kind of jump between like, you know, they'll explain a few things and then they'll go few words with a very, you know, Portuguese words and stuff.
So like for you, this is like English, Spanish, Portuguese, kind of all mixed together?
Yes. Sometimes. Like if you just ask them, how are you feeling? Okay. Then they come back, hey, you know, there's this, you know, I took this medication and it hurts, you know? And then if you say, okay, where is it hurting?
How? Then they suddenly kind of, you know, they regress to whatever language is most comfortable to them to explain their thing.
Gotcha. Yeah. I mean, yeah, you can see here that like the transcript, it actually correctly identified, you know, Schadenfreude because technically it's also an English word,right? But then like on the Chinese word, it just completely, you know, well, you know, you can blame partly my pronunciation.
Probably you can blame it a lot. But yeah, I wonder if like a native speaker, yeah, I don't have exact benchmarks on like, you know, how much like the transcript gets worse the more languages you introduce kind of in the same.
So I think like generally if you have two languages intermixed, it tends to perform okay. But like if you're like now having like three different languages, it just, you know, progressively tends to get worse. But I don't have exact benchmarks on like, you know, how many languages sort of, yeah.
Okay, cool. That's good to know. That's what I came up with.
But yeah, it would be worthwhile if you have like recordings of like some of that to like put it through our
transcription model and see kind of how it performs in like identifying that. That would be interesting. Yeah. Cool. Thank you.
I'll save this question for the end because it's kind of non-related. So I worked on a project where we used ElevenLabs for the voice track of our avatar. And ElevenLabs functioned well, but we had a lot of more downstream issues in terms of like lip sync and like I think someone mentioned like slugs and timing and other things.
So is there any plan for ElevenLabs to come like, I guess, further down the stack in terms of like avatars or is that even something you're thinking about?
Oh, interesting. So you, did you build like the lip syncing model and like avatar stuff on yourself?
I know. So we, like within the like NVIDIA Tokyo stack, so they have like a stack and they have their like Riva voice model and we kind of switched that out for ElevenLabs. So they have like the full stack of like the avatar and then ElevenLabs is just the voice portion of it.
Yeah.
Ah, okay. And sorry, which stack was that? The NVIDIA?
Oh, NVIDIA Tokyo. Yeah. It's like, yeah, it'll go from
like we did everything with the GPU, but then they have like the visualization and you can just plug your voice model or in Tokyo, like T-O-K-K-I-O.
T-O-K-K-I-O. See, I'm thinking about Japan. Is it this one?
Yeah.
Oh, interesting. Okay.
Yeah, I personally don't have
experience with that one. So I know that we're mostly working with partners like Hydra and HeyGen kind of for sort of the avatar side of things.
I don't know, Paul, do you know any? No. So this is something I would need to come back to you and like look into. It's interesting. So like you're saying out of the box, it uses like an NVIDIA model for speech generation or?
Yeah, yeah. But our client, which I assume is one of your partners, we can talk about that, but like our client couldn't use the NVIDIA model.
Gotcha.
Had a contract to use ElevenLabs model. So it's kind of.
Okay. Interesting. Yeah, sorry. I don't have a good answer for you thereright now, but yeah, this is interesting. We can go back to the team and see if there's any resources that we can give you in terms of like improving that.
No worries. Thanks.
But interesting use case. Thank you.
Allright, last question.
Last question. Nice. Hi Thorsten, thanks for the presentation.
Thank you.
I have a question regarding the transcription model around adding custom vocabulary. Like at the company I work for, we use a lot of three-letter acronyms. And let's say I want to have the model read out SAP as S-A-P and not SAP.
Is there a way to tell it to do that? And is there a way to like tell it to read words a certain way and kind of nudge the interpretation of what I say towards certain words that we use in our vocabulary?
Interesting. Yes. So you have this both, so you have this use case both on like the speech-to-text, so you need to correctly identify the acronyms, but then also you need the agent to reply back with the correct. So like for the reply back, we do have, have you seen the like pronunciation dictionaries?
So we have a way for you to provide, you know, pronunciation dictionaries with like kind of phoneme alphabets to actually, you know, identify specific, you know, basically like here, tomato, tomato, I guess, tomato, tomato. And so you can provide
for that, you can provide the pronunciation dictionaries to make sure the text-to-speech pronounces, you know, the acronyms and the words in the way that you want them to. Now, the other side of like the speech-to-text, that's an interesting case.
I don't think we have
a way to like
fine-tune that specifically for different acronyms. That's a good question. So, you know anything there? No,right?
You can do a normalization layer and then you talk to it and you put it in prompt.
Oh, interesting. Yeah. So you can, to a certain extent, you can do it through the system prompt where you put in kind of a normalization layer to basically identify things in the transcript that are acronyms and then basically have the LLM sort of massage that into what you want.
I think that's what you were saying,right? Yeah. So that could be interesting to see if that works well. Have you tried it out already or?
Where I was coming from is the company is using like their own custom chatbot called Jewel and it's like the unit of work. But whenever I read transcripts, it's oftentimes used as Jewel as the diamond.
Gotcha.
And so that's kind of the struggle that I'm facing.
Okay. Is this actually at SAP?
Mm-hmm.
Nice. I'm an SAP child myself. My father was early as, well, okay. Anyway, too much information. Cool. Yeah, thanks for that. Well, we can chat some more and see sort of if that's something we can get going. Sweet.
Yeah. And with that, we're at time. Yeah, thanks again. Thanks so much for joining. Please do, you know, connect, find the resources, fill in the form for the credits. Yeah, I'll leave this up in case you haven't had a chance to scan it.
But yeah, thanks so much for joining. Enjoy the conference. And we will also have a booth at the expo. So if you come up with some more questions, you can come find us there. Thank you. Danke schön.





