AIAI EngineerJun 1, 2026· 33:28

How to talk to statues — Joe Reeve, ElevenLabs

Joe Reeve, an ElevenLabs growth engineer, built an app that lets users talk to statues by pointing their phone at one—it identifies the statue via OpenAI deep research, generates a matching voice with ElevenLabs' voice design API, and starts an ElevenLabs agent conversation—all in 30 seconds. He created it in two hours on a Sunday using Cursor and a single prompt, posted it on Tuesday, got 50,000 impressions, then reposted about vibe coding and hit 1.5 million. Museums, auction houses (Bonhams, Christie's), and travel platforms reached out; one CEO tracked down his WhatsApp, saying a team of 10 had worked on a similar project for a year. The episode explores voice interaction patterns, the challenge of interrupting agents, the need for multimodal interfaces (voice plus visual), and the potential of vibe coding to democratize software creation, with Joe noting that music and captions are key to viral videos.

  1. 0:00Statue App
  2. 5:38Scaling
  3. 9:51Voice Problems
  4. 13:19Vibe Coding
  5. 19:22Info Density
  6. 24:50Voice Patterns
  7. 29:12Viral Videos

Powered by PodHood

Transcript

Statue App0:00

Joe Reeve0:16

Can I get a vibe check of the room? How's everyone feeling? Are we like, we really want to hear what Joe has to say about statues, or are we like, I just want to, like, chill and just sit in a quiet room?

How are people feeling?

Host0:27

Statues.

Joe Reeve0:28

Statues! Okay, that's good to hear. You're much more lively than statues, which I've spent a surprising amount of my life with recently, so. Um, this I'm going to— well, actually, I'll quickly introduce myself. I'm Joe. I work in the growth organization at ElevenLabs.

Hands up if you've— actually, I've got a slide for that. Hands up if you've ever used/heard of ElevenLabs. Okay. So this next slide, you'll probably be familiar with a lot of it. ElevenLabs does— basically, we're an audio AI foundation model company.

So everything from text to speech. You put some text in, you get some speech out. Transcription, the other direction. Music. We've got the first commercially legal AI music generation. We licensed all the training data behind an API. Sound effects.

Create voices. This one, by the way, just like pro tip, in case you're ever using ElevenLabs at a hackathon or something. This section, creating and editing voices, is like the thing that people really just don't use ElevenLabs enough for.

And this is the thing that the statue app that I built that we'll talk about is built on. And then agents. This is our fully managed agents deployment platform, which sounds like a mouthful of SaaS, and it sort of is, but it's also very cool in various ways.

So who saw the statue app? It went quite viral, at least here in London. Raise your hands if you saw it. Okay, that's fine, because I'm going to play you a video, because people like me love playing videos with our own voice out loud.

It's— I'll just play the first 30 seconds or so. I made an app that lets you talk to any statue you want using AI. So we've come to the British Museum to see how it works.

Guest2:10

I am Pharaoh Amenhotep.

Speaker 42:11

I am Demeter.

Guest2:13

Hoa Hakananaya.

Speaker 52:15

I am Hans Sloane.

Speaker 62:16

The guardian lion.

Speaker 72:18

I am the young rider.

Speaker 82:19

And I am the war horse.

Joe Reeve2:22

First, take a picture. Okay, so now I'm about to explain that bit of me explaining. So we, uh, basically what this statue does is it lets you take a picture of a statue. Sorry, this app lets you take a picture of a statue.

It then does an OpenAI deep research on the identity of the statue, generates a bunch of the sort of historical knowledge, and prompts for what it thinks the voices of those individual statues would have been if they were alive.

It then uses our voice design API, that really underutilized API, where you can put in a description of a voice and it will go and generate something that matches. And then it creates an ElevenLabs agent and then starts a phone call.

And that whole thing works in like 30 seconds. So you take a picture of something, you get all the, like, search research back from OpenAI, generate a voice, and start talking to an agent, to a statue, within 30 seconds.

Which is pretty fun. What sort of— if you're interested in, like, reading more about the details of how it was all built, you can scan that QR code. It's just a blog post. This was attached to the initial tweet that I made, which basically had the prompt for one shotting.

This whole thing I built in Cursor in two hours. It was sort of, sort of wild. So I built this in two hours on a Sunday, because I was sort of tired and bored, and then published the prompt through the ElevenLabs blog, made this video that you just saw, posted it on a Tuesday or something, and got 50,000 impressions.

It was like pretty, pretty good. Not bad. People on Twitter kind of liked it. And then three museums, or people who represent groups of museums, and a bunch of other, like, businesses, including like Tripadvisor competitors and stuff, who were coming and saying, "We've been—" well, actually one of them, the CEO called me.

He found my WhatsApp phone number somewhere. And called me up and said, "I've had a team of 10 people working on this for a year. How did you build this?" And so then the next day I reposted saying, "I've had a bunch of, like, interesting people.

I vibe coded this in, like, two hours." Not as a brag, just as a, "This is interesting. Like, vibe coding is so powerful. And if you are sort of experimenting with these interaction patterns, you can actually do something quite big.

Surprisingly big." It then went completely viral and went from 50,000 on the first day to 1.5 million on the second day. And it was in part because it got kicked off by the vibe coding. And then suddenly I got all these artists and creatives and everything from portrait museums to, like, Bonhams and Christie's reaching out saying, "We want to have people be able to talk to the items we want to sell."

I think what I would sort of— oh, and then this led to something called ElevenHacks, which we can talk about later maybe. There are loads of different things in this story we can talk about. We can talk more about the statue app and about ElevenLabs.

I sort of want to get more from you. We can talk about ElevenLabs generally. We can talk about what it means to do sort of growth, particularly API growth, at one of these companies, from the growth engineering point of view.

And we can talk about the sort of implications on culture, which I think is something that we're not, as an industry, we're not really looking at as seriously as we could be. Vibe coding generally, and again, what that impact is having on society, voice interaction patterns, or making viral videos, or anything else.

So it's sort of— yeah, please.

Scaling5:38

Host5:40

I mean, it's really easy to prototype these days,right?

Joe Reeve5:43

Yeah.

Host5:43

But it's like pushing it into production, that's the hard part. So maybe— I don't know, I'm sure there is about how you would take your really successful prototype to, like, how can you scale it to millions of users,right?

Joe Reeve5:54

Yeah, well, that's one of the things— this is going to sound very ElevenLabs salesy now. The nice thing is that pretty much all of, like, what I've done is stitch together existing APIs, which are designed to scale. So yeah, if I wanted to start doing user management, that sort of thing, that's relatively, I think, well understood, and you can buy another third parties for that.

But the hard bit, maintaining the agents and the voice design, that's all APIs that there's no way I can make a dent in the API volume, even if this goes absolutely gangbusters. So from that point of view, I think— and I think that's something that vibe coding is really showing, is that the glue pieces, and telling a good story about the glue, is in part the most important thing of the project, rather than the solving hard technical problems.

So yeah, obviously there's a lot more work to do to make it actually production ready, which is something that I've been talking to a bunch of the museums about. We might be doing that as, like, a nice ElevenLabs gives something nice to the museums.

But it's— the hard bit is actually not the user management and, like, Cursor again, can pretty much— you need to check it all, obviously— can pretty much one-shot that with Supabase or whatever else, for logins and magic links and that sort of thing.

It's mostly the relying on our APIs and our agents platform to do the heavy lifting, I suppose.

Host7:05

What about evals?

Joe Reeve7:06

Evals, I get— so that's one of the big things. I think the take a photo and just get your research back is not really the long-term solution for the museums. Really the important bit, the hard bit is going to the curators and saying, "Curators, can you figure out what is the actual narrative?"

Let's not just take random things you found from Google. Let's actually, like, put some thought and design into the content. So that's, I think, the piece that's the longer tail. The nice thing is a lot of the museums have, they see their core IP as these databases.

So they have APIs often, and we can pull that information out. The VNA has a public API for their stuff, so.

Host7:46

Yeah, I guess on this topic, maybe related to AI culture and voice interaction patterns, what is the interface then that you enable the curator to design the experience?

Joe Reeve8:01

Yeah, soright now there's no— I haven't designed anything for that. I mean, at best they could log into a dashboard, the ElevenLabs dashboard, and make edits to the system prompt and the knowledge base files that are in there.

The— I think probably in this sort of information management, the best case, or the best interaction pattern is probably editing text rather than speaking. Although obviously if you're choosing a voice, you need to sort of manage that. There's an interesting question there.

I've been talking to a chap, Jaygo, who used to be the head of the Americas at the British Museum and now runs the Sainsbury's Center, which amusingly is the location for the Avengers headquarters in the Avengers films. He is going through this big, long academic process of figuring out what should a voice sound like for an inanimate object.

So things like, where did the materials originally come from? It came from some mountain in China, or it came— and then the rock was shipped to Vietnam, and then it was carved in Vietnam, and then it spent the last 200 years living in a British museum.

So like, what should it sound like? It would have a little bit of a, maybe a Chinese origin with some Vietnamese twist in there, but then it's just lived around people with British accents, but maybe also not, because it's lots of tourists.

So actually thinking through from a much more philosophical point of view, what should objects sound like? And that's something that, from my point of view in ElevenLabs, is really interesting, because I think we have the opportunity to give all sorts of things voices.

Like elevators. A lift should probably— I mean, they do have voices. They're often quite discordant with what a lift is, I find. But it's quite likely, I think, that we start walking into lifts and saying, "I want to go to this floor, please," using voice to interact with them.

So what should they sound like? I think that's becoming a more important question. I don't know exactly what the answer is, but smarter people are doing that, so. Sorry, yeah.

Voice Problems9:51

Host9:51

Yeah, so it was really interesting to hear about your thoughts about voice as an interface to applications. In general, what problems arise? What do you think? How do users interact? What do they expect from this voice agent? And when does it not work?

What problems could we expect if we are going to build any voice interface?

Joe Reeve10:17

Yeah, I think currently voice interfaces, voice interactions have quite a large range of problems currently. But a lot of them are solvable. One of them is you basically have a binary. You're either interacting with voice or you're interacting in some other way.

And I still feel like the sort of interactive or generative UI plus voice is something that we still haven't seen. Particularly we've got— sometimes it's, and this is something I've experimented with, is like, if you think about a coding agent, you've got Lovable or something.

I want to be able to talk to my app, talk to Lovable, but not talk to the coding agent part. I want to talk to like a product manager agent that then goes off and triggers my coding agent to go and do things.

So what does that— so the voice interaction there is not direct the thing that I'm talking to, or the thing that's doing the work is the thing I'm talking to. I sort of want to be talking to a halfway house person.

So what are the problems? Aside from often the thing you end up talking to is not the thing you actually want to be talking to, there's also the, what are the parallel sort of interaction patterns in UI? And I think actually you were just showing me your app earlier, which does show the stuff that the voice agent's thinking and sort of extracting from the conversation, and allows you to interact with that at the same time.

That's something that I think we're going to see a lot more of, the sort of multimodal conversations where it's voice and visual. The other thing is people don't like interrupting voice agents, because they're too polite. People are too polite.

And I'm starting to learn to just, like, interrupt agents much more aggressively. And that actually makes the experience much better. But I don't know how to give people permission to interrupt. Sorry.

Host11:51

How do you solve the problem of, like, prompt guidance or skill loading? So, I mean, with a typical coding agent you can give, like, skills, this, this, this, you can throw everything at it,right? But I don't think it's the same interface, or is it the same interface?

Joe Reeve12:06

So the ElevenAgents platform doesn't really support the concept of skills, though it could. It does support the concept of knowledge files, which do get loaded in. So you could do it in that way. We also support MCP calling, so you can have knowledge embedded in those, or skills embedded in the MCPs.

The— I don't think that's core to voice or not. I think that's mostly down to, like, the interaction patterns of coding agents are quite lend themselves to skills. But you could have a voice agent that then has the ability to use skills.

And you can add voice capabilities to an existing coding agent in that way. So I don't know if that's related directly to skills, but.

Host12:47

Basically I want to talk to Claude as a human.

Joe Reeve12:51

I see, I see. Well, actually some people, some people have built the— particularly with OpenClaw, actually, there's ElevenLabs is quite a common interaction pattern for OpenClaw. People have built phone numbers they can call, and then it'll call them back and that sort of thing.

And then obviously you can, if you say, "Well, I want you to be able to load in skills," it just, like, learns how to do it and adds that capability to itself. So that's for sure possible. And people are doing it with their sort of OpenClaw setups and some Claude code setups.

Vibe Coding13:19

Host13:19

In this experience, more on the AI culture and the vibe coding, thinking about new interaction patterns to engage with history, with our built environment, like, where or how do you see that starting to— because I mean, we're so empowered now with this vibe coding.

It's still there's a barrier, but how do we build that engagement and do it well? I don't know what the museums are telling you. I mean, I'm sure this is so new, but.

Joe Reeve13:45

Yeah, I think, to be honest, the museums— I met with the CEO of the Science Museum, or co-CEO of the Science Museum, and they're asking the same questions. They don't really know the answers. They're saying, "Well, I mean, and the Science Museum, for example, is really good at going and trying stuff.

So they've got all these big tablets that kids can go and interact with. But in my personal opinion, a lot of the time that can sort of feel like sticking technology onto the thing, rather than it being a core part.

So some of the stuff we're experimenting with here is, obviously you've got to take a picture of a statue and you talk to it. We're commissioning a statue to be made that has the technology inside of it, and a speaker and phone and microphone, so that you can then talk directly to the statue without having a piece of technology in the way, without it feeling tacked on.

And that's something like with the red phone booth you may have seen here on floor three. There's a phone you can pick up inside of a K6 Red London phone booth, or British phone booth, and talk to an agent, talk to Sir Michael Cain.

So that's trying to put it into the real world a little bit rather than having it go through a screen.

Host14:47

So I also imagine, I mean, with vibe coding, like you can really see kids making their own games. Like you imagine a kid go to a science museum, they have their tool, and they just start creating whatever experience they want as they engage.

Like it just sort of explodes the possibility of.

Joe Reeve15:02

Yeah, yeah. I mean, I think even more generally there's a question of what's the vibe. Vibe coding, I feel like, still hasn't really gone consumer mainstream. It's like, even Lovable sort of feels like it's targeted at consumers for building, effectively building B2B SaaS apps.

You know, it's like, have Supabase and, you know, standard design components. But this is why I really love vibe coding, evening vibe coding events. They feel like the OG hackathons, because people show up and they've never even thought about writing code before.

Sometimes I go to them and I talk to people and they're like, I say, "So what's your favorite app?" Or like, "Who made the app?" And they're like, "Wait, people make apps? I thought they were just there on my phone, or on the app store,right?"

They hadn't even thought through the fact that people have to make them. So that's something that I find vibe coding events really fun, because people come in and they type in, they have no idea what, they have no idea what a hamburger menu is, or like an accordion is.

So they just say, "I want this, and I want this, I want this." And they get something completely wacky, because the other time just says, "Yeah, okay, I'll try it." When if they were talking to a software engineer, I would have said, "Ah, you want one of these and one of these."

So I think at some point we're probably going to have like, "What's the Instagram filters moment for vibe coding?" Or the TikTok moment for vibe coding. I think we're going to have something like that that makes social vibe coding much more.

But I don't know what it's going to look like, but worth experimenting.

Host16:18

Do you see anyone doing it well? Like in that sort of experimentation?

Joe Reeve16:22

There's SpielWork. There's sort of an app, a mobile app for vibe coding games, and it's TikTok swiping. And there's, I think, I can't remember what it is, there's a London-based game vibe coding tool that is, again, focused on games.

I don't know that games are really the thing, because they're quite complex. But there are a few people experimenting. But I don't think there are that many people really deeply pushing the boundaries of what's possible. They're mostly like, "Lovable, but on your mobile phone," is my opinion.

Does anybody else have any sort of, see anyone doing good consumer vibe coding?

Host17:03

Vibe coding to what end? In terms of like creating things, or copy, or for their work, or just in any sort of sense?

Joe Reeve17:10

I guess content creation.

Host17:12

It sounds like more into the vibe building in a sense, like vibe working.

Joe Reeve17:15

Yeah.

Host17:15

Vibe, you know, interacting with digital systems, the phones, whatever it is.

Joe Reeve17:19

Well, in your game you can like spawn things in with voice,right? And that's like a great, it's not quite vibe coding, but it's still interacting.

Host17:26

Because I think that's what it, like vibe coding still means you have an understanding of certain primitives. Like even what you just said, thinking about databases, I think when this goes mainstream, you're not, people aren't thinking about this way.

And that's what's so exciting for me, when you think about culture and what you've built, is when it gets to a point where just things that us as engineers wouldn't even, it's not even how we would approach the problem, and that's where some incredible creativity is coming.

Joe Reeve17:50

So the pattern that I think is closest to being a winner in this space is the Facebook Instant Games API, which doesn't exist anymore. They deprecated it. But it was in Facebook Messenger, you could play these games, and it had these primitives for social gaming.

There tended to be quizzes or like Fruit Ninja, and you'd compete with your group chats and things. I built, I bought for 15 pounds a Fruit Ninja clone off a website, instrumented it with the Facebook Instant Games API, which was this beautiful like JavaScript async/await, had a get user information, get friends, create a leaderboard, or get your position on the leaderboard.

Some very basic data storage, key value storage, and

async/await at show a rewarded ad and show an interstitial ad. And so those things allowed you to make very quickly and easily a social graph-enabled, ads-enabled experience for consumers. So I bought this thing, 15 pounds, instrumented it with Facebook Instant Games API, and went to bed.

The next day I woke up with 15 million users on this random, yeah. I mean, it didn't make very much money, but they were like 15 million users in Vietnam. Because obviously Facebook, they test everything out in the sort of lower value advertising regions and then rolls up.

But that was an amazing, because you've got the social elements of people instantly sharing it around, and that I think is probably the template that's going to, something along those lines is probably going to be the thing that wins on the social vibe coding.

But I don't know if that actually answers the question.

Host19:22

On the kind of like vibe working, vibe building kind of thread, I kind of feel frustration whenever I get a response back in voice. Maybe it's just me, but like, it's almost like the input, like the information density per second isn't quite high enough.

Info Density19:22

Host19:37

Like I use a lot of voice out to just get things out of my head. Faster than typing whatever.

Joe Reeve19:42

Or voice input, I guess.

Host19:43

Voice input, but then I still feel like I need, I don't know, diagrams, or text, or some like really high density thing back from the system. So I don't know whether, how do you guys think about that, whether I should curiously hit other people's views on that, whether they agree or disagree.

It sounds like there are some nods, but that's what I find myself leaning towards, where it's like information-rich input, and then I can just speak my thoughts as they come out, and it's almost semantically understood. Put into that information-rich format, and then my intent is like, you know, spawned out and across the system itself.

So that's the pattern that at least I'm seeing and feeling, and what I almost want to vibe interact with now, my email or like everything is, and I feel like that longing. It's not quite there yet, but I feel like I have this inclination to build in that direction and to try to get information-rich, you know, get response.

Joe Reeve20:34

Response, yeah.

Host20:34

But also sort of like the input being very free-form and just my sort of raw intent.

Joe Reeve20:40

Yeah, that's something that I feel. I absolutely feel, you know, I want to speak easily and quickly, and then receive maybe a little bit of voice, but mostly this like generated, maybe it's a UI, maybe it's just diagrams, maybe it's whatever app I'm in context of.

But yeah, I want to have a parallel input, parallel output, but like single input, my voice.

Host21:00

I think we both talked about it. I find that the, I don't actually get that much information necessarily from voice, but what I do get is like companionship. So it kind of like triggers that, it lessens the loneliness feel somehow.

If I'm like talking to something and I'm getting some information, I don't feel, I might be learning more if I'm looking at diagrams or text, but I don't feel as, I feel more motivated to kind of like continue tinkering.

So there's like some interesting modalities where you feel different things if you get the information coming in from all these different modalities, at least I find so. Curious to hear how other people think about whether they think differently, or whether that's something that they also feel as well.

Yeah, this is at least my view.

Guest21:40

Well, the visual cortex is much older than voice and text sort of, so seeing something is much.

Joe Reeve21:46

Yeah.

Guest21:46

Also, it just, when you ask it to be concise, I don't feel offended if it gives a concise answer. But in speech, if it gives a concise answer, it's like, "Oh, relax man, I'm just asking a question here."

If I asked it to be concise, then it just sounds rude.

Joe Reeve22:00

That's an interesting, maybe this is possible, maybe it's not, but what does skim listening look like?

Guest22:06

Yeah, exactly.

Joe Reeve22:07

Maybe like listening actually should also have two buttons, like forwards and backwards, and I can just tap, tap, tap, tap, tap, go forwards half a sentence until I need, I don't know, maybe that's, maybe we should build that.

Who wants to vibe code something with me straight after this?

Host22:21

How would it work? So you'd just easily, yeah, back and forward with audio, exactly on what you're doing.

Guest22:28

Just like a speed dial like on a podcast,right? 1, 2, X.

Joe Reeve22:31

Yeah, or like the old iPods where you can sort of spin forwards and backwards. I don't know, that's probably quite a nice listening interaction. And you sort of, I guess you want to scroll forwards in concepts, not necessarily in sentences,right?

Like what is the thing that your eyes look at when you're skim reading? It's probably, it's not the sentence structure, it's like the words, I guess.

Host22:49

Can we move on to the next thing?

Joe Reeve22:51

Yeah.

Host22:51

Yeah.

Joe Reeve22:54

And that's actually, oh yeah, this is, sorry, this is very interesting and exciting for me. Because if I'm talking to an agent and it just starts rambling about, you know, sometimes you get back three paragraphs of stuff, and I'm like, "No, not this one."

But I want to interrupt it and say, "Go to the next one." But then it's effectively saying, it's a new prompt,right? So then it's effectively saying, "Yes, okay, I'll focus on that next one," and then it'll write me three paragraphs about the second paragraph.

You know, that's not really at all what I wanted. Unless I say, "Be concise," and then it says something rude to me.

Guest23:19

Could you like make a summary for each paragraph, and then like if you hold the thought, then it will expand, and then you can like.

Joe Reeve23:27

Yeah, I think the Claude app has done some interesting stuff on the voice interactions, because they show you something different to you here. And they show the higher level sections, and then it goes into each one and you can tap on them.

So that's, I guess, getting a little bit closer to this. But I think that also means you're not interacting as though you would with a human conversation. I don't know.

Guest23:50

I'm thinking about that, like why do we not have this issue when we're talking to humans,right? Like how do we.

Host23:57

What are so many other cues?

Guest23:59

Yeah,right there, there are just like a cue, we have a sense, but there's so many other.

Host24:03

I think there are visual cues as well.

Guest24:04

Yeah.

Host24:05

I don't know, so that's something, because it's like, if I'm, this is what it's like, if there's an agent response coming in the audio, I don't know how long that's going to be. As I think one of us mentioned, is it going to be like, you know, a minute, or is it going to be 10 seconds?

I almost want to know, and if I know it's like really long and waffly, I'm just going to like cancel it. I have some other.

Joe Reeve24:23

But also you can tell when I'm about to interrupt you. So you go faster, and you like move, maybe, and you can tell if people are listening or not.

Guest24:31

Can you make a fast experience?

Joe Reeve24:32

Yeah.

Guest24:32

Instead of a conversation?

Joe Reeve24:35

Yeah.

Guest24:35

Face to face.

Joe Reeve24:36

I guess, oh, there's another thing which is interesting here, which is the interrupting. Sometimes I don't want to interrupt, I just want to say, "Yeah, yeah, yeah," or, "Oh, but no, go back." And like, it's like, you're always listening for the, "Uh-huh, yeah, yeah, yeah."

But you can't do that with an agent.

Host24:50

I think, so I'm willing to, you know, I mentioned this to Joe before, but briefly, actually I showed you this yesterday, but briefly, the idea is that we're able to sort of interact with a PS5 game, and also create what you want in that world as you're playing it.

Voice Patterns24:50

Host25:03

So you're, you know, you're on like a, you're going to shoot a battle and playing you, we're trying to like, you know, kill each other, you know, for the worst language. But like if I'm trying to be creative and, you know, generate a new getaway vehicle or a helicopter, I can just sort of save that experience and have that appear in the game straight away, and then I can sort of fly away.

So that's sort of the work that I'm doing, how I met Joe actually. But I think the challenge that I face is that relying on just the like voice interruptability is quite unreliable. So I've just got around this with just a simple whisper flow type.

Joe Reeve25:35

Push to talk type, yeah.

Host25:36

And then hold to talk, and then let go to finish. Which it augments like purely this audio stream with some other cue. And in the same way that I think you almost need some very light nudge interface on top of the audio that you're receiving.

Maybe as you say like, you know, you're saying something and then you see a little sort of like circle appearing, being like the agent wants to ask you a question. That would be an interesting experience to feel,right? You're not being interrupted by that.

You're being kind of like, you know, the agent wants to talk. And then maybe you either stop or say, "Okay, what idea do you have?" Because I even can feel you doing that to meright now, Joe. You want to say something.

Joe Reeve26:14

I'm just getting loads of ideas. This is great.

Host26:16

I can sort of like, we're almost like having an information, you know, communication on that level. But the audio, if I'm just listening to audio and not, I don't think I'd know that,right? As if I was an agent or if I was just like, just listening to the audio, not looking at the video stream.

Joe Reeve26:31

Okay.

Host26:31

I'll let you talk. Sorry.

Joe Reeve26:32

Well, I'm just imagining, I'm just imagining on that point that the agent wants to respond to you. It could be showing, "I want to interrupt and tell you about this thing." Or this one. And then suddenly that becomes what we've just done, but with even more context than doing it with a human.

Because it's signaling the topic it wants to talk about.

Host26:48

Yeah.

Guest26:49

I was just one question regarding the product, because regarding this topic, you need some sort of tool calling in the background to be able to actually understand that I want to interrupt. And how would that tool call be handled?

Will it sort of be streaming the audio directly to the RTC to the client, or through a back channel and then getting some information? Do you update the system prompt? How do you work at it?

Joe Reeve27:22

Obviously this hasn't been, as far as I'm aware, it hasn't been built yet. The way I would probably approach it is looking at the transcript and just keep analyzing the transcript over and over again. And say, "Do you have anything to add?

Do you have anything to add? Do you have anything to add?" Or, "What would you add?" Rather than being a tool call or being part of it. I'd do this as an asynchronous looking at the transcript.

Guest27:40

Are you changing the original prompt? Because the original prompt has sort of a plan, you want to do this, and halfway you sort of change the plan, and you go in and sort of, how do you steer the original agent?

Joe Reeve27:53

Well, this is actually an interesting thing with agents. They often we see agents as things that you can't really interact with the internals of. But effectively an agent is the sort of logic that says, "What's the next message?"

But it's also a transcript. And that transcript of the conversation is completely malleable. So maybe some other thing we could experiment with is actually allowing the ability to do the interruptions where it's talking, and then I can say, "Uh, uh, uh, yeah, yeah, yeah."

And currently if I have the, the way most agents' platforms will work is it'll generate its full text thing and then start generating the audio. And then if I talk, even while it's partway through, it's still reading out the audio, the text, it will just append my message to the end of that full message.

But we do have the timestamps. We know how far through the audio it's played. So we could actually just go and edit the transcript that's coming back and say, "Well, no, they interrupted at this point, so we're going to forget that the LLM even generated more text."

Guest28:46

Yeah, it's sort of, you have the system message and you have the user message.

Joe Reeve28:50

Yeah.

Guest28:50

And sort of you need a sort of a second type of user message that you are passing in. I understand that of course it's a user message, but it's also some other kind of steering here.

Joe Reeve29:04

Yeah. I mean, why don't you, let's go to the ElevenLabs booth, the expo floor, and just vibe code something and see. This is great. I can't wait.

Guest29:12

Now it's getting easier and easier to make tools and there's so many different ways to execute your ideas. And it seems like the main kind of differentiator now is getting your word out there. Like how did you approach making those videos, and like how long did it take you for the statue video to make it?

Viral Videos29:12

Joe Reeve29:30

That's a good question. Totally separate. So this is sort of the inspiration of this ElevenHacks thing. I learned to make videos, I'm not particularly good at it, you know, it's like still relatively janky. I learned to make videos through doing other like politics-related campaigning stuff before I worked at, before I even knew ElevenLabs existed.

It turns out that with videos, random things go viral, or with content, random things go viral and random things don't. Like things I think are going to go viral don't, and then things that do, do.

I think it's basically about practicing. Like editing a video, I find to, it's like the 80/20 rule. Editing a video to the standard of the statue app is actually quite relatively easy. But then to go beyond that, you suddenly, because I edited that on my phone, the editing itself took about 20 minutes, 25 minutes.

Going beyond the quality of that suddenly means thinking about using a desktop editing tool, which suddenly makes everything, even using CapCut on desktop, to me is like three times harder than using it on my phone. And it's also three times more expensive, the subscription on the laptop than on the phone.

So I think a lot of it is just doing stuff and trying and iterating. Big things that I found, adding captions really help to the video. Having a hook in the first, you can, I mean, you look at your various platform

analytics once you've posted a few videos. My videos tend to get between 6 and 12 seconds is the average, the median view time, and then people drop off. So you need to get your hook in there, because most people are going to drop off if they don't buy the hook.

So that's an important piece, front loading the like interesting piece. Adding music makes a massive difference. And this is something that was much harder, but now with ElevenLabs music generation, I will make the video and then add, sometimes I'll make the video with the narrative and everything, and then just experiment with completely different genres until I find music.

And I'll just put the music on, does it work? Yes, no. And then you can edit the different sections so that it times up with the sections of the speech. So you don't need to like figure out music first, find a piece of music, and then match your speech to it.

The other way, sometimes I will have a vibe I want to get across, and I will generate the music first, and then figure out what's my speech that matches that vibe. So it might be like an excited theme or a, I actually chose the music for the statue app

before I made the video itself, because I thought this is a fun piece of music for a statue type thing. It's like a bit of an imperial outside the British Museum, it sort of made sense as an attention grabber.

So, but I think music is a massive thing that people underrate. Because you just put it, it's relatively quiet, but it makes a massive difference to the feeling.

Guest32:24

So you just did it on CapCut, and did you think.

Joe Reeve32:26

It's literally on my mobile phone. I've got, I borrowed my wife's lapel mic, Bluetooth lapel mic, which costs 200 quid from DJI, which makes the audio much better. And then, yeah, just edited on CapCut. Super, super simple.

Guest32:39

You didn't think about like, this is the shot I want, and then like, I want to take this photo, so it's like.

Joe Reeve32:46

I mean, the one out the front of the British Museum, yes. Thank you. Yes, I did want that one. But the, and it was actually the second time I recorded the video. I went once and got some stuff that was like a bit boring, and then I went back and just took a bunch of silly photos, and that's what ended up being the video.

So, yeah. Thanks.

Yeah. Cool. Thank you very much.

Guest33:12

Thank you.