Intro0:00
So.
Good morning, everyone. Thank you for being here so early in the morning and to be able to visit a few ones who passed security to actually be here. I'm Guillaume, and I will talk to you about GenMedia in general,
like, so this is my life as Nano Banana sees it. Basically, I joined Google six years ago. I used to be a video game producer before. I initially worked on Stadia, the streaming game company that product that we killed, as so many other products.
And I've been at DeepMind for two years now. And I worked as a, what we call, a developer advocate. So if you're not familiar with what a developer advocate basically, my job is to make sure that whenever we release things, or whatever we release, you guys, the developers, have everything you need to work with our products.
So you need documentations, you need code samples, you need demos. Now, there are new things that you need as well, like skills and prompt guides and things like that. So making sure that you can startright away and it works.
And on the other side, when I'm talking internally, that's the reason why it's advocate, because I'm advocating for the developers and trying to bring some, let's say, common sense to the internal teams, and making sure that what we release makes sense in the real world and it's not entirely developed for Google by Google.
A very good example of that is the Imagen models. When we release, when we add, like, Imagen and Nano Banana, each model adds its own set of API, which doesn't make any damn sense. Like, a normal developer should be able to just swap the model name and it works.
And I've fought quite a long time for that. I never managed to win. But in the end, I think the Imagen brand doesn't exist anymore, so I kind of win by default. But yeah, you see the kind of work I have to do on a daily basis.
As I said, I'm mostly working on the GenMedia models. So we, like, and that's why the talk is about GenMedia. So if you look at the world, your phone, and anything like medias are everywhere. So you have, like, images, videos, sounds everywhere.
So it's really core in our world. And that was at the core of what DeepMind is building with our models. We, like, you're starting to hear with LeCun's new startup about world models, but that's exactly what we have been trying to build at Google since the beginning.
And our vision of a world model is that it's a model that, as its name implies, understands the world, but meaning it can ingest as many modalities as possible. So sound, videos, audio, sensors, whatever, kind of all five senses.
And then
outputs or, like, yeah, talks in different, as many different modalities as possible. So audio, text, and so on, and much more in the future. And we tend to have specific models. We have our image generation models. We have our video generation models.
But deep down, the goal is really to have one model that encompasses all of that. It's just that for release purposes, it's easier to ship specific models and to always update the main model and, like, have risks of breaking something else at the same time.
But, like,
a quick story. Like, when we released Gemini 1.0, it was two years ago. It feels like it was, like, years ago, like, eons ago. But the first Gemini model, Gemini 1.1, was meant to be multimodal, because all of our models have always been multimodal.
But since, I guess, the testing was not finished, or whatever reason I wasn't there yet,
they removed the image understanding, the multimodal understanding inputs from the model. So the 1.1 was not multimodal. And then 1.5 came. And then this one was multimodal in. And I think that was the first one that was doing that, which, once again, is crazy.
Like, it was a year and a half ago. And now you don't imagine working with a model that is not a multimodal one. But still, when you were using that one very often, you were giving it an image and it was telling you, oh, I'm sorry.
I'm just an LLM. I can't deal with images. And that's just because some of the training that was added at the end of 1.0, that was, oh, you know how to deal with images. But if you are asked, don't use it.
Though some of that training was still remaining in 1.5 and coming up from time to time. So that was, yeah, it was kind of annoying until we switched up to 2.0.
So as I said, we don't only have the Gemini models at DeepMind. So we have all of the image, video, music generations. We have a couple of other very specific models. We have the robotics one that is kind of a multimodal one as well, because the vision is very important for robotics.
We have a bunch of agents that we are shipping. I think this year is going to be the year of agents, the real one. Last year was everybody talking about agents. This year is the year where we are actually going to build agents.
And we have the open models. And I forgot to update the slide because it's not Gemma 3 anymore. It's Gemma 4 since last week. And then a bunch of researching models, like AlphaEvolve, AlphaGenome, WeatherNext, and so on, that are very, like, very specific for research purposes.
Just
on GenMedia, we ship things on average more than every month. On the world, if you just take all of DeepMind, we are shipping things every five days on average. Some weeks we're shipping two or three things. And if we add all of the small features, we're shipping multiple times per week.
So that's why I and most of my colleagues are very busy, because we always have, like, new things to document and talk about.
So very quickly about the new things that we shipped recently. Nano Banana, we shipped Nano Banana 2, which has new aspect ratios from, like, 520 pixels to 4K. It has search grounding, as you all know. But it also has, like, image grounding.
So you can ask it to search for images on the web and use those images as grounding so that it knows, like, yeah, it has better knowledge about what things look like. It's very useful for architecture stuff, for example, for animals and things like that.
There's a lot of rules internal that I was not able to all figure out. But for example, buildings, they have to be old enough. Otherwise, there are legal reasons we can't use the images.
But it's still very useful when it works. We have the Veo models. You've heard about Veo 3, Veo 3.1. We just released last week Veo 3.1 Lite, which is our cheapest model. I think it's $0.05 per second, so $0.40 for one video, which is very cheap.
And the idea is that you can iterate
on the prompt that way and then upscale afterwards. And we have Lyria, which is our music generation model that we released two weeks ago, and with which you can create either 30-second clips or full songs of three minutes.
I will show you demos of all of that afterwards. That's the point of the session anyway. And, but there's also another Lyria model that people don't know about that is called Lyria Realtime. And that's actually my favorite model.
And this one is basically, you create music as well. But it's not a diffusion model like the others, where you just give a prompt and you get something out of it. It's a predict model, which means it's basically a live model.
So it creates music and it continues to create music in real time until you stop it. And you can just send new prompts and it will, like a DJ, mix and, like, swap to the new music in real time.
And real time, two seconds. But then, and that's quite fun to play with. If we have time at the end, I will show you a demo. But as I said, this is meant to be a workshop. So the goal is for you to play with the models and for me to show you codes instead of slides.
So this is the content that we are going to use. That's this one as well, if you can read what I wrote.
And if that doesn't work, just tell me.
Are you all in? No.
Not.
No, it doesn't work?
Oops.
Let me triple-check that the link works.
It should.
Yeah.
Ah, no. I had a typo.
Ah. Cool. Are you all in? OK. So full disclosure before we start. This is using GenMedia models, which means they are all paid models. So running the notebook is going to cost you something like $1. And you can just skip the video generation, because that's most of the cost of that $1.
So the, yep.
I can zoom in a bit.
Book Illustrating11:26
So the goal of that content I prepared is to illustrate a book using the different GenMedia models. So what we're going to do is that we're going to take a book that is an open source one that I took on the Gutenberg online library.
And basically, we are going to use Gemini to come up with prompts and then the GenMedia to create the content for the prompts so that we will have images of the characters, images of the scenes, videos of those things, and so on.
And that example comes from what we call the cookbook. So we have this GitHub repo where we are posting examples of, like, both quick start guides to explain how to use a new model or how to use a new feature, and also full-fledged examples like this one on how to go further and to mix different features into content.
So if you're looking for ideas, that's a good place to check what you can do with the models. So
basically, introduction, building. So let's start. So the first thing is that you need to install the SDK.
You need the latest one because of music generation that was shipped last week, two weeks before. So actually, you need the next one.
So it's starting.
And I should have done that. And you will need an API key if you don't have an API key from AI Studio. And you need, as I said, you need that API key to be a paid one. If you don't have a paid API key, you can still use the image generation examples by using the Nano Banana 1 model, which has a free tier.
So you can just swap the model we are going to use for this one. And why is that so slow?
I think.
Small, neat ears and thick, silky hair. It was the water rat. Then the two animals.
I think it's good.
Guarded each other cautiously.
OK. So then I'm just loading the API key. And I'm creating the client, the GenAI client. And you can see there, I did these parts that we don't add in all of our examples. That is basically the auto-retry thing, because, like, if we are using Nano Banana 2 at the moment, especially in the evening when the US wakes up, the model can be overloaded.
So that part basically says that it's going to be automatically retrying five times after two seconds.
A bunch of imports. And then we are going to select all of the models that we are going to use. So in this case, 3.1 Flash Image Preview, which is Nano Banana 2, Gemini 2.5 Flash.
Wonderful day.
Let's use 3.3 Flash instead. Lyria Clip to create 30-second music clips.
And the TTS model, this one, the pro one.
That's all that.
Yep. And I added this checkbox that should have been forced by default, but I made a mistake yesterday evening, just so that you can, you should not be able to run the notebook by mistake if you don't want to pay.
Then I'm just setting limits here, because, like, a book can have a lot of chapters, lots of characters, and that can be quite long to generate. So that's why I set up limits on how many we want for the demo purposes and cost purposes as well.
So as I said, we are going to use an open source book that's The Wind in the Willows from Kenneth Graham that I don't remember reading. But I think in the UK it's quite well known.
So basically, I'm just downloading it from the Gutenberg project. And here I'm using client file upload, which is, you might know that we have basically two ways of using the Gemini models. We have the AI Studio Gemini API way.
And you have Vertex. And the main difference between the two, and I can maybe that's a good time to switch to this slide.
But we, you know, at Google, we like to basically create multiple products that are doing the same thing and confuse our users. That's kind of our motto. So that we're doing the same with AI. So we have a bunch of, but actually, when you think about it, it can make more sense.
So on the left, we have what we call the consumer's app. So it's a fully developed app that is easy to use for anybody who is not technical. So you can do plenty of things with Gemini. But you're lacking, as a developer, you might be frustrated because you're lacking control about which models, which features, which parameters are being used.
And on the other hand, we have Vertex AI, which is the exact opposite. It's meant for enterprise. So you have a lot of control. You can control on which data center it runs. You have control about your buckets, who has access to what, and so on.
The only thing is that it comes with, like, with great powers come great responsibilities. So it can be a pain for people to start there. So that's why we have the Gemini Developer API that are kind of a middle ground for developers, where you can just create an API key and then start using the modelsright away, which comes with security risks as well, that if your key leaks, anybody can just use it.
And AI Studio is, like, is kind of the same vein that it's meant to be a place where you can test a model and play with them as easily as possible. And the cool thing is that we have the same SDK between Vertex AI and Developer API.
So you can swap from one to another. So there's no wrong place to start playing with the models, because you can always change. And then we can skip that. And I can go back here, because actually, what I was going to say is that we actually have a few differences between when you're using Vertex and when you're using the Gemini API, because, like, the goal of the Gemini API is basically to hide all of the complexity from Vertex.
And one of those complexities is creating buckets, creating ACLs for the buckets, giving rights, and all of those things. So the Gemini API has that API that is called file upload. And basically, what it does is that you upload a file and then it's easily accessible from the model.
So I upload the file. And then I will use what we call chat mode. So basically, what it does is that it sends requests and it keeps the history of it so that it's easier to keep all of the context.
And in this case, it's going to be good, because we are going to feed the whole book to the model thanks to the large context window. And that's why it's going to be useful in our case, because I'm going to run that while I talk.
It's going to be useful, because we, like, for image generation, it's always a good idea to know what has been generated previously so it keeps the same styles and things like that. I'm also going to use structured output so that we, like, we have a structure about what the model outputs and we can talk the same language, which is going to be very simple.
I want it to generate prompts. And I want the prompts to have a name so that we know if it's a chapter, a character, or something, and the prompting itself. So I'm just creating, like, initializing the chat client with the response type JSON, the scheme that I'm providing, and something that we shipped yesterday that is service tier priority.
So don't do that yourself. I think you should remove that line.
What it does is that it basically, yeah, we shipped that last week. So we have three service tiers. You have the normal one. You're paying the normal price. You're, like, in the queue with everybody else. And we have another one that is called Flex.
And that's basically, I don't care if that takes a long time, but I want to pay less. So you're going to have a 50% discount, but your request can be delayed and so on, up to a few minutes.
And on the other end, you have Priority, where you are going to pay twice the price. But at the same time, you're kind of guaranteed that it's going to be fast, because you will have the fast track, like, in the airport or anywhere else.
So yesterday, I added that. And I'm going to use that to be certain that it works well for me. But for you, you might want to save a few bucks and not add that here.
How much more expensive?
Twice.
Twice. OK.
Yeah.
And I'm not sure it works with Veo anyways, which is the most expensive ones of the model. And not with Lyria yet.
OK. Yeah, yeah.
So anyway, so I'm starting with, I've created this chat. And what I'm going to do is to send it a first message that basically says, I'm feeding you the whole book. I don't need to do anything with the book yet.
But you have it. It's in your context. And instructions will follow. And then we are going to define a style. Usually, I just set nothing. And then I let Gemini come up with its own style. But I'm getting tired of having exactly the same style always.
So let's write something.
A colorful
building block style.
Wow.
Let's go with that. Let's see how it goes. So I'm just defining a style. And by the way, if you're using Colab, this Colab has those things that they call
Magix. And that's kind of nice to create those notebooks and to have the, like, forms that people can fill. And that just fits into the code.
Then some system instructions
to direct the model, like, into the kind of images that we want. Because the problem I had at the beginning when I was working on those examples is whenever you ask a portrait image and the model knows it's about a book, it tends to create, like, cover pages and to add a title.
And I didn't want those styles or to create ones with different panels. And I didn't want that either. So that's just what the system instructions are about. And then, like, we can start working. And basically, I'm going to ask it to create prompts for each character of the book.
So can you describe the main characters? Only the adults. Actually, we could remove that, because it was just from the beginning of Nano Banana when you could not generate kids' images in Europe. But it's not true anymore. So
you can create images from nothing with kids. But you can't edit images with kids. That's a current limitation in Europe. But anyway, here's
our prompts for each character. And then we can move to creating the images for each of them. And what I'm going to do is that I'm going to create another chat just for the images, because I don't want to mix the text and the images output.
But I want it to be a chat so that it will have all of the history of the previous images it created. So I'm setting up with the responsibility image, the aspect ratio we want, the same instructions we decided and the style, and Priority, as I said before.
And here we go.
So here's the model. That's nice. Here's the water rat. And that's actually way faster now than using Priority. So this works.
I could have done it in a better way and, like, make all of the calls asynchronous. That wouldn't have worked with chat mode. But that would have been a way to make it faster. I think it's good enough.
The toad. Oh, way bigger than the car for some reason.
It's a child book.
Yeah.
And then Mr. Badger.
Yeah. So
this is mainly a demonstration of what you can do. Like, every time I run it, I'm thinking, like, oh, if I was to optimize it, there's plenty of better ways to do that.
But still. And the badger. And then the author will be at the end. One of the things I'm doing in the code, if I go up, is that I'm also saving all of the generated images from the character in an array.
And I will explain afterwards why I'm doing that. And that's, once again, just appending them one after another. In the real world, I would save them in a better way.
And the author is here. So now we can move to the next phase, which is basically illustrating the book. So same thing, like, I'm going to ask the chat for each chapter, give me a prompt to illustrate what's happening.
It should be a single image, not a multi-tagged one. I'm trying to force it to describe the character again, even though I will have the images as references, just because it's still better. So let's go with that.
That should be quite fast.
Well, same thing, like, the way it works with chats, it's keeping an history that is basically all the previous messages. And it's sending back the history to the model every time we make a call, which can be, like, in this case, since it's basically resending the book all the time to the model, that's the reason why it can be quite slow.
Interactions API27:11
We actually released a new API a few months ago that are called the interactions API. Let's run that while I talk. And the main difference between the new API and the old ones is that the new APIs are going to be stateful and stateless.
And what it changes is that every time you make a call, you get an interactions ID. And you can reuse that interactions ID in future calls. And it will recover all of the context directly from the server. So you don't need to re-upload the same context again and again and again at every turn of the conversation.
And it's also making it easier
to fork the discussion. For example, like, you want to create a song with images and with a cover image. You can create the lyrics with one model. And then you fork it. And on one end, you create the images.
And the other end, you create the song. You had a question.
How long do you store the session?
Huh?
How long do you store the session? Is it?
That's a good question. I think it's two days, something like that.
It's still in preview. But there's good chances that at I/O we make it the default API. But yeah, I'm not using it enough to know. One of the cool things it does as well is that since it knows that it's context that you're going to reuse, it's automatically caching it as well.
So it makes it cheaper to run. But even though the normal API are also doing the same. So we have our chapter. So the first chapter is next to the river. So the model and the, I forgot his name, characters are together.
Then it's on the road with the toad. And then in the forest, in the snow. Oh, scary.
What is that?
So see, and basically, what I did there is that I used the fact that we were using the history
to trust the model to have all of the previous images of the character so that it would remember how they look like and how to create new images of them.
But there are actually better ways to do that. So I tried another way, which is basically to create
a new structured output that is basically the name of the chapter, the prompt, but also the list of the characters that are appearing in this chapter. And that's actually what I was doing if I wanted to do it at scale with, like, more than a few characters.
And basically, I'm going to ask the model to give me a prompt for each chapter, but also to give me to have, thanks to the list of characters, I will only give it as references theright images for the chapter.
So I'm going to basically get the same thing here.
Yep. And so in the first image, there should be the model and the water rat. In the second one, Mr. Toad, model, water rat, gray horse. And the third one, model and water rat again. So I created a very dirty script that basically searched through the character images that we saved earlier so that I can give it a list of characters.
And it's just going to give me the list of images to give to the model.
Did I run it? Yeah. And basically, I'm going to do the same thing and go to generate images for each chapter. But this time, instead of relying on chat mode, I'm going to use generate content. So the, like, unary call.
But I'm going to pass it the images that are from the characters that are specifically
in this chapter image. So it should give slightly better context to the model about
what to display and how to show them.
Let's see if it works better. I think I've. And
like I said, like, if I wanted to do it at scale, I would have a lot of improvement I would do about that. And one of them would be that I think I would generate more than one image per character.
I think I would have been one portrait image and then one full body image. And then maybe from the side, from the back, and ask the model to tell me which exactly, like, how they are going to be displayed so that I can give the Nano Banana exactly the reference we need
for the generation. So here they are.
And it's more or less the same, to be honest. But
except that this time, I don't know why it seems to be attacking the model for some reason. I don't think that's what the story is about. You might know better than me.
We can check the prompt. What does the prompt say?
No, it doesn't say anything about attacking the no, it's rescuing him.
Captain Captain?
Yeah. Yeah, that's basically how you can
use the model. For those who came, the content I'm showing, you can open it there. And it's a collab that is about taking a book and creating images and videos to illustrate the book and the characters. So we went through creating prompts for each character and then generating images for each of those characters, and then creating new prompts for each chapter, and then creating images for the chapters using the reference that we have from the images.
If you can't see what I wrote, it's goo.gli/cookbook-illustration.
And then we are going to try to move to the next phase with videos now. So we are going to use Veo to generate videos of those images. So I'm going to swap to the bigger model, because I can pay for it.
Video Generation34:15
You can stay with the cheapest one if you don't want to spend too much. And basically, I'm going to just do exactly, like, take the last chapter, take the last image, and basically send it the same prompt and the reference image.
So that's what I do here. So when I pass an image to Veo, that's basically it's going to use it as the first frame for the video.
And funnily enough, like,
most of the, like, a lot of the training for video generation model is actually image generation, because I think the most important part of generating a video is generating the first frame so that it knows where to start with and then what to do with it.
So.
So what's the best video model? Is it.
It's the one that doesn't have light or fast.
So faster, faster, faster version of the same.
Yeah, that's.
Even more expensive than the faster version.
Yeah. So Veo 3.1 is the main Veo model. And the other ones are basically, like, smaller versions of it that are running slightly faster and, like, basically doing less generation turns.
And yes, I said I was going to repeat the question, and I forgot. The question was, which one is faster, the better model among the three.
I guess the doors just opened.
Still very early. So up, we can see where it goes.
Wait with it. Innovate with this, Gemi.
I can't get out of this dreadful place. Oh, thank you, water rat. I thought I was good before.
Yeah, it saved. And there were songs. I don't know if we can we have more songs so we can see what.
Quick, give me your hand. We must get out of this dreadful place.
Oh, thank you, water rat. I thought I was done for.
That's not that bad, except the wrong character is speaking. That's the problem with basically using the same prompt when you create the videos and the image, because it doesn't have the extra content about what exactly is meant to be happening afterwards.
So that's why I add another, I added another example that is basically we are going to do the same, except we are going to add one more step that is basically asking Gemini to come up with
a new prompt just for the video and to explain what's happening afterwards. So that's what I'm going to do. I'm going to animate
this chapter image. And can you create a prompt with Veo about what's happening in the next few seconds after the initial image? And I'm passing it the last image so that it knows exactly where to start.
And while it runs, since we have yet again new people who came, if you want to follow on your laptop, that's a link to open the collab I'm showing. So goo.gli/cookbook-illustration. And what we are doing at the moment is illustrating the book that is called Will of the Willow, something like that,
and creating images and then now videos to illustrate what's happening in the book. So it came up with a prompt that is, like, in a colorful building block style, a water rat in his round, brown face and blue jersey, lowers his silver pistols and offers a reassuring pat to the model's shoulder.
The model in his black velvet smoking suit exhales a puff of white plastic vapors in relief, and as his pink snort switches, they turn and begin to walk together, blah, blah, blah, blah, blah. What's interesting in the prompt is that I feel like it got the style of the book.
So the prompt is kind of written in the same style, the same oldish style of speaking as well.
And we can see that it realized that
it's not attacking the model. It's just saving it and so on. And also, what we can see from there is that you remember I said some system instruction at the beginning saying that whenever it described characters, it should always try to describe what they look like and all.
So it added how they are dressed all the time, which also helps with the character consistency. So let's see if this one is better.
I think it's better, except there's no discussion.
So you can try with yourself. Just be careful, as I said earlier. Video generation can be expensive, so don't try it, like, 100 times until you get it. It can be expensive. Then we have this new Lyria model that we shipped two weeks ago that is our new music generation model.
And basically, as I said, it's just a model to which you can give a prompt, and it will come up with a song. And you have two different models. You have the clip model that is creating 30-second music.
Music Generation39:59
And, like, I think it's for $0.04 per song. And you have the full song model that is creating up to three-minute songs. And it costs, like, twice as much, so $0.08 per music. For the sake of the demonstration here, I'm using the clip model, so the fastest one.
So we are going to use exactly the same trick as before. So we are going to have Gemini create the prompts for the music generation. So I'm asking it to create instrumental songs for each chapter and to create them for Lyria, keep the consistency between the chapters, but at the same time highlighting what's specific in each chapter so that they are, like, I don't want, like, three times or four times the same song.
I want different ones for each song. And by the way, something that I forgot to say is that, as you can imagine, like, we have, like, multiple models. We have multiple teams inside of DeepMind. But actually, all of them are working together.
And a big part of the training data for the GenMedia models is being made with the help of Gemini. So basically, all of our GenMedia models are trained with, like, prompts that are written by Gemini. So that's also why Gemini is quite good at creating prompts for the GenMedia models, because, yeah, they have been trained to listen to him very well.
So that's why this kind of tricks of having Gemini write the prompts for you works quite well. And in any case, deep down, there's always a bit of rewriting of your prompts that are being done by the GenMedia models before it's actually starting to generate, just because otherwise, when people are sending, like, one-liners,
the models won't do anything interesting with a one-liner. And so usually, the longer your prompt, the more interesting it's going to be, and the more likely it's going to be following
what you're asking for. So I think I talked a lot, because it's, like, music generation is actually quite very fast. So we can see the different songs.
I think that fits with a pastoral suit
representing spring and flowing waters.
And see the next one, the open road.
Feels more adventurous, yes?
And then the dark forest, let's see.
And you can see the prompt here. Like, it comes with which instruments to use, how to use them. What's interesting with the Lyria model is that it's actually everything is managed in the prompt. You don't have that. You don't have parameters at all at the moment.
So if you want the song to be a certain duration, you can just ask in the prompt. If you want the song to be using a certain scale, you can ask it in the prompt. If you want a certain BPM, you ask in the prompt as well.
And the model is really good at understanding everything you ask in the prompt. And you can ask, like, it doesn't make sense for 30-second songs much, but, like, if you're building a longer song, you can say, during the first 30 seconds, that's what I want the songs to be.
And then it switched to something else after. Or you can say, this is the intro, this is the outro, this is the I always forget how to say that in English, but the part that is coming up multiple times in the song so that it knows exactly how to chorus.
Chorus.
It knows that this part needs to be repeated and that it needs to be the same thing and the same model. And if you want lyrics, you can either provide the lyrics
or just let it invent lyrics. So let's say we are going to change that. We are going to create songs with chapter.
Add lyrics
to describe what's happening
in the chapter.
And let's see how it goes with lyrics this time.
As I said, music generation is quite fast, like, especially the 30-second model. It takes a few seconds to generate things. The longer part here is actually sending the book again to the model and to ask it for prompts.
I think it didn't. I think it didn't work.
And worked long.
Mole flung down his brush and ran to the sun, away from the cleaning and work to be done.
And see this one.
Towed in his caravan of yellow and red,
dreamed of the dusty high roads ahead.
And, like, you can see that the theme of the songs is quite the same that's kind of the price we have to pay for using chat mode, because it remembers the prompt it did before. So we ask for new prompts, but it still has a memory of what it did before.
So I think it kind of anchored it to kind of use the same kind of prompt and describe the scene the same way. But still, it's very funny to work with the lyrics. And as you can see in the prompt, it's basically just, like, adding the lyrics in the prompt, and the model understands that these are the lyrics I need to play and add in the song.
And you can also say that this specific part
is said at exactly this moment in the song. This part
is being said at another moment. And if you check the output of the Lyria model, you actually have the old lyrics with the times. So you can create a karaoke app or something like that using what you get out of the model.
Text-to-Speech47:23
Let's try the third one.
Into the wild wood where the shadows are deep and evil faces through the hollow trees peep. Mole is in terror and lost in the snow till Ratty arrives with his pistols aglow.
Yeah, we could do music all with that.
So that's for music generation. And then we also have text generation models. And I'm going to show you something very fun. I guess you all know about the alt text-to-speech model, because everybody loves the NotebookLM integration that can create
a podcast. And that's actually great that you can select two different voices. So you can have two characters talking with each other. But I'm going to show you a trick with which you can actually create something that is basically
you can add more characters than actually two when you are creating discussions with the TTS model. So here's what I'm going to do. I'm going to ask the model to extract a specific dialogue from the book, just because I didn't want to copy-paste it.
So I told it that it starts with small, neat ears and thick silky hair and ends with his ears in the air. And I asked it to write it as a play so that it's basically a transcript of what which character should be saying.
And the trick is that I'm asking it to create a specific style of way of speaking for each character, even though it's going to be using the same voice, and to write the transcript that way. So when it's a narrator, I'm going to use one specific voice.
And when it's all of the other characters, it's going to be the same voice for all of them. So narrator is saying something, and then character is saying something, and then write the style between parentheses. And that's what is going to tell the model how to talk.
And that's something that you can also use to say, oh, it is saying this part very, like, whispering. And then this part is, like, has a lot of emotion in it. And you can play with the way the character talks.
But I'm going to use that to ask the model to come up with very different ways from the same voice to talk when each character is talking. And that actually creates the feeling of actual different voices for each of them.
And then I'm going to pass that to the TTS model and to ask it to read it, basically. One of the tricks and, like, I lost 15 minutes because of that yesterday evening. So
you cannot just give it the text to read. You always have to start with, read this text or something like that. Otherwise, for some reason, it doesn't know that it needs to read the text that is given to it.
And this is very complex. I think there's no simple way to set it up. But basically, what I say is, speaker narrator is using the voice Sulafat, and character is using Fenrir. And this is all text. So narrator talks, then the first character talks.
And that is going to have long poetic pauses. And then the second one is breathless and umbral stutter. And you can see that we can guess which character is which one, because it's the same way of speaking that is reused for each of those lines.
And it's still running. The problem with the TTS model is that it's a very good model, but it's not a very fast one. The reason for that is that.
Small, neat ears and thick, silky hair. It was the water rat. Then the two animals stood and regarded each other cautiously.
You can check.
Hello, mole.
Hello, rat.
Would you like to come over?
Oh, it's all very well to talk. The rat said nothing, but stooped and unfastened a rope and hold on it, then lightly stepped into a little boat, which the mole had not observed. The rat sculled smartly across and made fast.
Then he held up his forepaw as the mole stepped gingerly down.
Lean on that. Now then step lively.
The mole, to his surprise and rapture, found himself actually stepping onto the stern of a real boat. This has been a wonderful day. Do you know I've never, never been in a boat before in all my life?
What? Never been in a you never well, I what have you been doing then?
Is it is that so nice as all that?
We can stop that. But, like, you can see you couldn't guess that we are using the same voice for the two characters being, like, clearly we're steered into different directions. And you can use this trick to actually create, like, multiple characters, multiple voices for those characters, and make it seem less for users.
As I said earlier, this is meant to be just demonstration on how to do it. If you were to do it, like, at scale, I would not do this kind of, like, trick like that. I would actually create a full transcript with the actual names and then keep on the side a prompt for each character.
And maybe sometimes you still want to you don't want to talk to them exactly the same way, because sometimes they still need to be excited, even though they talk very slow and so on. But still, that's just to show how good the TTS model is at creating different voices.
And even though I asked it to force an accent, it didn't do it. But you can also play with, like, this character has an Irish accent, this one is English, this one speaks with a German accent or whatever.
And that's also a very easy way to create different feeling about the character using the same voice.
We're nearly at time for the questions. But, like, just to finish, like, we used a very large context window of the model to feed it a full book and to feed it multiple times a full book, because with chat, we just upload it all the time.
But since it's a multimodal in model, you can just it works with all the things and just text. So you can just feed it, like, an audio book. You can feed it a video as well. Like, you can play with that and not just get limited to text to illustrate on things.
So this is another example with another book, which is The Adventures of Chatterer's Red Squirrel. And we are basically going to do the same thing. I'm going to run all of it at the same time. And this time, I basically ask it somewhere to use a style that is futuristic, science fiction, utopia, saturated neon lights.
So it's going to be not the kind of squirrel you are expecting.
Yeah. And while it runs, I think oh, I said I was going to show you so we also have, like, I think you all know about AI Studio. But in AI Studio, we have a gallery with lots of example apps that we are building.
And I wanted to show you it's going to be in GenMedia. As I told you, we have the Lyria model that is creating music songs. But we also have the no, not this one. Let's go with the ah.
Realtime Music55:33
This one is better. We also have the Lyria Realtime model I was talking about. And basically, you are asking it to create to make music that is post-punk tunes and neo sound at the same time. But you can say, OK, I want more K-pop and slightly more drums.
I don't know what post-punk is, so I don't want it.
And see, you can hear the music changing. And let's go with something, like, more chill.
So I think, as I said, that's my favorite model, because I think it's underused. And there's plenty of things I can imagine doing. Like, as I said, I come from the video game industry. So one of the things I would have tried is, can you create music in real time for the player, depending on, like, where in which region they are?
Are they in a forest? Are they jumping? Are they cooking? Are they fighting? How much HP do they have? And so the music could change in real life, in real time. And so that's yeah, that's the kind of things you can try.
And for some reason, the link is not there. But there's another very cool example. Our colleagues who are working on this model made it's basically, you are in space, and each planet is a prompt. And you can move through the planet.
And depending on which planet you are close to, the music changes. So you can just move around the planets. And sometimes there are weird things happening, because, like, Christmas songs is just next to Viking Metal. So the mix can be quite funny.
So that's it for the presentation. I have some time for questions now.
Thank you.
Yeah, and for those who arrived too late, like, you can check the content afterwards. And as I said, we have this cookbook that is basically a GitHub repo where we are adding quick starts on how to use the models, kind of some tricks, and also examples of, like, more complex things you can build when you are mixing different capabilities and models.
Q&A58:32
I have a question. I don't know if you can answer it.
Yeah, I think for questions, we need Mike, so that's it. I don't need to repeat them.
One second.
One, two, one, two, one, two, one, two, one, two.
So thank you first of all very much for this nice demo. It was a lot of fun to follow along. I have a question. In our company, we are offering to all employees also some of the models. And I think we are still on Nano Banana One, because we can only offer models hosted in Europe.
And basically, all the new models are still in preview, so we don't have access to it. Do you know if this will change?
So the short answer is no.
OK.
But, like, I was expecting the question, because I realize it's a pain for everyone in Europe. As I said, my job is to bring the feedback from the developers and to try to make things change. So that's part of the one of the fights I'm fighting at the moment, so that we have some ways to offer a better exit for or better ways to use the model for people in Europe, because in Europe, we care about, like, data privacy and data sovereignty and all of that.
So I know it's a problem. So the core of the problem, in a way, it's the rule at Google Cloud that every preview model is only available on global endpoints. So that is unlikely to change. But what we are going to try to change is to release the model in global accessibility faster.
The problem we had with Nano Banana 2 and Pro and Gemini 3.3 as well is that we released models too quickly back to back. And so instead of, like, having Gemini 3 going GA, we released Gemini 3.1. And so we, like, kind of reset the preview counter.
So we need to do something about that. But yeah, I know. I hear you. It's kind of my P0 thing that I want to change.
Thank you.
And, like, have you run the notebook at the same time?
Yes.
Which time were you oh, wow.
Did someone else run it at the same time and with a different style and add, oh, maybe a different book?
I didn't.
No?
I did the Frankenstein one.
Oh.
And choose, like, a retro game.
Sure.
So I did the Frankenstein book and chose, like, a retro gaming style. And it was quite interesting. So
yeah, so the character looked very video game-like.
Oh, yeah.
And yeah, the main difficulty with books like Frankenstein is that sometimes the model is not going to be allowing things that are too, like, let's say, graphic that could be happening in the book. So it can be a bit toned down or worse.
It's not going to accept to show the image. Yeah. That's why I settled with kids' books for the example. That's easier, except when I was not able to make images of children, which was also another limiting factor.
More Demos1:02:21
Yeah. I can show you other I don't know what page afterwards is going to show. I can show you on other cool demos that we have related to GenMedia.
In the meantime, if you have questions, just raise your hand and we can
if we're going back here. See, that's
whatever. That's the futuristic neon style version of
Chatterer's Squirrel.
I can show you a bit more about Lyria because it's new. And you likely already know everything about Nano Banana.
So as I said earlier, when you work with Lyria, you always get two outputs. If you set the modalities to be their audio and text, you will get two outputs. And the first output is going to be the lyrics, and the second output is going to be the music.
And that's actually one of the few models where it's very interesting to use streaming. So when you do generate content here, you can use generate content underscore stream. And what it does is that you receive the first part first and then the second part afterwards.
So you get the lyrics first. So if you want to do something according to the lyrics, like creating an image or, like, giving the song a title, then you can do it while the music is still generating. And you don't have to wait for the full output to be there.
So you get the lyrics, and you get the timing. So
this sentence is going to be said at the beginning. And then after 4.8 seconds, it's going to say something else and so on. So you can hear.
The air is still and cold up here. The mountain tops are sharp and clear. And then a streak of gentle gold.
So yeah. And you can provide the same thing. And here, it's only, like, you can see in the last one, it's providing when it starts and when it ends, because it wants to have, like, 1.2 seconds without things said at the end.
You can also create music from images. So that's one of the things I forgot to do in today's demo. I should have also given the images from the chapter so that it would have been used as a reference to create the first image.
So I used this picture of, like, a grocery list for making a portfolio. And then it will come up with a song about doing a portfolio.
What was the prompt again? An epic song with opera voices about this quest.
See, it's becoming epic.
The scrolls are dry. The ink is ancient. The list is long. The soul is patient. Beef, shank, and ribs, the holy prize in the cellar where the shadow lies. The carrots and the turnips will.
And that's also nice that you can use multiple voices as well.
He's at the gate.
I'm going to skip ahead a bit.
Wait or chance. The time is. Die. The celery will. The celery will. The end is near. The final moment of the quest is here. Oh, king army movement.
That story is I like.
The
clothes, the clothes, the song.
So but then and I talked about the interactions API earlier, about being on new ways of using the API. So that's an example using those. So it basically was the same as the current API. So you give a model, you give an input, and you get response modalities.
Yeah, Philip, who worked on that, is not there, so I can say it. I would not have renamed
content to input, because it's going to confuse everybody. But that's how it is.
And then you get the output. And if we can check in the outputs, I think no. Yeah, we don't see it here. But there's and then it works basically the same way, except you get this interactions ID that you can use to chain things.
And then that's yeah, that's the one with images and prompting. What you can do as well is you can use the BPM part. So you can you want a song that is very fast or very slow. So that's that we use this.
And I even told it to use a reset accelerando illusion. So it gives the illusion that the music is getting faster and faster and faster.
So if you want music for when you do your sports routine, that's how you do it.
And as I said, you can give, like, a specific time. So the first 10 seconds are going to be fast acoustic guitar. And then it goes into piano for 20 seconds, for 10 more seconds. And then it's full band afterwards.
Actually, it's not following it.
OK, forgot what I just showed for some reason. And then but the easiest way is like this. You can use you can give the structure. So that's how I want my intro. That's how I want my verse. That's what I want my outro.
And that's a 30-second song. So it's not you won't have the chorus, but you can also add the chorus and the bridge and so on.
The darkness breaks. The shadows fall.
What is that about? Oh, yes, the sunset.
Hear the dawn's triumphant call. A golden light, a glorious sight. Chasing night with heaven's might. Our hearts resound.
Yeah. I need to make it a better, like, longer song, I think, for this example. But you can also chain all of that into everything together. So this is a full song where from the first two seconds, I have an intro.
I can tell exactly which scale to use, how intense it's going to be. And then it moves to another verse and so on and so on. And that's how you can get something very complex. So it starts very, very slow.
And then if we move,
there's the drums that started and so on. So we should be in this part. So it's still laid back. And
it's starting to be yeah, to add grooves and so on. So if you want to create complex things, it's better to use a full song model, because it's the short one is taking shortcuts to actually build something that is interesting in 30 seconds.
And as I already showed, you can provide the lyrics. So it's creating a song about Nano Banana.
Yellow peel, a tiny sweet. The Nano Banana, a tropical treat. But wait, it hums. It starts to create. Switching into AI.
But and what's funny as well is that you can use it to create things that are basically not songs. So you can ask it to create music, but without with very calm music in the background and just something reading a text or something like that.
So this and this one, I'm using the reasoning capabilities of the model to and it's knowledge about what Shakespeare is doing. And so to create a text that is basically
something that looks like Shakespearean.
And
all my glories to the grave must go.
Oh, heavy grace to be at last at peace.
And.
Grant my weary soul a sweet release.
And I didn't really tell it to read. So that's why there's still music. But you can really steer it into not having background music at all and just say things. And it works also in different languages. So you can just ask it to create songs in all languages that you want.
Sometimes there are a few words that are not pronounced theright way. It's still getting better. And this one, I tried to ask Ravik to use two different languages in the same song. And since it's the same as the TTS and live models that we have, it's really good at switching the language in the middle of the generation.
And once again, I'm using the model knowledge of things, because it's basically trying to explain how bubble sort is working in music.
And instruments or you saw it. So yeah, there's plenty of very cool things to build with the Lyria model. So give it a try. It's very easy to use because everything is in the prompt. So yeah.
Do you have any other questions? No? What do you want Paige to show you afterwards? She has seven more minutes to prepare something new in her presentation.
And also, just for clarification, I'm not sure if this was announced to all of y'all. But there are selectable issues in the building. So y'all are like the lucky valiant few that made it here earlier this morning.
Closing1:13:03
Most of the attendees were not allowed into the building. And so
we'll do the session that's coming up at 10:40. But we'll also be bringing back everybody in the afternoon who was going to be presenting in the morning to do kind of like a whistle-stop tour of all of the Google DeepMind things for the afternoon workshop.
So if you would prefer to come back for the afternoon workshop, you can. It'll just be at 1:00 p.m.
So and that was the example I was talking about. Just I don't know where Christmas songs are. But.
How can you drive us?
It just search for Space DJ. And it's available online.
Can you move the put the songs higher?
She shot this. That should be nice.
So the moor is not meant to do voices, but it can do some kind of, like, vocalizations like that.
This was not that bad at using the same voice and switching the style of music in every other song. And you have an autopilot. So it just moves around and creates music until you feel it.
Yeah. Southern rock, I think. Nashville songs.
Yeah, it's not moving fast enough. Let's move to another place.
Australian hip-hop.
Turntable, oh, turntable.
Australian hip-hop, I guess.
No, no Australian here.
But yeah, that's a very cool model. And the only thing is that the session ends after 10 minutes, so that you don't re-run it ad vitam. But I can feel it should work.
It's just missing a surf button. Speed metal.
So I will stop with that. But yeah, give it a try. It's a really cool model to play with.
OK. So yeah, I guess that's it. I will still be around if you have other questions, things you want to discuss that you didn't want to be on camera. So yeah. Thank you.





