Intro0:00
It's all good.
Can everyone hear me?
Yes.
Excellent. Awesome. Greetings, Valiant Few. Um, I'm not sure how many folks have heard, but there were some electrical issues in the rest of the building, so—so y'all were the ones who showed up early, which means that y'all are part of the few that get to hear the talks.
So if you don't feel lucky, just wait. Like, you know, there is—you are definitely—you are definitely experiencing something, something special this morning. And then, for anybody who wants to come back and hear more, you missed the Gen Media session a little bit earlier today.
We're going to be doing a whistle-stop tour of all of the presentations for DeepMind this afternoon, so you can come back and meet more of the team and kind of hear more about all of the talks and all of the technologies.
I also really, really love for sessions to be interactive, so I'm going to show you some demos. This is going to be very demo-heavy, as opposed to slide-heavy. And then, if you have any questions along the way, please feel free to shout them out.
It's always much more interesting if this is, you know, more of a conversation than just me, like, showing stuff over and over again. So, I don't think it's a secret—and also, I guess, introductions. Hi, everybody. My name is Paige.
I do—I'm one of the eng leads for developer relations at Google DeepMind. I've been doing machine learning for a really long time. I started in 2009 and was contributing to some of the early days of open-source scientific computing libraries, things like NumPy, SciPy, Scikit-learn, and then did product for a couple of years.
I'm back on the engineering ladder and really, really love that now it's really, really hard to nitpick what's product, what's engineering, what's design, and what's DevRel. And all of the roles seem to be conflated a little bit. So I don't think it's a secret that Google has been a little bit busy over the course of the last while.
Model Lineup2:18
Over the last month and a half, we've been releasing models so fast, I feel like everybody's got a little bit of whiplash. Gemini 3.1 Flash Live, which we'll take a look at in a second. Gemini 3.1 Pro and Flash-Lite, so respectively our largest and smaller model that are performant, efficient, able to do a lot of things very, very quickly and at low-cost profiles.
We actually just had AugmentCode, if anybody is familiar with AugmentCode, replat their entire agent system to default to Gemini 3.1 Pro, specifically for performance plus cost-related reasons. NanoBanana 2 for image generation and editing. Our embeddings model, which is supporting video and images and audio and text and code all in the same embedding space.
So you can say, "Show me all of the content related to cats," and it will show you not just video of cats, not just images, but also things like audio of a cat purring or meowing, or, like, books about cats, all sorts of stuff.
Lyria 3 for music generation, which you saw if you were in the Gen Media session just a little while ago. Genie 3 for world model building, so being able to dynamically generate new worlds based on user input. Our full-stack runtime for AI Studio, which includes things like databases and OAuth.
Gemma 4, which is part of our open model family. We're lucky enough to have a member of the Gemma team here at AIE this week, so definitely, Ian, raise your hand. Yep. Greetings. So the Gemma team, if you're interested in open models, would be excellent to talk to.
And then also Veo 3.1 Lite for video generation at a cost profile that's pretty compelling. So lots of different stuff. Just a show of hands, how many—how many folks have heard of all of these models before?
Excellent. And the DeepMinders in the back row, like, hopefully, hopefully, like, that was—yep. But if you haven't heard of any of these, by the end of the session you'll know all about them and hopefully know which ones you could use or consider for your projects.
So I don't think it's a secret that Gemini is kind of special in the industry. One of the reasons that it's very special is that it's multimodal both for inputs and also multimodal in terms of outputs. So it supports video, images, audio, text, code for inputs, but it can also output multiple modalities.
It can output text and code, but also audio, images. It can images and text interleaved. And most of the other models on the market are only capable of handling text and code as outputs, and only things like static images as inputs.
So it's pretty compelling to see what you're able to do. And via our APIs, you're also able to handle flexible kind of input formats. So you can have PDFs with embedded images. You can have, you know, different types of video, different types of audio that you can serve as tokens for inference.
But again, a lot cooler to see it rather than to just have me talking about it and waxing poetic. So I am going to go ahead and go to AI Studio real quick and pull up—pull up my personal instance of AI Studio, which is, you know, I always say it, but if you see anything embarrassing, please don't judge me.
AI Studio5:49
So this is—how many folks have used AI Studio before? Cool. Excellent. For folks who have never used it, you can access it at ai.dev or ai.studio or aistudio.google.com. It works just with your personal Gmail account, so you can get started for free.
You can select different models here off to theright. So you can see that there are different kind of pills here for the kinds of modalities that you might want to work with. So things like video for, like, video generation.
Veo 3.1, 3.1 Fast, and 3.1 Lite are all kind of in this tier section. You can also select the different Gemini models. So Gemini 3 Flash preview, Flash-Lite preview. I'm going to select that one just for the interest of time.
And you can also do things like toggle on and configure many of these different tools here off to theright. So you can specify things like structured outputs, code execution, which we'll take a look at in a second, function calling.
You can turn on things like grounding with Google Search, so just automatically incorporate that as a tool. Grounding with Google Maps, and also even things like URL context, which gives you kind of like poor man's retrieval. You can have a list of URLs and then incorporate that into the model's context window so it can use that to ground some of its outputs.
And as I'm sure all of y'all know, models are kind of limited based on the data that they have as part of their pre-training and post-training mixtures. So if they're only trained on data up to a specific point, that's all of the insight that they have out of the box for those kinds of data.
If you want it to be able to answer questions that happened after that date, you're going to have to give it access to tools, either through search or through retrieval, in order to do that work. And again, if anybody has questions as I'm kind of rambling along, feel free to raise your hand and shout them out.
This is a small enough group that we can—that it should be pretty fun. Cool. So I've turned on grounding with Google Search. You can also add—you can also add media. So you can connect to Drive, you can upload files, you can record audio, add camera footage, link a YouTube video, link sample media.
And YouTube just works via URL, so you can paste in a YouTube URL and have that be used for inference with the Gemini models. So as an example, I haven't tried this, so we'll see if it works. But we can take a look to see—to see if we can find a dinosaur video.
I love T-Rexes. So this—so this past weekend in the Bay Area, we had this thing called Bay Area Big Wheels, which—and it defaults to one frame per second. You can also specify different start and end times. So just for the interest of speed, I might specify start time as, like, maybe 0:00 or 0 seconds, and then maybe end time would be
maybe, like, 300 seconds.
And you can see that this ends up being around 27,600 tokens for 5 minutes of content. But I could say, "Create a table
with timestamps for all of the kinds of dinosaurs that come on in," you see in this video. No worries. It's all good. Make sure to include a fun fact about each dinosaur type, and then hit run. But what I was saying about Bay Area Big Wheels is that there's a big bendy hill in San Francisco with a whole bunch of very whiplash sort of turns, and everybody gets a little tricycle and rides down it.
And I did that this past weekend. It's an Easter Sunday tradition, but I was dressed as a dinosaur and was handing out dinosaur Easter eggs. So this is very on brand. And what's happening behind the scenes is we've turned on grounding with Google Search, so we have Search as a tool, which can help inform some of our fun facts.
We've got the video that's being pulled in, so the first 5 minutes' worth of content. We can see the different dinosaur types. Rexy and his parents obviously have a lot of appearances in this first episode, as well as a brachiosaurus, a velociraptor, and a pteranodon, which is a flying reptile.
And I love that it's—I love that it's calling out the true fact that pteranodons are pterosaurs, not dinosaurs. And you can also see the different—the different citations along the way from the URLs that are informing all of these fun facts.
You can also click "Get Code" to see all of the code that you would need in order to replicate the experiment that you just did in AI Studio. So it selects the appropriate model. It shows you how you would handle the URI for YouTube.
And then it also gives you insight into the prompt that you can use for the video in order to—in order to do the work in Python, in TypeScript, in Java, whatever your favorite language might be. And if you wanted to not use a YouTube URL, if you wanted to use your own video, you would be able to pass that to the model too.
It's just really, really handy to be able to pull in a YouTube URL, as opposed to having to do the process of downloading it and then kind of sending it off—sending it off yourself. And now I also want to watch this episode of Rexy, the little T-Rex.
This looks very cool. If you hadn't seen as well in the thinking config, you have different thinking settings for all of our Gemini 3.1 series. So minimal, low, medium, and high. If you want the model to spend more tokens thinking, you can turn on high thinking.
But I often just keep it on minimal or low just for time's sake. For Gemini 3.1 Flash Lite, you get a really nice price versus sort of
price—price, performance, and also speed profile for the models. So you're not having to make big, big trade-offs between them. So that is how you would interact with Gemini 3.1 Flash Lite within AI Studio for video analysis. One of the other slept-upon features, I think, in our APIs as well as in AI Studio in general is compare mode and also code execution, which we see here off to theright.
So if I turn on code execution, what we do is we give Gemini a sandboxed environment with Python and a whole bunch of data science libraries pre-installed, where it can kind of pull in those libraries as tools to kind of help solve arbitrary data science tasks.
And since this is sort of giving the model access to it in a sandboxed environment, you don't run the risk of having any of this impact your local environment, which is quite nice.
So as an example, if I select Gemini 3.1 Flash Lite preview, turn on code execution, go into compare mode, I might try to compare it against Gemini 3 Flash preview, also with code execution. And then one of the things that you can do, and we'll see if this works, is you can select a picture.
So this is just some Lego bricks. What I could do is paste this in. We can make sure that it's secure and safe for the corporate overlords. This image itself is around 1,000 tokens, but I could say something like, "Draw bounding boxes around all of the green Lego bricks using Python," and then maybe display the image with bounding boxes and hit run.
And what should happen is that we see a head-to-head comparison of the two different models. Gemini 3.1 Flash Lite was able to get itright out of the gate, which is pretty wild. So this super, super tiny model worked really, really fast, wrote the Python code to pull in the image, to analyze it, and to define the bounding boxes.
And then if you hover over the token consumption, the amount of dollars required to do this work is pretty wild,right? Like, so being able to—being able to pull in an image, do this kind of analysis, you could have also asked for things like segmentation masks.
You could have asked to count specific kinds of entities in the photo, again, using bounding boxes or something similar. And all of this was done at well under a fraction of a penny. So strongly, strongly recommend experimenting with the smaller weight models, especially turning on these tools to help them do their work more effectively.
And the—like, you can also see that Gemini 3 Flash preview got to the—got to the same answer. It just took a little while longer. And then the cost is slightly more, but still well under a penny.
Cool. So this—that's compare mode, again, using Gemini 3.1 Flash Lite, just with the addition of code execution along the way. For folks who might be interested in URL context, just because I know that this is—this is something that we've heard quite a bit about from folks that are using the Gemini APIs pretty regularly.
If you turn on URL context, you can do things like add URLs. So I'm going to pull in a URL for a blog post about Gemma 4, which was released just recently, last week, after the model's training data cut off.
I'm going to pull in, you know, a blog post about Genie 3, also after the model's training data cut off. And I could say something to the effect of, "Compare and contrast
Genie 3 and Gemma 4. Tell me how they're similar, different, or completely unrelated." And they're mostly completely unrelated, but we'll see what the model thinks. Hit maybe turn on medium for the thinking level, and then hit run. And what we should also see is that the model is able to give its output, but cite each one of the sources that it's using in order to make its assumptions.
So you can see the different sources down at the bottom, the two URLs that I had used. You can use, you know, many, many more than just two, but it cites each one of the sources in line as it's making assessments along the way.
And so you can use publicly available information, and then there are also tools within Vertex that allow you to do retrieval on custom documents that are internal only, without necessarily having to set up a vector database for retrieval.
And again, if you click "Get Code," it gives you all of the code that you would need to replicate what you're doing in the AI Studio interface.
Cool. So we've talked about—we've talked about the Gemini 3.1 series of models. You can also use—you can also use 3.1 Pro preview, which is our largest model series. It's also our most expensive. Our kind of, by virtue of being larger, it's also a little bit slower.
Gemini 3 Flash is kind of the workhorse. It's the model that we see most commonly used in production. And then Gemini 3.1 Flash Lite is also gaining traction, especially for people who are using Gemini 2.0 Flash or 2.5 Flash.
And then for our Gemini Live series, how many people have seen or heard of Gemini Live before?
Gemini Live19:12
Cool. Couple of hands? Like, thank you. Thank you, DeepMinders, for, like, you know, paying—like, paying attention and also for using the models. Gemini Live gives you the ability to share your screen, share a video feed, share audio with the model, and have a conversation with it dynamically, including things like adding custom function calls, generating automatic function responses, and then also doing things like grounding with Google Search.
So I can turn on grounding with Google Search again, just as a tool that the model can call. And then I can give the model access to my screen to help me either navigate an app, to help me answer questions about what it sees, and to do that in a broad variety of languages, accents, etc.
So as an example, I can start screen sharing.
Hey there, Gemini. Can you tell me what you see on the screen?
I see a Google search for Lego bricks and pieces. There are tons of images of colorful Lego bricks, sets, and bulk collections from various websites. On theright, there's a larger image featuring close-ups of red, yellow, blue, and green bricks.
Are you looking for something specific on this page?
And do any of y'all—does anybody in the room speak a language that is not English? Yeah. Guillaume, like, you would not—like, you would not be a reliable narrator. Like, but at least, like, last night we tried this and one of our colleagues yelled out, "Yes!"
very confidently, a language that they do not speak. And so we did not have anybody to verify. But I saw a couple of hands on the front row. What languages do y'all speak?
Italian.
Italian? Okay. So I'm going to—are you okay with verifying if the model is, like, saying true things in Italian? Okay. Cool, cool. Thank you for being our fact-checker. Hey, Gemini, could you tell me what you—could you tell me what you said again?
Just tell me in Italian. And then also, could you tell me what the weather is like in London today, also in Italian?
Certo. Per quanto riguarda i risultati della tua ricerca, vedo tantissime immagini di mattoncini Lego colorati, set e collezioni sfuse. C'è anche un'immagine più grande con mattoncini rossi, gialli, blu e verdi. Per quanto riguarda il meteo, a Londra oggi è previsto un tempo nuvoloso, con possibilità di pioggia.
Ti serve qualche altra informazione?
It was good? Excellent. Very cool. And so you can ask for different languages. You can also say, "Hey, Gemini, can you tell me a poem about what you see on the screen, but could you tell me it using a text and accent?"
Well, howdy there. Let me take a gander at this here screen. I see piles of Lego bricks, a colorful commotion, like a tiny town fixing to rise from a plastic ocean. Red and yellow, blue and green, in batches big and small, ready to build a castle tall, or maybe just a wall.
Folks want to buy 'em, sell 'em, trade 'em from all over the place, to put a smile, I reckon, on a little builder's face. Anything else you need help with, partner?
That's giving me flashbacks to my hometown. But you can change—you can change all of this dynamically, just by asking naturally within the flow of conversation. So you could imagine practically a scenario like, perhaps you have an entryway in a bank, and there's some sort of a screen, somebody comes in, starts speaking in Spanish, or starts speaking in their, you know, the language that they feel most confident in.
And the model's able to dynamically respond and answer their questions in a language that's familiar to them. Or you could kind of specify within system instructions what the—what language, dialect, accent, style you might want the model to adopt.
So if you only want the model to respond in a specific language or a specific style or with a specific tone, strongly, strongly recommend modifying the system instructions. And same as always, if I click "Get Code," you see all of the code that you would need to use to replicate the experiment that you did within the UI.
So you can see the media resolution settings, the settings for compression, and all of that kind of incorporated in naturally. I can also do things like share video feeds. So.
Hey there, Gemini. How many fingers am I holding up?
You're holding up two fingers.
What about now?
That's a thumbs up.
Yep. Cool. And so big, big kind of spectrum of things that you can accomplish with Gemini Live. And again, just a very, very low price point compared to other solutions that make you kind of stitch together the speech-to-text, LLM understanding, and text-to-speech pipeline with all of the video content inputs and outputs all by yourself.
We have another feature. Like, I always feel like whenever I'm describing AI Studio, I'm just like, "And also, and also, and also," you can do all these other things. We have another feature called Build, which if you've played with v0.dev or Lovable, feels very similar.
App Builder24:52
It gives you the option to kind of create and deploy and to share a whole spectrum of apps. And now we have even added support for things like databases and authentication. So you can add a database, you can add login with Google, you can add custom API keys that are all kind of kept secure for you.
And you can also, of course, create and edit existing apps. So you saw a little while ago some examples using music from Lyria 3, which is exciting. Guillaume, who created the Lyria Studio app, is in the back today and is, like, these are all really, really fascinating to play with if you haven't had a chance to experiment with some of the generative media models.
You can see some of the examples with NanoBanana 2 as well, and also with MediaPipe. So as an example, if I click on this app, you can see that it's requesting camera access. This is a game that's taking in kind of the location of my hand.
So I can grab
and kind of, we can all find out that I play this game really badly. But you can sort of play the game and then also inspect all of the code that's used to create the app itself. But for the purposes of this, I'm going to just show you how you can get started with creating an app from scratch, just based on anything that you could possibly imagine.
And I'm going to use database and authentication
so we can sort of add Firestore and authorize it, sort of the Google login with Firebase. And I will click this little speech-to-text microphone that we have here so it's easier than me typing out all of the details.
But as an example, create an app that allows me to upload
a picture of a bookshelf. The bookshelf should have a lot of books and kind of profile view so we can see all of the spines and maybe some information about the titles of the books, the authors' names, etc.
But the app should use Google Search grounding to add more information. So what we should get is, like, the author name, the title name, a description of the book, kind of what the category of the book might be.
And it should, the app should ask the user to log in with their Google login. It should save all of that information for the user to a database. And, you know, we should be able to have that persist.
So it's basically like you take a picture of your bookshelf and it automatically catalogs all of your books for you.
Which is a lot. Like, that in theory would have been a startup, you know, three, four years ago. But this looks reasonably correct. So I'm going to go ahead and click Build. And what's happening behind the scenes is you can see Gemini 3.1 Pro preview kicks in.
It starts thinking and planning about what would be needed in order to create this app. Since it's doing a lot, standing up a database, like thinking about authentication, it's going to take a while. And while it is, we're going to be kind of
going to show another couple of demos so we can let the model cook in the background. And then if it needs to, if it needs me to take any actions, there will also be, like, a little ping so we can hear it in the background
Genie 328:58
just in case, just in case along the way. So I am going to minimize this a little bit. And I'm going to pull up my other browser window. And we're going to take a look at project Genie. So how many people have heard of Genie before?
Yep. Excellent. So all of the hands in the back row, thank you. And then also a few folks here in the audience as well. Genie 3 is DeepMind's model for generating new worlds. So you can describe kind of a scene, describe a character, and then actively experience it with each frame generated dynamically.
No physics engine behind the scenes, no Unity, no Unreal Engine, just each frame generated dynamically pixel by pixel. You can navigate it using the arrow keys off to the left, so the WASD keys within the Genie app. And you can also change the video perspective using the arrow keys, do things like click the space bar.
But it's everything from this, like, volcanic landscape where you're navigating with kind of a rover, to things like navigating a watery landscape on a jet ski. And you can see that if you hit one of these lights, it actually responds as if there was some sort of a physics engine based on its training data and other information that it's seen along the way.
It also sounds like AI Studio might have done something. So we'll take a look at that in just a second too. And then you can also see things like hurricanes and what it would be like to experience a hurricane in Florida,
jellyfish, you know, and these thermal underwater situations. Just really wild and very magical sorts of experiences, anything you could, anything you could create. So let's take a look at what AI Studio is asking me for. So it wants me to enable the Firebase database.
And it looks like it's setting that up. So that seems good. Like, it's on track. And I'm going to head over to project Genie. And we're going to explore and create a world. If I could sign in.
We're very big on security.
For good reason. Yeah. And then so we have the, we have the option to create an environment, to create a character. And since I am feeling homesick after hearing that, like, Texas Twang
about the poem for Lego bricks, I'm going to say
Big Bend National Park in Texas in the middle of the summer, sun shining in the sky, but all of the rock formations are made out of Lego bricks. And
the sky has a rainbow, a quadruple rainbow. Why not? And that I can guarantee you is not, like, a situation that exists in actual Texas.
Ground is sandy and dusty. And then maybe the character is,
what would be a good idea for a character? Ostrich with a rocket blaster and goggles. Maybe make it pink. So pink ostrich. Cool. I don't think Texas has ever had that. So we'll see what gets created. Behind the scenes, Genie 3 is actually a composition of models.
So it's not just one model. It's NanoBanana, Veo, Gemini to help with prompting, all kind of stitched together, along with some really, really interesting approaches towards distributed systems and compute. Oh my gosh, this is amazing. Like, I immediately want a YouTube video about this guy.
Also, we see some Lego brick rock formations. So let's create this world. And then what we should be able to do is navigate through it again using the arrow keys, the arrow keys to change the visualization and the views and the WASD keys to navigate the little dude around the world.
So we've got the, like, a couple little options for the ostriches, each one moving. So you can see that it also seems to have given him, like, very, very muscular arms. Like, maybe it wants the ostrich
to kind of be like a military-grade fighter. But you can see it walking around, navigating the Lego bricks. And then if I turn around, let me see if I can find my way out of this rock formation. You can even have it investigate some of the scenes.
So we've got the rainbows. If I'm remembering correctly, if you walk towards this canyon, there should be a river at the bottom. So we can try to make him jump into the canyon. But all of this is kind of captured, again, just dynamically
by the Genie 3, by the Genie 3 model harness itself. So come on. Oh no, I didn't make it in time. But it's really interesting to see some of the things that you can build. One of our colleagues, Fofr on Twitter, so F-O-F-R, created a game where you're a fish and you have to escape a kitchen.
And you're just, like, bouncing along as a fish, you know, trying to get out before it's dinner time. So it's really, really cool to be able to see some of these things in action. Genie 3 is not currently available as an API just yet.
But the team is, you know, actively thinking about a trusted tester program. And today you can access Genie 3 through an Ultra subscription, though the Ultra subscription is only available with Genie in a select number of countries. So I strongly recommend taking a look at that.
Yep. Question.
From my understanding, this is like a, it's very cool, but this is like a video model essentially that does render to the frames basically. It's not like a 3D mesh.
No.
So you cannot ingest it into, like, Unity game engines.
No, you would not be able to create the 3D game meshes or pull them, pull, like, this ostrich dude in as an asset for a game. It is just the pixels. So we have seen people couple together things like the images that are generated with NanoBanana and kind of use additional techniques to turn them into 3D assets.
But that does require additional work. This isn't automatically creating the 3D assets for the games themselves. But it's a really, really good question. There are some other companies, there are some other companies that are taking different approaches for world model building.
So Fei-Fei Li's company as an example at World Labs is taking a different approach towards building out these environments that do incorporate more of kind of like the Unity Unreal Engine style asset generation. But I think it's longer term, as all of the models seem to converge on many input modalities, many output modalities.
We'll probably see all of that kind of converge as well. And so I wouldn't be surprised if in the future there would be an opportunity to have, like, video as an ingested thing for a model and then 3D world or, like, the code for it produced externally.
Yeah. Especially given that with Gemini today, you can already kind of give it an image and then say, please create, like, an SVG of this image. And it can do it pretty well, which actually might be a fun demo.
So, like, but one I've never tried before. So, like, let's see if it actually works. And hopefully it's not just me pretending that it does. So, but what you can do is if we, making sure I'm still sharing my screen.
Media Generation38:30
Cool.
You're kind of doing writing advice of the.
Well, so, but that benchmark has gotten saturated,right? Like, so I'll take the Lego bricks, the Lego bricks photo that we had just used and say something like, create an SVG of this image. SVG representation of this image, which is a very, very simplistic prompt.
I could probably get a lot better, I could probably get a lot better results by asking Gemini to expand upon this prompt as opposed to me just kind of, like, spitballing a really, really simple one. So if we don't get great results, we'll ask Gemini to rewrite our prompt to improve it.
And so we can see the thinking kick in. One of my most favorite hackathon projects ever, they created, they used NanoBanana actually to take an input image and then to show step by step how you would be, how you would draw it with the different stroke marks along the way.
But we can see that it's thinking through the perspective. It's defining the bricks. It's thinking about the dimensions of the bricks. It's calculating a grid, defining some colors. Since I turned on the thinking level to be high for the Gemini 3.1 Pro model, it's doing an awful, awful lot of thinking about simulating the rotations and the transformations.
We can see that happen along the way. It also sounds like AI Studio has an update for the bookshelf cataloger. So let's take a look at that while the SVG is generated.
Firebase terms accepted. Let's retry to see what it's doing. And it looks like it was able to create some code for the TypeScript, the CSS, et cetera. I wonder if because I started using Gemini 3.1 Pro in a different tab, maybe it got a little bit tired.
But we'll see. I also really, really love that you can experiment with the generative media models in AI Studio. So if you were here for the earlier session, you saw Guillaume share a lot about Lyria, about our NanoBanana models, about Veo 3.1 Lite.
And so as an example, with NanoBanana 2, you also have the option to do things like image search grounding. So you can turn on image search and it will kind of reverse the image search and bring back things that are tightly aligned with what you're asking for.
So as an example, I could add sample media for this, like, this cute little dog. Maybe sample media for, let's see, sample media for, hey there, greetings.
Sorry.
Welcome. Like, the sample media for maybe this outside, this very, very nature-friendly location. And then say something like, show me the dog in the middle of the natural park with a can of Celsius. Which if you have never had Celsius, like, bless your heart, like, that seems like a great life.
Celsius is, like, a notoriously disgusting, or at least from my perspective, it was pretty disgusting, but very popular at hackathons caffeinated beverages that tastes a little bit like battery acid. At least to me. Like, I'm sure it tastes delicious to many other folks.
It's also very low calorie. So it's a little bit like a Red Bull alternative. But I've given it a picture of a dog, the picture of this natural scene. I've turned on reverse image search. So it should be able to pull in details about what a Celsius can might look like.
And it's thinking through the assignment. It's got my little dog in the natural scene with a can of Celsius. And you can also, if you hover over the token consumption, see that in comparison to the NanoBanana, the kind of pro model or pro tier model, it's much, much more cost effective than previous iterations.
So if you're interested in using the NanoBanana series, NanoBanana 2 is a good one to get started. And just as always, if you click get code, it gives you the code that you would need to replicate whatever you just did in the AI Studio UI, just using TypeScript or Python or whatever it might be.
You just switched off thinking to be fast.
This is true. Like, if you want the, just as always, if you change the thinking settings to be minimal or low, the model will give you a response much more quickly. Whereas if you ask it to think, it will spend a lot of time generating tokens for planning and for reasoning about the task that you've described.
Cool. And so let's go back to this SVG representation. It looks like we've got a first pass. So I'm going to copy. I'm going to go to an SVG visualizer, just an online one, and then paste in that.
And it looks like we've got our Lego bricks. They're a little bit distorted, but they look pretty reasonable, honestly. And then the, as a reminder, the picture that we were trying to replicate is this one. And it was able to get all of the different kinds of Lego bricks, just not in theright configuration setting.
So it's really, really cool to see that you can kind of pull in an image. And then with this was a very, very simple prompt, but with a much more detailed prompt, you would probably be able to get a much better representation.
I'm also curious, like, if I turn on code execution, like, I wonder if it would be able to have, I wonder if it would be able to invoke code execution as a tool call in order to do that more effectively.
And so we'll see that in a second.
So it does look like it was able to, it does look like it was able to pull in an appropriate library to think through the, to think through the process of generating SVGs
or an SVG for the image.
And it's even doing the segmentation. This is very cool.
And for folks who came in a little bit later, code execution is a tool automatically invocable via the API that gives Gemini the option to kind of create a sandboxed Python environment with a whole bunch of data science libraries pre-installed.
And it can invoke those as kind of subtools within the environment.
Wow.
Awesome. Very, very cool. So we're also still building out the bookshelf visualizer. It looks like it's creating the Firebase blueprint as well as some of the rules. And so if we go back to code, we can see all of this getting generated along the way.
Another thing that I strongly, strongly recommend folks take a look at if you have interest is our video generation series. So we have a new model called Veo 3.1 Lite that also gives you the option to create really, really nice stock footage backed with audio,
as well as basically anything that you would be using the larger tier series of Veo to do, just with the model itself. So as an example, let's go to, let's go to Gemini and ask it to help us generate a prompt.
I'm just going to turn on thinking to be low and say something like, create a prompt for a video generation model to generate stock footage for a vegan basketball-themed food truck. Make sure that the food options are warriors-themed,
which is a San Francisco, which is a San Francisco team. And then hit run.
And then what we're going to do is we're going to take this output prompt and then put it in, put it in Veo 3.1 Lite.
Hit run. You can see that the output resolution is set to 720p. You have a couple of different options for output resolution, not 4K, which is something that you would need to use kind of a higher tier video generation model for.
You can also specify different aspect ratios, so 16 by 9 or 9 by 16 if you want more of a mobile app experience. And you can also sort of configure the video duration. So if you want eight seconds versus if you want, you know, something a little bit more concise, like four or six seconds, you can pull that in.
As well as this is a paid tier model, so you would have to attach an API key in order to use it. The handy thing, or another handy thing about AI Studio, is that if you expand the settings off to the left, you can see there's a section called Get API Key.
And if you click Get API Key, you can create, you can create one that's acceptable for free tier use, just out of the box without having to, oh my gosh, Chef Curry. This is amazing. Chef Curry!
Splash Brothers.
And that does look like tofu, like tofu barbacoa with kale and with avocado and with edamame. Like, I would, oh my gosh, I love this. No kidding. Like, and I am absolutely going to send this to somebody I know because their dream is to start like a vegan basketball food truck.
As well as like a custom vegan nut butter business, which I think would be a really, like, apparently nut butters have like a 50 to 60% margin. So if any of us need like a hobby plan, like maybe cultivating some of these culinary hobbies is a good one to take.
Another thing that I want to make sure to mention, we talked about it a little bit, and we have Ian from the Gemma team also available in the back. He'll be coming back later this afternoon to discuss as well.
But we just recently released our Gemma 4 series of models, which are extremely, extremely powerful. So they're able to punch far above their weight in terms of the parameter size and the compute footprint associated. But you can use them via the APIs in AI Studio as well for free.
So if you want to be able to test out the Gemma series of models, you can have this kind of try before you buy experience within AI Studio before downloading them to your own infrastructure. If we don't necessarily have a spare GPU at home, hidden out in your closet, you can just kind of, you can ping it through the AI Studio interface as well.
And if you click, I'm going to do another prompt and then pull in just, pull in just an example image. The Gemma 4 models also support multimodal understanding, so they can analyze audio or video or images. You can say something like, generate a brief description of this image.
Turn thinking level to minimal.
And then the Gemma models are pretty fast as well. So if you need a lighter weight model accessible via an API that you can work with for free, or if you need a model that you can download, use on your own infrastructure, fine-tune and run for free with an Apache 2 license, the Gemma 4 models are an incredible option for you to try.
They also run on mobile devices for the smallest versions. So you can have one locally downloaded to your Pixel. The next series of Pixels, like Pixel 10, should have Gemma already added to it. And then Chrome as a browser is also incorporating the Gemma models.
Lyria52:55
Cool. So we've seen the vegan warrior food truck. We've seen Genie 3. We've seen our open model family, some Lego bricks and pieces. It looks like the AI Studio app is still cooking a little bit. And one of the other things, one of the other things that was mentioned was, one of the other things that was mentioned was the Lyria model, which is also available via AI Studio.
So if we go to audio, you can see a couple of different models that are available to try via API. So Lyria 3 Pro preview, Lyria 3 Clip preview. So as an example, if I click on this guy, you can see the kind of some of the automatic templates that you can use with it.
So Acoustic Folk, 90s All Rock, et cetera. But I really, really love this app that Guillaume built, which you can find in the gallery and can also remix to your heart's content. And it incorporates different sound configurations. So if we preview this guy, you can see an option to create your own sound.
So a clip, maybe electronic, danceable,
vegan food truck, vegan basketball food truck, and Legos. And then we talked about Italian. What language do you speak, sir, in the front row?
Oh.
Yep.
Spanish.
Spanish? Excellent. Spanish.
Lyrics in Spanish. And then create.
And we should see the clip start synthesizing. It looks, that does look pretty Spanish. And we'll see what it means for electronic and, oh my gosh, that's amazing. That is so cool.
En el food truck vegano, la fiesta empezó. Con mi equipo de básquet,¡qué buen sazón! Piezas de Lego de todos los colores, construyendo un mundo de sabores. Baila con el ritmo de la capital. Goza una energía monumental. Es la fiesta vegana.
Es la fiesta vegana.¡Vamos!
This is, you know, like, so, well, so we've got a video for it. We've got a theme song for it. Like, clearly this is something that we should all be, like, like our post-ASI plan is now, like, we're going to start a vegan food truck that's basketball-themed and Legos.
But this is our Lyria 3 model. We had a session, a great session about generative media just before this, led by Guillaume. So if you missed it, it should be recorded and you can watch it, you can watch it afterwards.
And then we'll also be talking a little bit about it in the workshop later this afternoon. But again, all of the code is kind of available for the app, so you can experiment with it and test it out.
Bookshelf App56:28
And then we'll also take a look at the, oh, cool. So it looks like our bookshelf cataloger is done. I'm going to go ahead and sign in with Google. So it should ask me to log in with my personal Gmail account.
We're going to continue.
So it signed in as me, which is great. We're going to upload a photo. So I'm going to find a bookshelf with books on it,
like a smaller one to make it a little bit easier.
Let's go with, I've tried.
So we'll see what this one looks like. Yep. So this one has some, this one has some, like, handwritten style text that I want to see if the model will be able to pick up on. And also you can't really see some of the author names.
So I want to see if it'll be able to sort of figure out what the book title is, even though I can't see everything on the spines. I'm going to upload this photo that we just downloaded.
And it shows the latest upload.
It's figuring out the book details and it's adding all of them. So it's figured out the different kinds of books, the name of the authors, the descriptions of the books. And then if I sign out and sign back in again,
it should be able to also persist. Yep. So it persists all of the books that I had on my shelf. And then if I wanted to share it with all of y'all
and copy the link, make public. So public, anybody can access. If anybody wanted to, like, QR code generator.
Yeah. If anybody wanted to try out this bookshelf app themselves, you can access it by trying out the QR code there and going to it, which is pretty wild,right? Like, it's also a one-button-click deploy to cloud run. Though, like, in the interest of not burning up my quota too awful much, like, that is, I will refrain from doing it for this app in particular.
Robotics & AR59:08
But those are most of the things that I wanted to show. So let me go back to the slides again. I hate slides. Like, I'm pretty allergic to them. We'll see how this works.
So we've talked about Lyria. Another thing that you can use the Gemini Live model, so that real-time interaction model that we were just, that we were just playing around with, is in robotics. So this is a robot called Pupper.
It is completely, like, open sourced. You can 3D print all of the parts. It's running Raspberry Pi. All of the software is open sourced, but it's using the Gemini models behind the scenes for things like object detection and to be able to respond to its environment.
You can also run Gemini Live with the Pupper. You can use it to kind of flexibly tell the robot what to do. And the way to orchestrate this isn't having Gemini Live control the robotic actions. You would have it kind of build the plan and then invoke models that might be local on the robot in order to do things like pick up specific items.
But you can use Gemini to build the plan to accomplish those tasks. And then also things like augmented reality. Gemini Live is great at giving directions, at kind of responding to things that it sees, at describing, you know, how to do math that might be on a whiteboard.
And even, you know, enabling things like real-time transcription of, you know, if somebody's speaking to you in Chinese, being able to transcribe just in English what the person is saying. So lots of really, really cool things are capable with these multimodal systems.
And with that, it feels like a good place to stop, to ask for questions, and to also, I know I'm the only thing standing in between all of us and lunch. Like, hopefully, hopefully get us all to the cafeteria or the session with the food a little bit early.
Q&A1:01:08
Does anybody have any questions? Did anybody learn anything new?
Cool, cool. Yeah?
Are there any plans to use some sort of data set for the Google AI as well?
So, so yes. They're not so much a codex, but there is a plan to have an AI Studio app, which Logan has alluded to at least a few times on Twitter. So stay tuned. Stay tuned. It should be interesting to see.
And the team's very excited.
Yeah. Excellent. And you can also use the Gemini APIs with all of the things that you know and love, like OpenClaw. We have a colleague, Ali, who is very emotionally invested in his Telegram plus Gemini setup and uses it all the time to invoke, like, workspace actions and to, and coupled with Google search.
So definitely, especially given the free tier for the Gemma models and for some of our Gemini models, Gemini plus OpenClaw is a good path forward.
Cool. Excellent. Well, thank y'all all for coming. Thank y'all for being early as well. And then I hope to see you tomorrow and later this afternoon.





