Intro0:00
Uh, while people get settled in, I'll do a quick survey, I guess. Yeah, feel free—go, take my seat. Um, raise a hand. How many people use Google Photos in any way? Are you familiar with the application? Cool. How many people have used the editing feature from Google Photos?
Cool. If you're less, that's fine. How many people have used, like, Dolly or any of these new generative image editors? Okay, not that many. Cool. Okay, well, my name is Kelvin. Uh, I am an engineer on the Google Photos editing team.
Really quickly about Google Photos, I think most people are familiar with it. We are the home for your memories. We offer auto backup. When we first launched, the product really was built with machine learning in mind,right? We want to use ML to make your life easier.
Um, so if, for example, you—if you back up your photos with us, we index your photo, we run OCR, so you can do really powerful machine learning search. So you don't have to organize your photos anymore. You go on a trip, you come back home, the photos get backed up.
We go, "Hey, you went on a trip, here's an album of your photos from that trip." Um, you can just search for stuff like "receipts from this restaurant" and we'll find it for you. We're all taking way too many photos nowadays, way too many videos.
Nobody's got time to sit down and organize any of it. So let us do that for you. We have about 1.5 billion monthly active users, and we do hundreds of millions of edits per month across our clients. The team I work on is called the Compu-Computational Photography Team.
Comp. Photo1:33
I just call it the editing team, it's easier to say. Uh, it was started in 2018. We had a pretty basic image editor at the time, but we wanted to really focus on this idea of using the compute on your phone,right?
It doesn't have a great sensor, but it does have a lot of compute,right? And we want to use that to make really great image edits for any images. It doesn't have to be captured on the phone, it could be old devices too.
And the idea is, with computational photography—so for example, you have, like, this photo on the left that is kind of taken at sunset, and the subject is very dark. Traditionally, with DSLRs, what you would do is you would capture multiple images at different exposure levels, and then you combine them into theright image, which is called the HDR image.
But that's a lot of work. You need a tripod, you need to know how to do this, you need to know how to use Photoshop. With machine learning and compute, we can just do it with one image. So capture it anywhere you want, bring it to Google Photos, we run machine learning, we generate this photo, we kind of re-bring in all the brightness and variance in the image.
And the reason we're doing this at Google with the Photos team is we are able to vertically integrate. We control the hardware from Pixel,right? We're able to use stuff like Edge TPU to really do accelerated compute. But we also have internal research.
Um, this is 2018. Back then, I think Hugging Face was founded, but it was not a place where you can just go on Hugging Face, find a model that suits your need, and then go off and build an application with it.
Like, you can do that now, which is really great, but you couldn't then. And we're able to work with our researchers in Google that are experts in computer vision and machine learning to build the feature from the ground up,right?
Um, like, "Hey, can you iterate on the model? We'll iterate on the application." Let's keep going back and forth until we hit something really good that's easy to use. I'll really briefly talk about our tech stack. We have three main clients: Android, iOS, and web.
Those are all pretty natural to each of those clients. My team also owns this shared C++ library that does all the model inference, and it's all on-device. Um, we integrate across all the clients, we integrate with our research partners, um, and then we run inference on-device using, um, TensorFlow Lite, which is now called LiteRT.
I'll briefly show off some of the editing features highlights that we built from 2018 for the next couple of years. The first one is this post-capture segmentation,right? You capture a photo of a portrait. Oh, you decided I actually want more bokeh in the background.
Segmentation3:41
I didn't have a lens that could do bokeh in that environment. We're able to just do it for you after the fact. Another one, you know, you capture a nice portrait, the lighting's not perfect, maybe the sun's in your eyes, it's kind of, like, making you washed out, or it's, like, overcast and you have shadows on your face.
Also able to fix that after the fact. You know, you want lighting on your left side, lighting on yourright side, no problem, we'll do it for you. Another one that's really popular on one of the Pixel launches is Magic Eraser,right?
You may be at a popular spot, there are a lot of tourists or distractors in the background, you want a really clean photo of just the core subjects. Um, no problem, don't worry, you don't have to wait for the scene to clean up or anything like that, just take the photo.
We can auto-detect, like, what are the subjects you care about, and it just cleans the background for you and inpaints it for you and gets something really beautiful that you can print off or show off to your, you know, close friends or family members.
Diving briefly into, like, you know, the first big feature we launched, um, which is the post-capture segmentation. This is really simple, actually, relatively. It's a UNET convolutional neural network. This is, like, pretty traditional ML stuff in the computer vision space,right?
And it's able to just focus on a specific use case where it works really well, which is single portrait, single subject portrait segmentation. So you see below, we separate that into a foreground, which is the white, and the background.
And this is where the strengths of ML really come in. Like, you can do this with computer vision without using machine learning,right? But it doesn't behave as well, and you need experts who are experts in that system to tune it.
With machine learning, you just get better performance, you just build a benchmark of the dataset and you do the training. And it always returns a result. There's no error handling. Like, the beauty of models is it always works,right?
You give it any input, it'll run, and it's like, "Here's the output." And that's nice. There are some challenges. Models are really big. They're giant, um, static files of floats,right? In this case, this model was 10 megabytes, which is a lot of code.
And that's not something we can bundle in the application. We're very sensitive to our APK size. So now we have to download the model after the fact. Then you have to do model management, IP protection, since we're doing on this on-the-user's device, we have to make sure the model can't be extracted and used somewhere else.
Uh, also, I'm sure everyone here who's worked with AI and ML has to deal with evals. Like, how do you know this model does what you want? How do you know the next model is better and does better than what you want?
And that just takes a lot of time. You have to build a benchmark, you have to run the benchmark. Um, you have to make sure the benchmark actually reflects your real-world usage, because if they separate, then the benchmark is useless,right?
I think of it as, for traditional code, you have unit tests, and that's how you make sure you don't have regressions. The benchmark is the equivalent of unit testing for your model. And you need to maintain it, and it takes time.
And then the pro is also a con. The model always returns a result. Like, it'll—people deal with LLMs, so it's like, "Hey, tell me when you're not sure, so I can do something else." And the LLM will be like, "What do you mean?
I'm always sure." Right? Even in this case, it'll always return you something. And for the example, even this segmentation,right? The subject has really long hair, and you'll see the mask it produces is not perfect. It's not capturing her hair, the finer strands.
So when you really do blur, or you do some really sharp segmentation, you will notice that it's not perfect, there's blurry edges, and that sort of thing. And we can fix that post-model. We'll be able to run kind of more image understanding traditionally and be like, "Hey, that's a hair, let's follow the strands of the hair," get it really perfect after that.
Magic Eraser7:23
Uh, a couple years later, we launched Magic Eraser. This is in 2021. At this point, we're really doing a system of models. You know, now people call it, like, orchestration in the LLM world. But really, we have a few things here.
We are detecting distractors, we are segmenting them, we are then running an inpainter on them, and then we have, um, custom GL rendering on your device to make all this seamless. So visualize the mask, animate the mask away, bring in the inpainting area, so on and so forth.
More models, more things to con. Your system is now more complicated, you have larger models, they're getting to hundreds of megabytes now, even on-device. Um, the failure cases are more obvious. If you ask to inpaint something really large from the scene, very visible in the foreground, um, even a model will struggle with it.
Um, at least in 2021, as Paige mentioned, now we are able to do much better stuff. So, quick summary of our learnings from those couple years,right? ML really does give you great capabilities that you wouldn't be able to do with traditional image understanding.
It's great that you can shape these features that are very easy to use. The models themselves have very consistent latency on a specific device,right? Once again, the model just does the same thing, it doesn't matter the input. We're going to run this number of flops, it's going to go great.
The devices themselves, on Android especially, have a huge variance. You know, the latest Samsung and Pixel, really kind of comparable to certain laptops. The really older phones, not at all comparable. So you have to deal with that. And then, once again, going back, like, we really had this great relationship with our researchers, some of which is now in DeepMind,right?
Being able to talk to them early. It's like, "Hey, we have a use case. Our users want to be able to do certain things. What do you have that lets us build something on top?" And we're able to go with them for years at a time, really, to see, like, "Oh, you made it better this year."
It's not quite ready for our use case, but that's great, let's keep working on it. So that's all really good. The downside is, once again, the unpredictability of certain edge cases. Sometimes you'll look at two images and you're like, "They're the same to me."
You put one in the model, great result. Put another one in the model, terrible result. And you're like, "Why? Why?" And the researcher will be like, "Who knows? We should, you know, collect more data, train a new version of the model, we'll let you know in a couple weeks or a month if it's fixed or not."
That's just the nature of working with machine learning,right? No one can look at the system and go, "Oh, there's the bug, let me fix it, let me push the fix, it'll be up." That's totally different from software engineering.
So that means slower iteration. You got to keep in mind, especially for us, we launched a Pixel, we have hard deadlines. We're launching in two months. That means we get, what, two iterations of the model at best. Can you promise a launch?
Generative AI9:45
Can you go to your VP and go, "We are ready to launch"? That's a totally different mindset from machine learning. So, what happened in 2022, 2023? What's the big excitement? Some of you probably know, certain things launched, they covered it too with Dolly and ChatGPT.
AI. AI. AI. AI. Generative AI. Generative AI. Generative AI. AI as AI. AI. AI. AI. AI. AI. AI. AI. AI. AI. It uses AI to bring AI. AI. AI. AI. AI. AI. AI. AI. AI. Generative AI.
Yes. So that was just a snip of I/O for that year. Uh, really, everyone's excited about AI. No more ML now, it's all AI. Larger models, more ambitious, more ambiguous, really, that's the world we live in now,right? So that's why—well, that's not totally why, but, like, a large reason why we decided to embark on building this new Magic Editor experience, which is now using the largest state-of-the-art models.
Um, but it also solves another problem. We've been building these great features, specific features, for years at this point. But discoverability is an issue. Like, the user still has to know, "Oh, in this case, I want to use Magic Eraser.
In this case, I want to blur the background." AI is great at going, "Hey, you know what? In this case, you should try using Magic Eraser." Like, it solves that, and we want to combine both of them. So from a product level, I think the previous speakers, like Paige, showed off, you know, this really amazing, fantastical ability, um, to generate, like, anything you want.
Like, some things you can capture in the real world, some things you can't. For Photos, product-wise, we want to be more grounded,right? We are the home for your memories. We don't want to generate something really weird that kind of would stand out if you were to share it with your friends or family.
We obviously will have a prompt, everything has a prompt now. Um, but we want the prompt to be more specific. Like, we want to still have the ability for users to select and visualize what they're talking about, so you can interact with the model, kind of as, like, a co-editor or a co-pilot, so on.
Server Shift11:55
And of course, we have to use the best model available, because we want to be more ambitious and do more, like, what's possible at the edge. So we're no longer constrained to on-device.
That means some challenges. Now we have to worry about servers. Before, everything happens on one device, and there are a lot of advantages there. But there's no way, back then, or even now, we can fit the best models on a mobile device specifically.
Two, the problem space is really big. Like, it's great to say, "We want to build the best GenAI image editor." And you're like, "Great, what does that mean?" Right? LLMs can do a lot of things. Editing images is a subset of what they can do.
And then, within that space, there's still a lot of specific things they can do. So you got to narrow what you are actually tackling. Because, like, image—generative image editing is not a specific user problem. The user doesn't go, "I want a generative image edit."
They have a use case, and you got to, like, meet them where their use case is. And I think the other speakers also talked about, like, trust and safety is a big one for this sort of thing, preventing deep fakes, being responsible, so on and so forth.
So, client-server. Um, this is actually new for us. I'm coming from an on-device, local-first background. I think most people are using AI server-side. So up until this point, I have never had to do server capacity planning. The user brings the compute, it is great.
I don't pay for power, I don't pay for compute. You know, a billion people want to use it, five billion people want to use it, great. Nothing changes on our end, we push the same code. But now we have to worry about it, especially because we use accelerated compute, TPUs, GPUs, you really have to do a lot of planning, and this takes a lot of time,right?
Second, latency is a concern now. Before, we had zero network latency on-device. Now you worry about network quality. Is the user in a remote spot? Is the data center overloaded and they're backed up? Is the user really far from the data center?
And now you have to go, like, a round trip around the world, and that adds a ton of latency. And then the other thing is testing is now hard,right? Before, our models were small enough that we could run our, um, tests, including the models.
Now these models are way too big. We can't run them in our automated testing suite for our daily regressions. Now they worry about, like, do we recapture their responses, do we try and have a test server, none of these are really good options here.
Focused Vision14:11
I do want to spend a lot of time talking about this ambiguous problem space. There's this old comic from XKCD. This is, like, I don't know how old, but old one that makes sense. Basically, it's like, "Hey, some problems in computer science are hard to explain why it's hard."
You know, the first one's like, "Let the—is the user in a national park? No problem, GPS, that's been solved, problem." Like, easy, anyone can do it now. And then it's like, "Check whether the photo is of a bird."
And back then, that'd be like, "I need a computer vision researcher team, give me a couple PhDs, I'll get back to you." Now this is a solved problem. So the world has changed,right? And that's the power of GenAI.
Which is great. Like, I love going to hackathons now, because anyone can sit down at a hackathon and be like, "You know what? I have this idea, and I can build it. It will definitely work in some use case."
And that's great. And when—so when your PM goes, "I want to solve this problem, because—and I used Gemini or ChatGPT, and it worked in this case, that means we can build it." As an engineer, you go like, "Did it work 5% of the time?
Did it work 10% of the time? Did it work 50% of the time or 80% of the time?" Those are—it all worked in some case, but those are vastly different things. You cannot ship a product that is 5% reliable.
If it's 50% reliable and you can work with researchers to get it to 80%, then let's talk. But you really do want to constrain the problem you're solving. I think someone talked about yesterday the comment that, like, "Prompt is a bug, not a feature."
And I agree with that,right? You want it to be easy to use. A user doesn't want to look at a prompt and think about what they want in a very detailed way. Paige talked about, like, prompt editing, and we do some of that too.
The user doesn't want to write a whole paragraph on their phone. We want to extract their intent and go, "Great, we think we know what you want, let us give you what you actually want, not what you say you want."
Right? And then, once again, anything is possible means nothing is out of scope, which means you cannot design an engineering system around it. A system, by definition, has constraints.
So really, what we focus on is reducing ambiguity across all of our functions. So product goes, "Hey, talk to users, what are the big demands they have?" Right? And they identified a few things here. One, the ability to move things within the image, to relocate,right, seamlessly.
Another, reimagining some of the scenes. So this is, like, the left side is the real part, it was a very gray sky, kind of boring. Reimagine the background to be something more exciting. And then we had eraser before, but now we can do better erase, because we're no longer constrained to the device capabilities.
So here, it's not just erasing the drink itself, but also the reflection of the drink, so it actually feels more natural. And then we work with research to highlight those, or train specific models for those use cases, so they're more efficient, more accurate, um, easier to use.
And then we build the UX and the software engineering on top to guide users to those cases, so they work really, really well, so you can rely on them. And then for the other use cases, we kind of just take the advantage of LLMs, which tend to hallucinate.
But in our case, we'll just give you multiple responses. The hallucination is a feature. The good thing is we work in a creative space, there is no correct. It's not like code. There's no compilers,right? It's like, if you see something you like out of these choices, great, take it, go with it.
Trust and safety, obviously, this will never be perfect. This is not a world where we just go, "This is the correct answer 100% of the time." You want to aim for certain precision or recall,right? And you have to manage the expectations of your users, of journalists, of media, of, you know, your internal stakeholders to be like, "We are doing what we can.
We're preventing the really bad cases." But this is something we're going to always keep going on. It's ambiguous. The language itself is ambiguous. The prompt could be like, "The view from the mountain was sick, you should make it look like that."
Lessons & Future17:55
What does sick mean? With the LLM, make sure it interprets sick in the way they meant, and not in the "they were unhealthy" way. Like, that's just how human language works, and that's why we have code. So my learnings on working with AI, once again, is like, to me, AI engineering is just software engineering, but with machine learning or ML on top.
And to me, ML is this great power, but it also adds a lot of randomness into your system. And as engineers, our job is to reduce that randomness and kind of bring things back to a deterministic, repeatable way, so you can actually have repeatable outcomes that, like, provides value to your users in some way,right?
So build evals, use your evals, make your evals faster to run, and then once you have a useful value in production, go from a large model that can probably do way too much, and it's doing way too much compute for your use case, and replace it with a smaller model, faster model, more efficient model.
Whether that's through model distillation, or you find a more efficient model, or you can even replace over traditional engineering without ML, all great stuff. The main point is to be able to move faster,right? Faster iteration means more tries, means you get better improvements in your product.
And then for what's next for us, we are—we announced last week at the Google Photos 10-year anniversary, we're rebuilding the editor from the ground up to be AI-first. So once again, really fulfilling that vision of, "We are meeting you where you are."
You can tap on your image, we'll surface the relevant edits, you can kind of adjust where they are, and you can use AI as part of that tool if you want. But, or the AI can use deterministic tools to do better edits.
Uh, I'm almost at time. This is more my personal view. Like, where will we be in two years? I don't know. You know, you talk to different researchers, they have different views on AI progress. I think Paige talked about, you know, like, the Gemma Nano on-device model now that was just released.
Is this as good as the pro model of the last version of Gemini? I hear that too. That's a very exciting world to me, to bring things back on-device. Or who knows, maybe Gemini 4 and these things keep increasing, and the Nano version is one step behind.
That's an interesting part too. Or maybe we have no progress. But either way, being able to iterate quickly and having really good benchmarks so that you can know what to change is always important, so we should just focus on that.
And I think that's my time. If you want to contact me, that's my contact, and I'll be here afterwards to talk.
Thanks, Kelvin.





