Intro0:00
Thank you all for coming today. I really appreciate it. Uh, thank you for coming to this talk, which is Black Forest Labs: FLUX, open research, and the future of visual AI. I'm going to start quickly with a quick intro of myself.
I'm Stephen Batifol, sorry, I'm a developer-relations engineer at BFL, and I want to start with two questions. First, who here knows about BFL? Raise your hand. Okay. Who here knows FLUX? Okay, about the same people, actually. But for the people that don't know BFL, I have a quick intro so you're not lost.
But BFL, at a glance, we are the team behind Stable Diffusion, Latent Diffusion, and the FLUX models as well. Our team has more than 200,000 academic citations, and we don't only build models; we actually also work with enterprises and customers with them.
FLUX.1 & Kontext1:07
So some of our customers are Microsoft, Adobe, Canva, Mistral, and many more. And the way we started is we started in August 2024 with FLUX.1. FLUX.1 was the first breakthrough, you know, that was the model that was the big competitor to Stable Diffusion back then, and that was really the breakthrough where people were like, "Oh, this is a really cool model."
We released it in open source in the first place, so really that was the one that was text-to-image only, and you could run it on your laptop. That was a game changer as well. And the anatomy was really good in comparison to the other models, and especially in comparison to other models that were way bigger.
So this is where we really had a breakthrough, and Clem from Hugging Face actually gave us a shout-out back then. This is fairly old, but FLUX was actually the model that was the most liked on Hugging Face back then.
This is not true anymore, but back then that was really the thing, and that was really good, and really big, actually, for a company that was, you know, coming out of nowhere and just released this model. We then released FLUX.
Kontext, which was the first open-source editing model in the world. That was like the combination of text-to-image and image editing as well. This one now, what I'm showing you here, what I'm showing you here, you know, it's obvious because now we have editing models everywhere.
But back then there was a big breakthrough where you could do both text-to-image and image generation at the same time. If I have an example here, we have this input image, and then you remove the snowflake from the face, you know, you can see you have the character consistency.
But you can also then move this person to be in Freiburg, which is where our headquarter is. You know, she's chilling, taking a selfie in the streets of Freiburg, and then you can, you know, do some local editing where you can change the background to be snowy, and then you also have snow on her face and everything.
It was also the model that was really one of the fastest back then. You know, if you remember, this is the time where you had the first GPT image, where it would take like 40, 50 seconds to generate or edit images.
Whereas Kontext, if I remember correctly, was like 7 to 8 seconds. It was also really useful to tell stories. I've seen a lot of use cases from our partners, from our customers, where they would start with an input image and then create a storyboard like we see here.
We have the famous seagull here, you know, which has the VR headset drinking a beer in a bar. But then you can actually create other things. They can have a friend that is joining and that is then drinking with them.
Now, you know, I guess they got a bit tipsy and they're wearing hats in the bar, then they're going outside, and then, you know, this is a story you could create. And that was really useful, actually, for video model or for animation models.
You know, you would give those images as input frames or as end frames, and then the video model could then create different content.
FLUX.2 & Klein3:58
In November, we released FLUX.2, which is our steps towards what we call visual intelligence. FLUX.2 was, uh, our base still, our base, our best model, sorry. Those are our samples, which I don't know if you can see clearly, but in my opinion, they are like really amazing samples, and it's impossible to basically tell they are AI-generated.
If you look at the hands, if you look at the veins, if you look, you know, at the bracelets of the person on the left, there's personally no way that I would tell it's AI-generated. Same for the turtles you see on theright side, same for the dog or cat, actually, that's in the bath.
That would be very, very hard to get a sample like this, but those are AI-generated. And then you have more as well, so it's not only the people or animals. You know, you could do like some proper product photography.
You can see it with a waffle here on the bottomright. Or you can make some very cute images, like this person, you know, on the left that is ri- driving the moped with some balloons. And this is what we released in November.
But it's not only an image generation; it's also an image editing model at the same time. You can see on the left we have six images that we give to the model, and my prompt was literally like, "Create an outfit with those images."
And then the model is intelligent enough to actually, you know, make things that make sense. Like the jacket, you know, is worn properly, same for the tie. And on theright side, it's a bit more of a simple use case where you have the sofa, and then you have to imagine, you know, maybe you're an e-commerce website or you're a sofa maker, and then you want people to imagine, you know, what it looks like in your flat or what it really would look like if you were to buy it.
And those use cases are really, really important, and those are like the main use cases we have currently for FLUX.2. But it also takes, yeah, up to 10 images simultaneously, so you can really edit a lot of images at the same time, and then you can create magic things.
It's very good at character, product, and style consistency.
And what I want to make clear is that BFL, as a company and as a research lab, our first operating principle is to release state-of-the-art models. This is what we want to focus on. This is what we want to do as a company.
You know, we want to raise the bar on quality with every release we do. So we did it in the past, you know, with FLUX.1 when it came out. We did it with Kontext. Our FLUX.2 was our best image model to date.
It's the first one we released that was actually multi-reference as well. It was state-of-the-art in the open-source world. It was state-of-the-art for text-to-image and image editing. In January, we released FLUX.2 Klein, which is a step towards, like, interactive editing and interactive generation.
It generates and edits images in less than a second. I'll talk a bit more about it later on during the talk, but I think the fastest it can do, if I remember correctly, it's 500 milliseconds for editing and 300 milliseconds for generation.
So basically real-time. But this is not it. We also have more things that are coming, and this is where I want to talk about today. So I mentioned it. We are a research company first. We publish things in the open.
We publish paper. We really want to make sure that the field is moving forward with us, you know. This is our big focus as well, so it's state-of-the-art model, publishing things in the open, and that's what we want to do.
Training Background7:28
But I want to first take a step back and tell you a bit, you know, about, like, how do you train models, and especially generate models that are generating content, generating images, and everything. You know, when they generate things, when you train them, they actually don't understand what they're generating.
You know, they don't understand that my glass here should be actually undistable. I shouldn't go through it. Because you train them, you have images, and then, you know, you add some random noise to those images, and then you just try to denoise them.
That's what you do. That's what those models are doing. And when you denoise images, you never learn, you know, that my glass shouldn't go through here. You never learn that, you know, you're sitting on a chair, you shouldn't go through it.
So what do you do is that you use, you do what is called like representation alignment. So you use an external model that actually knows about this, and that is an encoder that is like an image encoder that is teaching our model, "Hey, he is currently sitting on the chair.
He shouldn't go through it." And those models are external, and they are really, like, trained to segment images, whereas our models are trained to generate images or generate videos or generate audio. And you try to align them to be on the same objective so that our generative model actually learns, "Okay, you shouldn't go through the chair.
My glass should stay on this table." And this is great because it really improves the way generative models are working. We can see here on theright, you know, it is 70 times faster to actually converge and to reduce the loss when you use this external alignment.
So you're like, "Okay, this is great, but as usual, if something is working well, there are also counterparts to it." So the first one is that you have a scaling ceiling. You imagine you're working with a model that is external, that has been trained.
It's at a checkpoint. You're not changing it anymore. What if you train a new model and you have a generative model that you want to scale up? You're still, like, limited by this encoder that you have on the side, you know.
You're never actually scaling up fully with it. Also, those are specialized modalities. You have an encoder, for example, DINO V2, and the other one that I can't remember, it's specialized in images only. What if you want your model to generate images, audio, video, and more?
You would have to have encoders for all of those. And you can imagine then you would have like a very Frankenstein setup, you know, nothing would really make sense. And the objectives are also misaligned, so I've shared it before.
We want to generate content. We want to generate images or audio. The other one is here to segment things. So, like, they have different objectives, and you're trying to make them work together. And it works great, but it's also not perfect.
Here, you see on theright side, we have DINO V2 and DINO V3. DINO V3 is a better model technically, per se, than DINO V2. But when you train your model, you're actually getting worse performances. You know, DINO V3 is here inright, in red and green, and so you're getting worse performances.
So you're like, "Okay, like, this is supposed to be a better model, and yet when I do train a model to generate things, then it gets worse." And there's also, like, not really any rules as to why, you know, certain encoders should work or otherwise shouldn't.
So how can we solve this? How can you teach, you know, a model representation directly without this external encoder? This is what we released about a month and a half ago now, which is a research paper. You can read it.
It's called Self-Flow. It's in the open. We released it to really make sure, you know, we're moving the field forward again, and it's not only us benefiting from it. And it's basically a scalable approach to training multimodal generative models.
Self-Flow11:05
So they use self-supervised learning, so you don't need any other models, you know, to train it. And I'm going to try to go in a tiny bit more detail into it, but we combine representation learning and generation in the same flow.
And you see here on the left, you have videos, images, audio. You have different modalities. What do you do when you usually train a model? You add some noise. You add some random noise. You try to denoise it, and then you align it with the encoder, you know.
How do we do it then? We actually add two different kinds of noises that are both random, and they're both different. The first one we're adding is actually we're adding a lot of noise to the assets, so this is the one you see at the top.
And the other one we're adding like a low amount of noise. This is what you see at the bottom. And the idea is that then we have two models that are actually working together. We have the student one, which is always getting the images with the most noises and is trying to denoise them.
And then the teacher one, which is basically a more stable version of the student, is always getting the low noises images. And then the student one is actually trying to learn two things at the same time. It's trying to minimize the loss for the generation and the loss in representation.
And this is how then you actually work across different modalities. You know, this is then you only have one model. You don't have anything external. And if you actually scale up your model, then you're scaling up your student, you're scaling up your teacher, and you don't have to worry about the encoder that you have on the side anymore.
And this is where we're working on. This is something, you know, we are currently using for different models that we're training. And this is, you know, where we believe the future is going to be and to get rid of those encoders that we have.
We actually trained models. So those, disclaimer, those are research models. They're not meant to be released in production. But we released actually one model on all those modalities. On the left, we're comparing flow matching, which is the usual way of training models, with ours.
And you can see we are better in audio, so this is what you see in orange on theright side. And then we're also better in images, so the dash lines is the baseline. And then we are like the full line where we can see we are also better at images and also better at video.
So with this approach, without having the encoder and the external model that you may struggle with, you actually get better at every modality that you're training your model on. It's also converging faster. You can see on theright, you know, the baseline is converging.
It's actually hitting a plateau, whereas we are converging faster, and we're still, you know, decreasing the loss. And I'm pretty sure that if we were to go towards 2 million steps, you know, the baseline would really plateau, and then it wouldn't really get any better, maybe actually get worse, whereas we would still go down in loss.
And this is the difference between the two. If you used FLUX in the past, or if you used different models, you know, to generate, you may have noticed the text might not be perfect, or, you know, things don't really make sense.
This is what you see at the top, where on the left it's like the future is FLUX, but you can see, you know, you have like some letters that are missing, or maybe you have two letters. Like on words, for example, you have two Ls instead of one.
Whereas with this approach now, and the Self-Flow approach, you can see at the bottom everything makes sense. There's like, they learn representation, you know, they learn that FLUX then for the letters should be like one next to the other.
And same on the mirror, same on the tree. And this is where we believe this is the future. But on top of this, we can see some comparisons here. On the left is the baseline, where again, the letters are wrong.
On theright, you can see that the letters are correct. Here is the same for the anatomy, where you see on the left, you know, you have like a face that's looking a bit odd. Let's put it that way.
And on theright, this is the one with Self-Flow. And again, this is not like a production model where, you know, you expect the face to be like perfect, but you can see that the anatomy is way better than what you have on the left.
But I want to show you as well some different generation if it loads. Yes, thank you. This is also possible. This is also possible for video generation. So this is the same model that has been trained on images, now also can generate videos.
On the left, you see the baseline. It's a weird way to do a push-up. Let's put it that way. Whereas on theright, it's a perfect form. You know, the arms are correct, the hair as well is correct, and nothing is wrong in it.
And this is, you know, a way to actually fix all those artifacts that you may see usually in generations. So same here for the birds. Oops, my bird. Yes, thank you. It's the same here for the birds where, thank you, where you see on the left side, you have the baseline.
There's a lot of flickering. There's a lot of like, you know, weird things happening because the model was using this encoder that was trying to align things. Whereas on theright side, with Self-Flow, it just does it perfectly, and like the bird, you know, is walking on the floor, and there's no flickering or anything.
But it's not only about images or videos or audio. You train those jointly, so you can also actually generate things jointly. We have here an example of a video and audio sample where the idea is that we have someone that is saying hello from the Black Forest.
I will just play them, and you will hear the difference. Again, this is not a production-ready model, so it's not like perfect, but you can hear the difference, hopefully.
Oh, from the Black Forest for a fact. Oh, from the Black Forest. Oh, from the Black Forest for a fact.
And so this one was the baseline, where if you hear it correctly, if you try to pay attention to what he's saying, you hear like, "Hello, from the Black Forest." There's a bit of like weird things at the end.
Oh, from the Black Forest for a fact. Oh, from the Black Forest.
Whereas on theright side, you can see, you know, the prompt is really just say, "Hello, from the Black Forest," and then it ends here. And yeah, this is the same model that was trained on those images that we've seen before on video and on video and audio.
Oh, oh, from the Black Forest.
Thank you. But this is cool, and this is great, but what if you could also teach robots on how to use this? This is also the same model. This one is trained on actions and not only on images, video, or audio, so it can also predict actions.
And what I'm going to show you now, it's a robot that is trying to pick up a can and make it closer to us. On the left, this is a baseline. Again, you see some like flickering. You see like the arm is doing weird things.
Whereas on theright, for the same amount of steps, you can see Self-Flow, the robot is picking up the arm directly and like bringing it closer. And this is where we're going as well as the company. This is where we're really interested is like not only image generation or video, but it's also doing actions and doing more things toward physical AI.
The Road Ahead18:26
And there is more. It's also how do we make our models faster? Because this is really important for us. This is a demo of Klein, which is, you know, like near real-time editing. You see it on theright side.
This is, you know, generated with Klein on Korea, where you see the edits. And this is not a video model. This is, those are images that are always editing in real time.
Not only they are faster, they're also actually at least on par or better than other models. And I'm almost out of time. Oh, they added five minutes. So I don't know if I'm, okay, cool. So I'm not out of time, so I can chill.
So yes, here on the left, we can see, you know, we have Klein that is 4B and 9B that is compared to the other open-source models. So it's at least on par, while the latency, you know, it's like 0.5 seconds, while Klein is like around like 15 seconds, you know.
And if you are like on par and you're like way faster, then this is really, really good for us. Same for image to image. You can see the editing. For Klein 9B, we are at like a tiny bit more than 0.5 seconds, whereas Klein is still at around 15 seconds.
And then same for multi-ref. You add it, but we're still at less than a second, whereas Klein is more towards the 20 seconds. And this is what is really, really important for us because you really want to actually generate things in real time.
This is where we believe the world to visual intelligence. This is where we're going as a company in the future. And why does it matter? It's like, I mentioned it, real-time generation. So you can imagine your random mockups as fast as you think.
You know, you don't have to wait. You don't have to wait like 10 seconds, 20, a minute or two. You do things in real time, and you can guide them in real time. This is also where we're going.
You can think, you know, interactive visual engines for gaming or films, where you really render a movie as you prompt it. On top of this, there's also world models. The idea of world models and behind it and why it matters for us, it's you train your models to understand and simulate geometry, relationship, and like different interaction of the world.
And you may be like, "Okay, that's cool from a research perspective. Why do we care?" The reason is robots. That's why we care. That's why, you know, robotics and automation. This is where we're taking BFL, and that's why we also want to go towards world models, is to train agents in those generative world to scale self-driving and automate every manufacturing.
I think that is it. Thank you very much.
Can we take, yeah, I think we can take questions.
Q&A21:09
Can you share something where you base your training on for the world models?
You mean the data?
Yes.
Can't really. This trade secret, I mean, data is very sensitive, as you can imagine, so I can't really share this. We're partnering with a lot of people, though, for it.
How do you store the state of the world in those action prediction models?
Mm-hmm. Well, this is what the model is learning. It's basically like those representations. You know, the model is learning that in itself as like the state, and it has like some kind of memory. And this is the way we do it.
What is that some kind of memory? Is it like in the context window, or is it external and it gathers it, or?
No, it's, yeah, it's the context window that you have. You know, you train, and then you have the tokens, and then they'd be like, "Oh, look, I've moved. Here is where I should be then next."
And can you then run it for long, or does it do some compaction?
Define long. What do you call with long?
Indefinitely.
Indefinitely. That I'm not sure. I mean, there's always going to be a limit. So you may have, you know, like a sliding window, but this is the way we sew it. Thank you.




