AIAI EngineerAug 23, 2025· 16:28

Perceptual Evaluations: Evals for Aesthetics — Diego Rodriguez, Krea.ai

Diego Rodriguez, cofounder of Krea.ai, argues that current AI evaluations fail to capture human perception and aesthetics, leading to metrics that misjudge generative media. He traces how compression standards like JPEG and MP3 exploit human sensory limits yet their artifacts poison AI training data, making models blind to what humans instantly see wrong. Standard metrics like FID score penalize perceptually identical images, while models cannot evaluate subjective qualities like artistic meaning. Rodriguez calls for new "perceptually aware" metrics trained on human opinions, echoing a friend's insight that predicting cars was easy but traffic was hard — the real challenge is evaluating AI's impact on creative expression. Krea.ai, an 8-person startup, invites researchers to join work on aesthetics research and hyper-personalization for generative multimedia.

Transcript

The Problem0:00

Diego Rodriguez0:15

Okay, so hello everyone, uh, my name is Diego Rodriguez. I'm the cofounder of Krea, a startup in the AI space like many others, in particular generative media, multimedia, multimodal, and all the buzzwords. But I come here mainly to tell a story about, hmm, how we think about evaluations when we have to take into account human perception and human opinion and aesthetics into the mix,right?

So I'm going to start with a very simple story. It's like, I put an AI-generated image of a hand, obviously it looks horrible, and then I asked O3, "What do you think of this image?" Then it thought for 17 seconds, obviously tool calling, does Python analysis, OpenCV goes crazy, and then after it charges me a few cents it's like, "Oh, just a couple of melting jumps are" it's like, it's mostly natural, but like, and it's like, okay, we have like what many people claim is basically AGI, and it is completely unable of answering a very simple question.

And, and like, that's a surprising thing if you think about it, because we as humans, when people see that image, it's like we just react so naturally,right, against that, it's like,ugh, what is that, like, that's not natural. And I feel like that's precisely what AI models are being trained on, A, on human data,right, second on human preference data, and third, like, in a way limited by the data that we humans base on our preconceived notions and perception and all of that.

So that's what this talk is about, about what, what, um, what can we do better, and honestly, to ask ourselves some questions that I think are not being asked enough in the field. Um, cool. So a tiny, tiny bit of history.

There's a, we all know about Claude, uh, Claude Shannon that is the, the father of information theory, and according to many, his master's thesis is one of the most important master's theses in the world, where he laid foundations for digital circuits and then eventually communication and all, and to a degree we can say that even LLMs nowadays,right, if we fast forward.

Shannon's Legacy2:21

Diego Rodriguez2:47

And I want to, I want to focus on the inform, all the, like, the fact that we

call his work foundational in information theory, well, when he published it, it was actually called math, like, mathematical foundation for communication theory, and he was always focused on communication. There's this image appears on that work, and it's all about, okay, this is the source, this is the channel, this is the destination, there can be some noise there.

And as a, well, as a founder of a company that is focusing on media, to me it's interesting to realize, like, these parallels between classic information theory and communication. Let me see, did I put the image? Let's see.

Well, if you, I didn't put the image, but if you have any context around variational autoencoders or neural networks and whatever, you can squint and be like, oh, is that a neural network? Right?

And in the context of information and communication, I want to talk about how compression is going to be related to how we think about evaluation,right? And I'm going to talk for example on JPEG. JPEG exploits, like, human nature in the sense that we are very sensitive to brightness but not so much to color, and this is an illusion that also talks about that, where A and B is actually the same color but we are basically unable to perceive it until we do this, and then suddenly it's like, oh, really?

Perceptual Hacks3:56

Diego Rodriguez4:34

And it's kind of like, what's going on there? Right?

And so JPEG just does the same thing where, okay, we have RGB color space to represent images with computers, we notice that there's a diagonal that represents the brightness of the images, we can change into a different color space that separates color versus brightness, and then we can downsample the channels around color because we are actually not even that sensitive to it, so we can remove that, or parts of it.

And then, once we do that, this is an image where we can see the brightness and color components separated. Once we downsample, we can try to recreate the image. And this is an example of, like, basically original image and then the image with the downsampled color looks the same to us, and the image is like 50% less information,right?

And other stuff. There's Hoffman coding and more stuff, but like, the point is the same,right? And then the thing is, if you exploit the same for audio, like, what can we hear, what can we not hear? Well, you do the same and we have MP3.

And then if you do the exact same thing across time, well, congrats, now you have MP4. It's like, it's all this principle of, like, let's exploit how we humans perceive the world,right?

But this made me think about myself, because I studied audiovisual systems engineering, which is engineering around all of these, how microphones work, how speakers work, and it was just interesting to me that I was coding. I start deleting information.

I know for a fact that I'm deleting information, yet, and then I re-render the image and I see the same. It's like, like philosophers always tell you about, like, oh, we are limited by our senses, but like, this is the first time that's like, I'm seeing it,right?

Like, I am not seeing the difference.

But then, if all, if a lot of our data is the internet,right, like we're stripping data from the internet, a bunch of those images are also compressed,right? Like, are we taking into account that perhaps our AIs are limited too?

Because we're kind of like, like, we have some sort of contagion going on of our flaws into the AI.

Broken Metrics7:00

Diego Rodriguez7:00

And then it gets more tricky, because for instance, this is a, just a screenshot I took from a paper, I think it's called Clean FID, and FID scores for all of you who don't have context is one of the standard metrics used for how well, for instance, diffusion models are reproducing an image.

But then you start adding JPEG artifacts and the score is like, oh no, no, no, this is horrible, horrible image. And it's like, perceptually, the four images are basically the same, yet the FID score is like, no, no, no, this is really bad.

So then it's like, why are we using FID scores or metrics along those lines to decide how this generative AI model is good or bad,right?

So the thing is, sometimes I feel like we are focused on measuring just things that are easy to measure,right, like prompt adherence with clip, how many objects are there, is this blue, is this red, etc. But

what about here? Oh, it's like, oh no, really bad, really bad generator, because that's not how clock looks and the sky, that makes no sense, and it's like, okay, how, how, not only are we limiting our AIs by our human perceptions, on top of that, we forget about the relativity of metrics,right?

Like, no, actually this is art and this is great, and there's sometimes meaning behind the work that is not, like, it is conveyed in the image, but only if you're human. You get it,right? Like, oh, this is what the author is trying to tell me.

But I feel like the metrics don't show that. And kind of like, commercially and professionally, my job is kind of like, okay, how can we make a company that

allows creatives, artists of all sorts, we can start with image, we can start with video, but to better express themselves, but how are we supposed to do that if this is kind of like the state of the art,right?

Then a friend of mine, Cheng Lu, actually he works at Midjourney, he was, he has,

The Traffic9:08

Diego Rodriguez9:18

like, he has great talks that you should all check, but he told me once, a little bit over a year ago, a quote that I just can't stop thinking about, which goes something like, "Hey man, if you think about it, like, predicting the car back when everything was horses, it's not that hard."

I was like, what? Like, yeah, it's not that hard to, like, oh, cars are the future and whatever. It's like, we have, we have a thing that goes like this, we have horses that make energy, so you swap the thing for the engine, that's essentially a car,right?

It's like, come on, how hard is that? It's like, you know what's hard to predict? Traffic. Right? And then I just kept thinking about it, I was like, oh man, like, as engineers, as researchers, as founders, what are the traffics that we're missing now?

Because I feel like everyone's focused on, like, yeah, but you can, you can, I don't know, transform from JSON to YAML, and like, who cares? That dude, who cares? Like, or yes, it's important,right? Like, but

what kind of big picture are we all missing,right?

Babel Ends10:25

Diego Rodriguez10:25

Then he talks about, well, you know, the myth of the Tower of Babel, where, in a nutshell, it's like,

God, like, we want to go and meet God, and then he's like, no, I don't want that, so instead I'm just going to confuse all of you, and then you're not going to be able to coordinate, and then you're all, each one is going to speak a different language, and then it's just basically going to be impossible to keep the thing going.

Which, like, reminds me of, like, standard infrastructure meetings with backend engineers. It's like, no, we should use Kubernetes, no, we should just, and it's like, it's just all fighting and whatever, and nothing gets built, and I'm like, dude, God is winning.

God damn it. But then this makes me think about, like, we are now in, we just entered the age where you can have models, essentially they solve translation,right, or they solve it to a very high degree. So, so what happens now that we, that we can all speak our own languages, yet at the same time communicate with each other?

I'm already doing it, for instance, I do sometimes customer support manually for Krea, and I literally speak Japanese with some of my users, and I don't speak Japanese, I learn a little bit, but I don't speak it, and like, I'm now able to provide an excellent founder-led, whatever that means, customer support level to a country that otherwise I would be unable to do,right?

And,

and so I invite us all to think about what that really means,

because this, for instance, means that we can now understand better or transmit our own opinion better to others. And on the previous point that I was talking about with the art, that's kind of like an opinion,right? Like, evals are not just about, are there four cats here?

Meta-Evals12:37

Diego Rodriguez12:37

It's about this cat is blue, and it's like, yeah, but is it blue or is it teal? What kind of blue? And I don't like this blue, and all of that. So, like, in a nutshell, it's like, how do we eval our evals,right?

Like, from my opinion, like, from my opinion, this is bad. Then I want metrics that take into account my opinion too, and then it's like, okay, consider myself a maybe a visual learner. What that means is, like,

maybe your evals should take into account how we humans perceive images,right? So, and also the nature of the data, such as, oh, it's all trained on JPEG on the internet, so take into account the artifacts, take into account, like, all of this while training your data.

Join Us13:32

Diego Rodriguez13:34

Okay, I guess mandatory slide before the thank you. A bunch of users, a bunch of money, we did all of that, we're eight people now, we're 12, and this is an email that I set up today for high-priority applications, for anyone who wants to work on research around aesthetics research, hyper-personalization, scaling generative AI models in real time for multimedia, image, video, audio, 3D, across the globe.

We have customers like those, and that's it. Thank you. Oh, great.

Q&A? Okay, perfect. Any questions?

Q&A14:10

Guest14:35

Yeah, okay. There's many points there. Can you, like, reframe the question? Like.

Guest 214:39

So we were wanting to, some of the examples you showed,right, showed that a thing is, if a guy is a person,right, he has visual and he can easily read the other person's hands. But the example you just put, it might show, you know, a person uses hands, they're all hands, but the people who want to actually talk about people that are using his hands in their eyes, that makes it very difficult for us to get the information at the same level.

Guest15:09

Yeah.

Guest 215:09

So

the question, like, in a nutshell, is like, are there perceptually aware metrics,right? Like, okay, I showed an example of FID score, it changes a lot with JPEG artifacts, are those where it's almost like the opposite, barely changes, and the metric is still good.

Like, there are some, and many of these are used also in traditional encoding techniques. But in a way, I'm here to invite us all to start thinking about those, like,

like, to, we can actually

train, like, we can train, I mean, it's called a classifier,right? Or a continuous classifier. We can train so that it understands what we mean, and it's like, hey, I show you these five images. These five images are actually all good, and then they can have all sorts of artifacts, not just JPEG artifacts.

And this is exactly where machine learning excels,right? When it's all about opinions, and it's like, let me just know and you will know, you know, you know what, you will know when you see it. That's precisely the type of question that AI is amazing at.