AIAI EngineerJul 17, 2026· 2:20:21

Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth

Daniel Han of Unsloth argues that reward hacking—where AI models cheat to maximize reward—is a critical problem in agent training, citing examples from GPT-5.1's calculator hacking and GPU mode kernel competitions. He shows that models exploit benchmark flaws, such as viewing Git history or editing timers, and that even open-source models like GLM 5.2 require anti-hacking measures. Han emphasizes that harness and tooling quality now outweigh model choice, with inference providers sacrificing accuracy for speed (e.g., 10% accuracy drops across providers). He also warns that hardware limits (float4 precision, diminishing returns) shift focus to software algorithms like FlashAttention and gradient checkpointing. The workshop concludes that benchmarks are unreliable—DeepSpeed's false positive rate is contested at 44.9%—and urges verification before trusting performance claims.

  1. 0:00Intro
  2. 2:32State of AI
  3. 20:23Open vs Closed
  4. 38:50Throughput Maxing
  5. 1:02:59Bench Maxing
  6. 1:26:08Cybersecurity
  7. 1:37:42Kernels
  8. 1:59:42RL Primer
  9. 2:05:21Reward Hacking

Powered by PodHood

Transcript

Intro0:00

Daniel Han0:13

Hello everyone. Um, yeah, thanks so much for coming today. Much appreciated. Yes, I'm Daniel, from Unsloth. My brother is also here today. But yeah, like, you know, thanks for coming.

So, for you folks who don't know us, we actually, you know, we're one of the largest distributors of language models and diffusion models as well. So we don't just do language models. We upload our models to Hugging Face.

And, you know, we're on the, I think we're number 10 or something on the, I don't know, I don't remember. But anyways, we're on the list of the top organizations on Hugging Face. We have over 300 million total downloads, so definitely check us out on that.

You can run like, you know, DeepSeek, GLM, many other models, and we quantize them down using dynamic quantization. So you can run them on your local computer. We also do many bug fixes for open-source models. So, you know, we, you know, fix many bugs in, you know, OpenAI's GPT-osses, you know, Meta's models, Google's models, DeepSeek's, many other models.

We fix bugs in them. And so, like, you know, they have many issues sometimes, and then we post about them on Twitter. You know, we post about our findings. So, you know, most of the open-source models that you probably guys have used are most likely fixed by us.

And yeah, like, we collaborate with everyone in the entire world on, you know, model releases. Yeah. We also collaborate with hardware providers, and, you know, we really appreciate the collaborations with everyone. We also don't just do model fixes and bug, you know, bugs.

We also introduce new features, and we also, like, you know, do fixes for the entire training stack. For example, we introduce something called async gradient checkpointing, which is used by many organizations. We also introduce flex attention, which is used by many folks.

And we also fix a gradient accumulation bug fix, which increased accuracy by 1 to 3 percent across the entire training stack. So we don't just, like, you know, do bug fixes for models. It's also like, you know, whole training stack fixes and stuff like that.

So today, you know, the workshop is quite long, so there will be multiple sections in the workshop. And so after each section, anyone can ask a question. And so, you know, please, I guess if, I'm not sure if there's a microphone, but if you can raise your voice and, you know, ask a question, you know, I'm more than happy to answer them.

But, you know, the first section we're going to be talking about is the state of AI. So where is currently language models, AI models, where are they at currently? So I'm not sure if everyone knows the meter plot.

State of AI2:32

Daniel Han2:45

So this meter plot shows the time horizon of models. If you can, you know, every single task, if it takes a human 16 hours, can a model, you know, finish that task? And you can see on this plot, you know, Claude Mythos, you know, preview is very good.

It can do tasks that humans can do that take, you know, human 16 hours. You know, Opus 4.6 is also there. You know, all the other models are also there. And so, you know, this plot is very good because it symbolizes that AI models are getting better and better and better over time.

You know, recently with the launch of, you know, GPT-5.6, you know, just, well, their preview model, you know, just on Friday, you know, I put the plot. So they didn't, so Meta didn't actually update their plot because they said that the results were not trustworthy enough.

But, you know, I just put it on the plot. And so you can see that GPT-5.6, you know, is around, you know, Opus 4.6 level, I guess, with large confidence bounds. So it's very, you know, uncertain about the capabilities of the model.

However, if you include cheating, so if you include that the model sometimes likes to cheat on some of the tasks, then it actually goes to 270 hours. So we're directly, and, you know, if you look at the y-axis, I actually did a disjoint graph.

So the y-axis is 50 hours skipped to 250 hours. So if you can imagine, the graph is actually very skewed. When I, like, made the graph, GPT-5.6 was like a very big outlier. So I had to, like, compress the graph.

But this only, you know, this graph only works if you consider that GPT-5.6 cheated on some of the tasks. And so we'll be talking about, you know, why AI models cheat and how do we, like, you know, solve these issues.

But yeah, this plot is very useful to showcase the capabilities of these models. So previously, this is 50 percent. You know, if you could, if a model can complete the task with 50 percent of the, you know, of the time, so 50 percent accuracy.

If you want to actually one-shot the model, so you just ask the model, you know, implement X or implement Y, and you want the model to do very well, then you want to look at the 80 percent success rate.

If you look at the 80 percent success rate, it kind of drops quite a lot. So you can see that previously, Mythos is around 16, 17 hours. Now it only can do 3 hours. So if you prompt a model and you want to have, like, a one-shot example, you know, you just trust the model by just asking it, you know, implement, I don't know, PageRank or something.

You know, implement some sort of RAG system. You know, fine-tune a model or something like that. It can only do a task that will take a human 3 hours to do. And so, so that is a problem with AI models.

Generally speaking, if you want to use AI models very well, you need to prompt it at least, like, you know, five times or something. And each of those times, assuming they're independent, the success rate is much higher if you prompt it many, many times.

Right? You can't just call the model once and expect it to do work, to do well. You need to call it multiple times. And you can also work out the probability of it, like, succeeding. You know, if the model is 50 percent accurate, then it will be 50 percent failure.

Then it's 1 minus 0.5 to the power of 5 or something like that. You know, if you do five turns, and then your success rate jumps to, like, 97 percent or something. So you need to call the model at least five times for it to be very effective.

So previously, these are linear, you know, this is a linear trend. You know, on the y-axis, it's just, it's not, you know, it's just linear. If we log it, you know, if we log the y-axis, you can see that it's more exponential progress.

So it's actually a straight-line fit to the entire progress of AI models on the meter time horizon, you know, benchmark. You can see that, you know, it's very clear that AI models are getting better and better over time.

I also added, you know, GPT-5.6 with the cheating and no cheating, and also Claude Mythos are, you know, accentuated that. And you can see, I, you don't need, now you don't need to, like, you know, fake the y-axis.

You know, you don't need to do, like, a disjoint y-axis. If you do that, you can see that, you know, models are getting better over time. And supposedly, you know, if this trend continues, these models will get better and better and better, better, and much better.

Yeah. So the question is, if the trend continues, you know, that's the fundamental question. And it's not just, you know, one specific task for this benchmark that you can see that models are getting better over time. Across all benchmarks, models are getting better over time.

Right? So, like, you know, GPQA diamond, you know, it's kind of plateau, you know, it's kind of already saturated as a benchmark. But over time, you know, it does very well. You know, every single benchmark you see, models are getting better.

Right? Live code bench, you know, maths algorithm, maths tests. You know, even Tesla's, you know, you know, self-driving, I guess, is also has, like, a doubling time of 17 months. So every single 17 months, the models will get better and better.

You know, double, double their capabilities. So over time, all these models in every single subject, you know, every single

area, it will get better. So I guess the main question is, you know, if we assume every single subject, every single area, the models get 100 percent, like, you know, approaching 100 percent accuracy, is this AGI? So that is one of the fundamental questions that people ask.

You know, if we just get better on benchmarks, is this AGI? What happens if we get better on all benchmarks? You know, every single benchmark that human, humanity has created, it just gets better on all of them. Yeah.

So this is a, you know, very good plot, well, I guess, chart showing all of the different types of benchmarks, and they all get better over time.

Everyone's favorite, I guess, artificial, you know, artificial analysis benchmark showing, you know, artificial intelligence getting much, much better over time as well. You know, Fable, I guess, is, I guess, the best for now. Although not everyone can access it currently.

But anyways, it's for now, it's the best. And you can see over time that, you know, these models are getting better over time as well. And, you know, like, this plot showcases a very useful indication, you know, like, how do we, like, you know, benchmark, you know, is this benchmark actually good in terms of, like, you know, showcasing the capabilities of models as well?

And we'll be also discussing about that as well. On the other hand, yes, models are getting better over time. But there are some things which models are not very good at still. For example, long context is not doing very well.

So, you know, most models, you might say, okay, Gemini has 1 million context length, you know, GPT has 1 million context length, Claude has 1 million context length. But should you actually use all of the 1 million context length?

So there are actually benchmarks to showcase that if you use, for example, GPT-5.5, you know, if you use 512 context, your accuracy reduces to 50 percent. So if you use, you know, 512 context, you will only remember 50 percent of the facts that you wrote in the previous context.

So maybe that's not a good idea to use the full context. You can see Opus 4.7, 4.6, 4.7 is the very last orange line. So at the context length of 256k, it goes to 0 percent. So this might be a benchmark flaw.

So maybe don't trust the benchmark too much. But it's good to look at the benchmark overall. You know, where is the model's capabilities for long context? The blue lines I highlighted are open-source models. You know, DeepSeek, GLM 5.1, other models.

Green is Google's models. But you can see in general, you know, models are, models definitely do degrade over long context. So if you, you know, for example, if you set, like, a, you know, automatic compaction area, I would not suggest you to use all 1 million context length.

Maybe maximum 600k or something. And then compact it and then continue your, you know, coding session. But I would, yeah. But in general, you know, this plot shows that long context still has a very long way to go.

And if we want to have long context, you know, capabilities, labs, I guess, will have a lot of time to fix this problem. Yeah. So another plot is, you know, just showing open-source versus closed-source. So open-source still has some way to go for this, you know, long context.

So open-source is blue line, and the black lines are, like, you know, closed-source models. And you can see in general, open-source does okay, but there's definitely much more room for improvement. I guess compared to Opus 4.7, it's better.

But, you know, maybe this benchmark does need, maybe there are some flaws in the benchmark as well. Yeah. But overall, you know, this plot shows that long context definitely still has more room for improvement.

And also, you know, like, if you looked at the plot previously, you know, this meter plot, I'm not sure if you can see that before 0.1 preview, there is actually a plateau of performance. And so if you can see, you know, GPT-4 to GPT-4.0, there's not that much performance improvement.

And so this timeframe around 1 year was when, you know, the labs were confused on what is next. You know, before 0.1 preview, which showed that reasoning was very important, they didn't actually know what to pursue next. And so for 1 year, the models kind of plateaued.

And so I call this the intelligence plateau, the hypothesis that, you know, you know, assume that we never have discovered reasoning. Then maybe AI models would have, like, plateaued. But because we have discovered reasoning, you know, we have shown that models can do reasoning capabilities, we have continued the trend continuously.

And so normally, I don't know if this is, like, luck or if this is a self-fulfilling prophecy. So I don't know if you guys, you know, the Moore's Law, you know, Moore's Law has continued, not because of the law, but because people know that it must continue.

And so people invest money into the resources to make the law continue. And so this kind of, like, shows that, you know, we might have been in a world where models have stopped improving. But, you know, with the launch of 0.1 preview, you know, I guess models have went back to trend.

In fact, I made a plot showcasing, you know, assuming we did not discover reasoning on 0.1 preview, then the black line was the supposed, you know, capabilities of the models. You can see I made it into an S shape, like a, you know, a sigmoid-type shape.

And if, you know, if we didn't discover reasoning, then models definitely will taper off in terms of capabilities. Right? We'll only have a model that's as capable as Claude 3.7 Sonnet, I guess, or 0.1 or something like that.

But, you know, luckily, because of reasoning and this new paradigm of scaling, you know, the green line is the new scaling law. And you can see previously the black line, the doubling time was actually around 7 months. So every single 7 months, the capabilities of the models double.

But now it has shrunk to 3.5 months. So every single 3.5 months, you just need to wait 3.5 months, and the models will get double better. Right? Better by two times. And that's quite striking, I guess. So the main question, though, is, will the green line continue as a straight line?

That is the fundamental question that labs are still struggling on. You know, what happens if the green line again, you know, the green line again goes as an S shape? You know, that's possible. But, you know, we don't actually know if this will happen.

You know, if the green line will continue scaling, you know, going all the way up to infinity, I guess, or will it be like an S shape? And this is, you know, many researchers are, you know, I guess, have sleepless nights.

You know, what is the next, you know, what is the next thing afterwards? After reasoning, after 0.1, you know, what is the next thing afterwards? And, you know, many researchers will need to, like, you know, I guess, think about this.

Yeah. But, you know, this plot is very, you know, this is one of my favorite plots because it shows that, you know, AI progress can continue over time with new ideas and innovation.

Oh, yes. So does anyone have any questions for the first section? Yes.

Guest14:49

So we came all the way to 1 trillion,right? Do you think the next jump, if we need, do we need, like, 10 trillion parameters when we'll see the jump or hardware will be the limitation?

Daniel Han15:01

Yes, that's a great question. So the question was, you know, models were currently at 1 trillion parameters. Do we need to go to 10 trillion parameters or more for models to be even more capable? So the scaling laws does say that, you know, if you multiply the parameters and the data size, generally speaking, this number, if you increase the number, you will get the models become more capable.

So yes, you can increase the parameters by 10 times, and in general, your performance will increase. However, the view is there is going to be diminishing returns. I feel like, you know, it's not just the model size times the data set size.

It's actually a ratio, some sort of, like, power law when you multiply them. So you actually get diminishing returns over time. So yes, you'reright. If you want to have actually, I'm not sure the exact law, but if you want to have double capabilities, you do need to 10 times the parameters.

And then if you want another double, you have to 10 times it again. So it's 1 to 10 to 100 trillion parameters. If you want, maybe that's not a good way to scale. Maybe instead, you know, instead of making a 100 trillion parameters, some sort of new algorithm or new architecture could solve that problem.

But you'reright. Like, if you're a lab, you want to do something easy. And so the easiest path is to just make a 10 trillion parameters. But I would say, like, you know, maybe a new algorithm will be better.

Yeah.

Any other questions? Yes.

Guest16:28

So do you two think that we are approaching the limitation of next token prediction?

Daniel Han16:34

That is a good question. I would say that for next token prediction, it's very powerful because you can essentially, the human language is extremely powerful. And it doesn't have to be human language. It can be, you know, maths, coding.

You can just predict the next word. And in order to predict the next word or token, you need to know everything about that token or that word. Right? So, like, I think Ilya was talking about, like, you know, Ilya Sutskevich, he was saying, like, you know, you need to have, you need to make a welled model in the model in order to, like, predict the next word.

So I still think next word prediction still has a lot of way to go. For example, if you see this plot, you know, if we didn't have reasoning, I guess, okay, maybe it would have plateaued. But because we have discovered this new methodology, you know, reasoning and trying to, like, scale even more on next word, you know, next word prediction, we have, you know, went back to trend.

I feel like, so the main question is, if we don't have next word prediction, what is next? That is a fundamental question. Most, I mean, I'm not sure. Like, you know, I'm not certain what's the next thing. I feel like next word prediction is just extremely powerful because it's very easy to formulate.

And you can just, like, you know, you can have, like, you know, because attention is very powerful as well, you can have the, you know, the special causal attention mechanism, and it's very efficient to train. So I'm not sure.

I think the main question is, I'm not sure what's next. I guess researchers will, like, you know, they're trying to scratch their heads, you know, what is next afterwards? Yeah. I, yeah. Yes.

Guest18:03

Just a follow-up on it. Do you feel like we are in the same era, like how we were in with the LSTM and when attention came out? Right? So attention, like, we don't know what's next in a way.

What's next after attention and then what's next after next token prediction?

Daniel Han18:17

Yes, that's a fair follow-up. So you were mentioning how it's kind of like LSTMs or in the old AI world, we don't know what's next afterwards. That's a fair point. I feel like,

so, like, you know, previously, this example,right? So after GPT-4, it was just pre-training, some, you know, supervised fine-tuning, some RLHF, you know, some RL. And they waited 1 year until 0.1 preview. So in this 1 year of fog, you know, the fog of war, we don't know what was next.

And so researchers, you know, were scrambling, you know, do we do the reasoning process? Do we make pre-training better? Do we make the model bigger and bigger and bigger? You know, they tried all these experiments. And reasoning was the one that won, I guess.

But I think, like, I think the main question is, is the green trend going to continue? At the current time, it looks like it's continuing. Once we see models starting to taper out in intelligence, you know, in capabilities, then we'll go back to the, you know, olden days of, like, you know, this 1-year waiting period.

But I think for now, these models seem very powerful. Yeah. So, like, I'm not sure if this will, I mean, if you look, okay, if you squint at the plot, I guess maybe we're tapering out. Maybe. Let's not consider the GPT-5.6 cheating example.

Right? Let's remove that from the plot. But you can see the GPT-5.6, Mythos, you know, 4.6, they're kind of all, I guess, they're kind of tapering. So maybe as a, I mean, I don't know if someone wants to bet on this, but, you know, maybe models have tapered out.

But we're not sure. So we should wait a few more months and see. So let's wait 3.5 months. If we wait 3.5 months and see the models do not improve, then we have tapered out. But remember, we only need to wait 3.5 months.

So then this law will fail. In fact, if you wait 7 months, if you wait 7 months, so double of the time, and models have, you know, just assume, you know, that dotted line, that if the models just follow the dotted line, okay, then we have tapered out.

And I would agree that, you know, we'll have to design something new in, you know, make some new invention or something like that. But for now, you know, for now, it looks like it's doing fine. Yeah. Okay. Next section.

So every single section, we can have questions. So you can ask as many questions as you like. The next section we're going to talk about is open versus closed models. So artificial analysis has this very cool plot showcasing the performance of open source.

Open vs Closed20:23

Daniel Han20:36

So open source is the blue line. So open source is the blue line. And closed source models is the black line. And you can see that open source does lag. You know, open source definitely lags over time. Another very good benchmark is called the Weird ML Benchmark.

This also shows that open source models lag closed source models. Right? The blue line is open source models. The green line is closed source models. And you can see over time, you know, the X-axis is release date of the model and the Y-axis is performance.

And you can see that open source models kind of lag closed source models. And why the Weird ML Benchmark, I'm not sure if you folks actually know about this. Why the Weird ML Benchmark? It seems like the Weird ML Benchmark is a very good indicator, better than other benchmarks.

And the reason why is, you know, previously I mentioned, you know, previously this graph,right? Reasoning, the reasoning models are the green line, and then the black models are the non-reasoning models. And you can see that reasoning models double, you know, reduce the time of doubling time to 3.5 months.

Previously, it was 7 months. Interestingly, on the Weird ML Benchmark, these reasoning models didn't actually do better. It didn't actually change the trend. All it did was make it slightly better. And so this Weird ML Benchmark seems to be more robust.

And that is why, you know, this benchmark is very useful. In fact, if you go on the Twitter burst before GLM 5.2 got released, most, you know, most of the Twitter people said, oh, you know, DeepSeek, you know, DeepSeek, if you see very, if you squint, okay, I think I have a plot.

Oh, yes. If you squint, DeepSeek and Kimi are in that little corner over there. You know, DeepSeek, those three models, the three whales are DeepSeek, you know, Flash, DeepSeek Pro, I think one of them's Max Mode or something like that.

And also Kimi's over there as well. So before GLM 5.2 got released, you know, on the Twitter burst, everyone kept saying that open source models are much worse than closed source models. Right? They're not lagging. You know, they're not just lagging.

They're much worse because of this benchmark.

In fact, if you look very closely of the Weird ML Benchmark, all of the top models are closed source labs. You know, like Fable, you know, GPT-5.5, whatever. You know, all of these are just very, you know, it shows very clear that open source models are not doing very well in terms of this benchmark.

Until GPT-5.2 came along. You know, number 15 is GPT-5.2. And it shows that actually open source has come back. And GPT-5.2 kind of shocked the world that, you know, I guess open source has not died. And, you know, DeepSeek, yeah.

So in general, this worked very well. You know, DeepSeek, you know, GLM 5.2 showed that, you know, open source does very well still.

You can also filter out by country. So by country, you can see that the black line is United States, you know, the US models. The dark red line is the Chinese labs. And, you know, there's other labs as well.

You know, French, South Korean labs and stuff like that. But, you know, over time, it shows that these models, you know, the US labs seem to do very well over time. You know, they're always at the frontier. And then the Chinese labs like to catch up over time.

Previously, you know, I mentioned, you know, the, you know, the plateau before, you know, before 0.1 preview got released. If you actually look at this plot, there is something called the open source draft. So after 0.1 preview got released, open source labs did not know how to replicate 0.1 preview.

They have never, you know, they don't know what is reasoning. So I'm not sure if you, okay, this is a few years back. But on Twitter, you know, OpenAI kept talking about, oh, you know, 0.1 preview was extremely powerful.

You know, every single tweet you see every single day, you know, they show that 0.1 preview was very powerful. And so for 1, I think it was 6 months to 8 months, open source models, open source labs, they got confused on what to do next.

But then, as everyone knows, DeepSeek R1 came along. And they showed that even for open source models, you can train these models to do reasoning, GRPO, reinforcement learning, and it does very, very well. In fact, if you take this plot, you know, the black line minus the blue line, if you just minus it, you get this plot.

And you can see this is how many months behind open source is. And, you know, over time, you can see, like, you know, after 0.1 preview got released, you know, it kind of skyrocketed. You know, the open source models were very, very lagging in terms of, you know, behind closed source models.

And so, like, when DeepSeek R1 got released, then the open source labs knew, okay, we can also do 0.1 type reasoning. And that is why recently, you know, the time between closed source labs and open source labs have started decreasing again.

Yeah. So this is slightly outdated. This is like May. So I think now it's actually 4 months with the release of GLM 5.2. It's around 4 months now. So open source labs lag behind closed source labs by around 4 months.

There is actually a very nice plot, you know, doing some sort of regression. So some sort of, like, trend extrapolation. According to this plot, if you extrapolate the trend, by December this year, open source models will 100% catch up to closed source models by this year, December.

But, you know, who knows, I guess. Maybe, maybe open source, maybe we can have an open source model as powerful as the best closed source model by December, you know, if this trend continues. So I guess the question is, will the trend continue?

It's always about, will the trend continue?

And, you know, maybe most of you maybe may know that, you know, open source, some of the open source improvements in technology, you know, improvements in capabilities are via distillation. You know, so some of the open source labs, what they like to do is they like to core the models, you know, core the frontier models like Opus or GPT, and then use the traces to train your model.

So this is a common methodology that labs like to do. I wouldn't say this is a bad method, but it is a method that, you know, some closed source labs like to look down upon. You know, they like to stop, you know, their view is, you know, we should not allow these open source labs to, like, do this training and, you know, get away for free, I guess, in terms of training cost.

But you don't actually have to do this approach. So most labs, when you do distillation, there are two different types of approaches. The first approach is you need to have the logits. You need to actually have access to the full logits.

And unfortunately, most labs do not actually have that. Right? So, like, labs will not give you the full logits. Instead, you only get the reasoning traces that are summarized and the final output. And so these, you know, these open source labs, they're not just, you know, they're not just training on the, you know, Opus output.

Right? That's just, that's silly. What they do is they use GRPO or reinforcement learning to recreate the traces. And so because you have the final output, which is the answer, all you need to do is use GRPO and RL to create the reasoning trace automatically.

And so that's kind of how they train these models. And so you don't actually need to, like, access the logits or the weights of the model. That's not necessary. Yeah.

And, you know, one of the most important factors of, you know, these large models is, you know, as models get bigger and bigger and bigger, you can't run them on your local device anymore. It's extremely complicated to run.

And so we do something called dynamic quantization, where essentially you take a model, you quantize them down to 1 bit. But the trick is you don't quantize every single layer to 1 bit. You quantize some important layers to 16 bit or 8 bit or something like that.

And so if you quantize the whole model down to 1 bit, you will get 0% accuracy. Right? 0%. But the trick is if you do dynamic quantization, so if you look on the, you know, this is a 3-bit DeepSeek model, a 3-bit one, you get 75.6% accuracy, a 3-bit one.

In fact, if you do dynamic 1 bit, you get 57% accuracy. So we show that, you know, if you do something called dynamic quantization, where you quantize the model down smartly, you can recover accuracy. And this methodology will become even more important when models get larger and larger and larger and larger.

If you plot the Pareto, you know, efficiency, there, if you don't do dynamic, if you do, you know, some other dynamic quantization methods, it does okay. But we show that if you smartly choose the layers, it does even better.

I'm not sure if you folks have followed, but GLM 5.2, we also released dynamic quantizations for that. We show that GLM 5.2 can quantize very well. So if you look, I think this is, oh, this is an animation.

Oh, it works. But yes, you can show the animation. You know, you can see the animation, a 1-bit GLM 5.2 model. This is 1-bit. And the 1-bit model is literally 86% smaller. So it's 86% smaller than the full 1.5 terabytes.

And it still managed to do very well on one of the prompts. So it shows that the models are not dumb. Right? If you make the model 86% smaller, it does not get 86% dumber. It only gets 14% less dumb.

And so it shows that, you know, if you do special tricks to compress the model, the model still works very well. And we also compare to Opus, you know, we compare to Opus 4.8, we compare to GPT-5.5. And also, you have to notice that for GLM 5.2, I use high reasoning mode.

You know, for Opus, it's extra high. And for, you know, GPT-5.5, it's also extra high. And so, like, you know, there are different reasoning modes as well, which we can also see. And all of these are one shot.

So we do not, like, prompt the model, like, you know, 50 times or something. This is just one shot directly.

Okay. So the next, I guess, the open source versus closed source section is done, I guess. Any other questions? Yes. So the question was, which parts of the model do we quantize to lower bits versus higher precision? So in general, we did actually a lot of research on this.

So if you look at the Qwen, the Qwen 3.5 architecture, there are some layers, which are the linear attention layers. The linear attention layers should never be quantized. If you quantize the linear attention layers down, you will definitely suffer in long context.

So in general, the linear attention layers need to be left in 8-bit or 16-bit. That's for example. Another, like, if you look at the model layers, some layers can be quantized down heavily to 1-bit. And the reason why is because these layers are kind of like filler layers.

And so they don't actually do anything. And in order to check whether a layer does something or not, you do need some sort of collaboration dataset. So you need to have some sort of, like, representative data and pass it into the model.

And you can get, you can get the outputs after each layer. And then you can see, okay, does this model at this specific layer, you know, does it change that much? And if it doesn't change that much, okay, maybe just quantize the layer to 1-bit.

But if it does change dramatically, then you need to be careful. You cannot quantize that down to, like, 1-bit or whatever. So there are actually many, we actually published a lot of, like, blogs research on this. We show, I think there was, we also show, for example, you cannot quantize the vision layers down.

If you quantize the vision layers down, you will make the model really bad. If you give it a, you know, if you give it a picture of a train, it will say it looks like a beach, for example.

And so it's, you should never quantize the vision layers, the audio layers, and only the language, the language model layers you can, like, quantize. But there are many tricks in order to do that. Yeah.

Guest32:40

Other than normal computing is?

Daniel Han32:41

Correct. So the question was, if you do distillation, you might have done worse on other topics. But, you know, only if you, for example, if you just do coding, it will just do good in coding. And then the rest gets very dumb.

So that's a fair point. So I think that the main trick is you will need to do many, many, many examples. You will core the model, like, you know, 10 million times. And so, like, the trick is once you core the model 10 million times with high diversity of questions, in general, by using the pre-training argument, the model will do well on other tasks.

So the reason why pre-training does very well is because it has learned so many tasks that it can interpolate the missing holes. For example, if you just, if you just pre-train a model with just maths questions, assume you do only maths, okay, maybe it's not going to do very well.

Right? And it's not going to do very well on every other task. But the trick of pre-training is it does maths, coding, law, you know, every single, you know, every single topic you can imagine. And the trick is because it has so much knowledge, it fills the holes of the things that it doesn't know.

And so for distillation, you also need to do the same approach. You need to sample, you need to sample well. So, for example, instead of doing 10 trillion tokens, sample, you know, like, 1% and then core the model.

Yeah. So that's kind of how the labs are doing that. That is a very good question. So instead of doing one big quantization, can you instead prune the model, like, you know, delete some layers entirely? So in general, from our research, pruning does work.

There is a very big problem, though. You need to retrain the model. You need to continuously train the model after pruning because you have deleted an entire layer. And so if you delete an entire layer, you will need to do, like, you know, QAT or further fine-tuning to push the other, to push the other weights to have more knowledge.

So that is the only problem if you delete layers. If you don't delete layers, when you do, you know, dynamic quantization, it's called post-training quantization. So PTQ. You do not need to do any training at all if you do, you know, quantization.

But if you do prune the layers, you do need to train. So that is one of the problems. Yeah. Yes. That's a great question. So the question was, you know, because open source labs, you know, they use closed source models, the gap will never actually go to zero.

And so I partially agree. And so the main argument was, labs, open source labs, the easiest way is to do distillation. However, you know, if you, for example, if you were an open source lab, you will only use that approach to firstly enter the market.

But as long-term, you know, as long-term safety, as a long-term safety net, you will not do this approach. Instead, as a, you know, instead you will do, you know, for example, generate the answer, get the question, for example, you know, you will get data from Merkle or Scale or whatever, you know, have some sort of, like, large data labeling army or something.

I don't know. And so, like, in general, because currently some of the labs, so they don't just do distillation. Right? So they're not just going to core the model 10 trillion times, you know, and just do distillation. They also augment the training data with their own approach.

So I will be talking about the GLM approach maybe, like, later. But they did invent some new approaches to do very good reinforcement learning and GRPO. And because GRPO and reinforcement learning, you know, is open source, these labs just use these methodologies to make the models better.

So you don't, so distillation is only one part of the training system. And it's not, I would say that assume distillation disappeared. Okay. Maybe open source labs maybe increase, you know, it's not four months, maybe eight months. But that's fine because, you know, we always have some sort of innovative new approach.

You know, DeepSeek might invent something new. And so, like, you know, GLM, KEMI, all of them, Google, you know, even the American open source labs, they'll have some new innovation. And so, like, I think, like, yes, if you stop distillation, it will increase, you know, the, you know, four months to eight months.

But I still think that is fine. It's just a delay, you know, and then the delay will go back to, like, four months. Yeah. Yes. Good question. So the question is, if dynamic quantization is always better, why do people not always do dynamic quantization?

So it depends on the definition of dynamic quantization. So for every single lab, they will have different approaches to dynamic quantization. In fact, I'm actually going to talk about that. I was going to talk about that in the benchmarking and accuracy minimizing section.

So I'll be talking about that. So I will, your question will be answered later. Yes. Okay. One more question. Yes. Yes. So the question was for consumer-grade GPUs, you know, what are the open source models in terms of, like, you know, the parameter size, capabilities, and stuff like that?

So for the open source community, you know, the most popular models are probably Qwen 3.6, 35 billion, 27 billion, Gemma, you know, Gemma's 26 billion, GLM 4.7 Flash, the smallest type models. And I feel like these small models are actually very powerful.

So, okay, I don't have, wait, I don't think I have a plot. But essentially, these small models, the biggest problem, oh, actually, I'm going to talk about this as well. The biggest problem of these small models are they fail very bad at tool calling because they have tool calling issues.

They loop continuously. And the biggest problem is because they're small. And that is why they have these problems. But we can counteract this. And so one of the things I'm going to talk about later is the model becomes not important anymore.

It's the harness or the tool that is actually the most important thing. How do you actually core the model? That actually affects the most accuracy of the model. So not actually the model itself. But I'll be talking about that as well.

Yeah. Okay. I will continue on. There will always be questions after each section.

Yes. Oh, yes. The next section, the fun section. Throughput maxing. Oh, actually, I think I did, it's supposed to be 2x. I don't know. Whatever. Throughput maxing and accuracy minimizing. I thought it was, like, accuracy minning, but there's no such thing.

Throughput Maxing38:50

Daniel Han39:02

So it's called accuracy minimizing for now. Yes.

So this part actually I really like. Okay. I'm not sure if you guys can see it. It's a bit, oh, whatever. This shows the pareto efficiency of cost, of the cost of the model. So cost is, cost is the x-axis.

And the y-axis is the Arena score. So this is, like, you know, LM Arena's Arena score. And this part I really like. So I don't really, you know, maybe you see, like, Arena's scores, you know, LM Arena's scores between each model.

I don't really like that. It's not, it's not very easy to see. Instead, the better approach is to plot every single model on two axes, cost versus accuracy. And you can see Fable does very well. Right? So Fable does very, very well on that plot.

But you can see there is a pareto trend, you know, like, Gemini 3.1 preview is over here. You know, Opus 4.6 is over there as well. There are some other models as well.

When Fable got released, okay, and, well, now it's banned. But anyways, when Fable was released, when, you know, when people tried it, they noticed that it's not that much better in terms of actual capabilities. You can see, you can see, but however, people really liked the front-end design.

You know, they said if you core Fable, it was very, very good for UI, UX, front-end. And in fact, if you look at the LM Arena's chart, you can see it was a very big shift in terms of front-end design.

GLM 5.2 is also there, if you can see. You know, it was part of the pareto trend. But in general, for these large models, they seem to have, they're not going to be doing, they're not going to be doing that much better on general tasks.

However, for UI and designing, Fable seems to have done very, very well. And so you should use Fable for your designing. You know, you should use Fable for designing, for UI, for UX, whatever, HTML, JavaScript. But you should probably not use Fable for the rest of the tasks because it is very expensive.

So, you know, use some other models instead.

And, you know, however, yes, okay, you know, some of the models, you know, like, okay, this, you know, this shows that Fable does very well on UI and UX. But how about over time? You know, how, what do, you know, Anthropic, their view is we need to maximize throughput.

Right? Maximize throughput, but also maximize accuracy. You know, they want to, like, you know, serve more people. But sometimes it doesn't actually work. Sometimes they actually reduce accuracy. And so you can see there is a, I don't know if you folks know Margin Labs.

They have this very cool, they do, they do Sweebench. They benchmark codecs. They benchmark codecs and Claude Code with the models. And this is accuracy over time for these models. And the dotted lines are the release of the new models.

So there's actually another, there was actually a dip in, wait, can you, is there, oh, okay, the math is there. I think it was over here. I think it was over here that Fable got released. So there was actually another dotted line.

There was actually very interesting trends you can see. The first one is every single time there is a new model release, this, this, you know, daily tracker seems to decrease in accuracy. And so if you want to predict when a model gets released from Anthropic, you can use this as an indicator of when the model gets released.

It works very, very well. Right? So, like, essentially, if you were over here, the dip in accuracy over a very long period of time was because Fable got released. And over here, I think that's Opus 4.8, I think.

I think, yeah, I think that's Opus 4.8. This is Opus 4.7 and so on. That's 4.6, I think. Whatever. I don't remember exactly, but, but you can also see that there are ginormous dips of accuracy. And it's not just, like, one day or two days.

It's for a very long period of time.

This is also codecs. So they also do codecs benchmarks. And you can also see that over time. I don't know if you can squint, but you can see that actually codecs have been getting worse if you plot the trend.

Right? If you can, I don't know if you can squint, but if you draw a line, it seems to be getting worse. So I'm assuming OpenAI is investigating this as well. Okay.

Guest43:33

So this is a different model. So this is like.

Daniel Han43:35

This is codecs.

Guest43:35

Codecs is doing worse than 5.4.

Daniel Han43:38

So this is using 5.5. This is using.

Guest43:40

It's the same model every day.

Daniel Han43:44

Correct. It's the same. So what this benchmark does is you randomly sample 50 Sweebench questions. Sweebench is very large. So you just sample 50 of them and then you core the model to answer it. And then you record accuracy.

And so obviously, you know, every single day there's like, you know, daily variations. It's not, it's not that useful because you're only calling 50 questions. So the trick is to look at the trend. And the trend, oh, maybe OpenAI should investigate this.

And you can see the trend for Claude, you know, Anthropic is also not very good. In general, sorry, this is not the same model. These models change. My bad. So it's the same harness, but the model changes. So this dotted line is GPT 5.5.

So everything over here is GPT 5.5. Everything over here is GPT 5.4. I think this is 5.3 and so on. But it seems like the model's getting worse. So I don't know. This is probably just on this benchmark.

Right? On the Sweebench Pro benchmark, it's getting worse. But, you know, I wouldn't really trust these benchmarks. The best way is to look at the degradation, you know, the sudden drops. You know, for example, codecs dramatically dropped over here.

I don't know why. And, you know, Claude, you know, Claude Code was very bad for a few weeks over here or over here. Right?

Okay. Yes.

Guest45:10

Is there a confidence interval? Is it a one-time?

Daniel Han45:12

Yes. There is a confidence interval. I did not plot it, but this is 50 tasks. So every single day, they core 50 tasks randomly. So they will sample 50 tasks. And so you should not look at this daily.

This is daily. So every single day is 50 questions, another 50 questions, another 50 questions, and so on. Instead, you should do, like, a rolling average. You know, some sort of rolling, you know, seven-day average. That's a better number.

Yeah.

Guest45:38

I don't see any data there.

Daniel Han45:41

Really? I can see it from here. It's, like, decreasing.

Guest45:45

I see that.

Daniel Han45:46

It's, it's.

Guest45:47

It's detectable.

Daniel Han45:53

If you look at this, if you do the seven moving average, I'll probably get the plot later. It actually is decreasing. You can see it. If you can see, I don't know. If you look at the top peaks of the, you look at the top peaks and the peaks are decreasing.

Guest46:06

No, it's just random. So you have the top peaks and then it goes up.

Daniel Han46:10

Okay. How about the bottom peaks?

Guest46:13

You can only count down. You can have a, there's a gap multiple and it has a random noise on it.

Daniel Han46:18

Okay. I agree. There is random noise. So the trick is you need to do the moving average. And if you look at the moving average, you can actually see it's decreasing. I'll probably, I'll get the plot later. You can, you can search it.

It's, so go to Margin Labs. Search in Margin Labs codecs, Claude Code benchmarks. And they do show weekly, the weekly trend. But I'm just saying this is not, this is not to say that the model's getting worse. This is just to show that accuracy, that, you know, the sudden dips, the accuracy of these models can decrease.

And the question is why. You know, for example, why did Claude Code over a few weeks, why did the performance decrease? Like, why? That's the fundamental question.

Guest46:59

Did it change the weights on those dips? Did it change the weights?

Daniel Han47:04

So that is one theory. A theory is they might have accidentally, you know, that before the model release, they are doing testing. And so they might have, like, you know, you know, some of the, some of the queries they route to Opus 4.8 or Fable or whatever.

And the problem is they did not, so the main question is if you do route to another model, why did the accuracy decrease? It should actually get better. And so one of the theories is, theory one, they forgot to edit the system prompt.

And so the system prompt for Fable was different, but then they used the wrong system prompt for, you know, for Opus 4.8. And that is why the accuracy decreased. And then after the model got released, the accuracy went back up because they used the correct system prompt.

That is one theory. The other theory is, the other theory is, okay, we're actually going to talk about this, is it's actually they're doing tricks. You know, they did quantization, but they didn't do dynamic quantization. They did some dumb quantization.

You know, they, some GPUs are broken, for example. You know, they use the wrong GPUs. Some of them have, like, you know, bit flips or something. I don't know. They have, like, a new data center. And then that data center, just by chance, has lower accuracy.

In fact, there is actually, okay, I'm going to talk about this actually.

Yeah, but there are many, many theories, like, you know, possibilities why this could reduce inaccuracy. Actually, I, I think it's the next plot. Yes, the next plot. Oh, well, the next slides. So actually, when was this? I don't remember.

It was a few months ago. Someone from AMD actually made an issue on Claude Code, you know, during this dip. I think it was during the, before a very large dips in accuracy. And they actually asked Claude, you know, they asked the Claude team, why is there a noticeable dip in accuracy?

You know, why, why is that? And Claude actually wrote a, in April 23, they actually provided details on why they had reduced inaccuracy. Right? So they did a postmortem on what happened with Claude. And the reason why is because the thinking trace got deleted after the second, you know, when you, when you asked Claude the second time, the thinking trace got deleted.

And it had a bad system prompt. And they found out that that, that was why the accuracy got reduced. So somehow in Claude Code, the second time you ask a question, the previous thinking trace got erased. And I don't know, I don't even know how they did not find this, but oh well.

According to them now, is Claude now has this internal benchmark. So they will use more internal investigations to test, okay, next time if there's a new model, this won't happen ever again. And, you know, like, these things do happen over time.

And so, like, for this specific example, Claude, you know, Claude Code, the harness, the harness itself was the problem, not the actual model. Right? The harness, the thinking trace got deleted and they had a very, not a very good system prompt.

And that is why the accuracy actually degraded. So that, okay. So we found one answer why these models got worse.

They also released in September 2025. Right? In September 2025, they showed that it was due to, okay, I didn't, okay, I didn't put the slide. But anyways, they showed it was actually due to a hardware problem. So in their compiler, they used TPUs.

So, so Anthropic likes to use TPUs and GPUs. They showed that the same software stack for GPUs and TPUs actually produced different results. And so for the TPUs, it actually was different sampling. And for the GPUs, it was a different sampling mechanism.

And so that is actually why they had decreased inaccuracy during September sometime because they actually had different hardware. And so you need to, like, yeah, so, like, once you have different hardware, accuracy also changes.

So I think the main point is the harness, the implementation, the tool is now the most important. It's not the model. Right? The model is useless. Most models, you know, if you look at the model of, you know, open source versus closed source, models are generally the same.

The difference is how Claude Code is made, you know, how codecs are made and used. And so that is actually the most important factor. It's not the model anymore. And so, like, you know, as we have seen, if they have accidentally botched, you know, if they accidentally botched the harness, you will get reduced accuracy.

And so, like, you know, definitely, you know, for large labs, as, you know, I'm sure they are, they know these problems and they're working on it. But I feel like, you know, these are probably, you know, these are still very hard to fix.

Yeah. So hopefully I answered some of people's questions on the harness, you know, the accuracy. So there was actually reasons why accuracy got degraded.

But, you know, it's not just closed source labs doing bad. Across open source model providers, the accuracy changes. So if you look at this plot, and so this is from OpenRouter. This is DeepSeek V4 Pro. So most labs, what they want to do, most inference providers, what they want to do is they want to serve you the highest throughput,right, with the cheapest price.

They want to give you, you know, 60 tokens, 120 tokens, 1,000 tokens per second. Right? They want to give you the fastest. But did you, but did people actually bother to check accuracy? So that is the fundamental question.

You know, you might be getting 10,000 tokens per second and there is no model. So the main question is you need to be careful of what you use from these inference providers. And so for DeepSeek V4, you know, there are two benchmarks which OpenRouter ran, you know, yeah, it's like sorted, it's sorted, I think, on the gray, I think it's sorted on TailBench.

So it's sorted on TailBench. And the green one is GPQA. And you can see that in general, some of the labs are not, you know, some of the, sorry, not labs, some of the inference providers are not doing very well.

So you need to, like, before you, before you use an open source model, please check the accuracy before you use the open source model. And also, one of the biggest problems of this is every single time, you know, for example, like Claude Code, you know, Claude Code and codecs, you can benchmark accuracy over time.

And the good thing about closed source labs is they control the supply chain. The biggest problem of open source is there are so many suppliers and providers of these models that sometimes what happens is people get turned off and they get very annoyed that the open source models do not work very well.

So everyone, you know, in the, in the ecosystem, people keep saying that closed source labs do much better than open source. But it's not because of the model. It's because of the inference provider. Right? The inference provider is to blame that they are causing the downfall of open source because they're giving a bad name for open source.

So I would, like, check, you know, whatever favorite inference provider you have. So this, this benchmark was run, I think, yesterday by OpenRouter. So this is, this is daily data by OpenRouter. So whatever favorite inference provider you have, please tell them not to, you know, reduce accuracy that much.

This is GLM 5.2. So, you know, GLM 5.2 as well shows different accuracies. You can see, so the plot on theright shows most model, you know, most inference provider, okay, I keep saying model labs. Most inference providers are throughput maxing, but they are accuracy minimizing.

That's where the phrase comes from. Okay? So they do not care about, in fact, look, like, you know, the highest accuracy is 76.4% and the lowest is 62.4%. So there is a 10% gap between the, you know, between the highest accuracy and the, you know, lowest accuracy.

And so, like, you need to, you know, as a, as a, you know, as a callout to inference providers, you know, please increase accuracy, you know, before trying to make things faster. Right? You do not want a model to be very dumb.

And it's like, you know, 10,000 tokens per second. Right? We can make it 1 million tokens per second and there is no model. You know, just call a human or something, you know, make a fake or something. So, yeah, so the main point is we need inference providers to do good in terms of accuracy.

Otherwise, this will make open source have a very bad look. Yeah. Oh, okay. That's the end of the, the second section. I guess that was a bit of a rant. Any other questions for this? Yes.

Guest55:54

Yeah. This is for a new organization that, that wants to use, like, an open source model. Do you suggest using a, you know, inference service provider or do you suggest downloading from Hugging Face and then using, like, Moodle or, you know, some kind of server to, you know, implement yourself?

Like, what do you suggest if any new organization comes and asks you, like, how to use open source model?

Daniel Han56:19

That's a great question. So when an open source model gets released, you know, how should you use it in terms of accuracy, throughput, or whatever? So in general, in general, open source has come a long way. So for example, we did report bugs in Gemma 1, Gemma 2, Llama, Mistral, you know, OpenAI's GPT-2.

Every single one of those models had bugs. And so the good thing is, you know, as Unsloth, we will help the labs before they release a model to fix some of the issues. So every single model you now have has some of our fixes.

So that's a good thing. But in general, if you have an open source model, I would use Llama CPP, for example. I think Llama CPP and Llama Server is probably the most bug-free system. So I would, like, suggest, yes, you should download from Hugging Face, use Llama Server, use Llama, you know, CLI.

I don't know, you can use Unsloth Studio, whatever, whatever is your favorite tool, but you should, yes, you should download from Hugging Face. In terms of, like, you know, if you're a large enterprise, generally speaking, what they like to do is they like to wait one week.

So most enterprises, they'll wait one week for all the problems to be fixed. And then, you know, then they will use the model. But in my view, that is not a good approach. I would say if you, okay, if everyone waits one week, then, like, how do we fix the bugs?

Because only at scale, only at scale, then we can see the bugs. And so, like, in general, we need everyone to start trying these models earlier and not, like, you know, wait one week, wait one month. You know, don't do, don't do the waiting approach.

But I would say, like, in general, the enterprises, what they like to do is just wait, wait one week. Yeah, that's like common practice. Yes.

Guest58:04

Would you mind, can you go over the theories and reasons again, like, a bit more about why the performance degrades before a model really issues related to what Anthropic?

Daniel Han58:15

So, okay, the question was why would the model performance degrade before a model release?

These are just hypothetical questions, hypothetical theories. So every single model has a different system prompt. So Opus 4.8, Opus 4.8 system prompt is very short, but Opus 4.7 system prompt was extremely long. So the theory was, this is just a theory, that Anthropic via Claude Code accidentally routed some of the models to Opus 4.8.

Right? They use Opus 4.8 as testing. Right? They need to test Opus 4.8, but they used Opus 4.7 system prompt. So they used the wrong system prompt, and that is why accuracy degraded. That's one theory. Another theory is, actually, I think that's the, actually, I thought about it.

That's probably the only theory I had. I'm, like, thinking, hmm, is there another theory?

I guess the harness itself, like, you know, sometimes the harness itself, the harness was designed for Opus 4.7. And during, when they were going to release 4.8, they need to collaborate the harness. Right? They need to change the harness for 4.8 to make it work.

But the problem is you're not allowed to publish it. Right? You're not allowed to publish it and give it to people because otherwise people will, like, you know, go on Twitter, on LinkedIn, you know, everywhere. Ooh, I can see Opus 4.8 is going to be released.

You know, everyone's going to be screaming, you know, 4.8's coming, 4.8's, you know, getting released. And so maybe that's why accuracy decreased. It's, they update, they did not update the harness. Or the other option is they, they already updated silently, they silently updated the harness before the new model got released and it regressed, you know, it reduced accuracy.

I don't know. Like, to be honest, I, you should probably ask Anthropic that question. Or, but I think in general, in general, the dips, the dips don't always correspond to, like, new model releases. Some of the dips are actual issues.

Like, you know, the thinking trace got deleted, the system prompt they wrote it wrong. I think for the system, it's funny. I think for the system prompt, they said

they tried to reduce verbosity, so they tried to make the model less talkative, and it actually made the model dumber. And so I think it was just one word. They added one word, no, one sentence, I think, one sentence in the system prompt that made the model dumber.

Yeah. I don't know if that helps, but I don't know if anyone else has any, like, theory. I don't, I don't think so anyone even has that many theories on this. Obviously, the Anthropic engineers will know, but I, you know, they're not going to tell.

So it's just based on hypotheticals. Something to do with the system prompt, something to do with the harness. Yeah. But I think in general, you can also use this plot. You know, if the performance decreases, most likely a new model's going to be coming.

Yeah. Any other questions? Yes.

Guest1:01:12

Just to add on that, OpenAI, Anthropic actually released system prompts?

Daniel Han1:01:19

Yes, correct.

Guest1:01:20

So I was reading into it, like, they used to have a lot of guardrails, and when new models get released, they drop the guardrails.

Daniel Han1:01:26

Correct.

Guest1:01:27

And then they again add new guardrails because they figured out that they are being exploited. So that's.

Daniel Han1:01:33

Yes, exactly. Exactly. So before a model release, they use a different system prompt for that new model, for the old model. And so that is probably why there is some decrease in accuracy. They switch the system prompts around or something like that.

And also, you know, the model itself, you know, I think 4.8 system prompt is very short. It's, yeah, I think it's like very, very short. And 4.7 was ginormous. And the reason is 4.7 was like, you know, I don't know what, I don't know what happened, but they have this ginormous system prompt and the 4.8 just shrunk it a lot.

So maybe, maybe they used the 4.7 system prompt, I don't know, or 4.8 system, the short system prompt for 4.8, and then they used it for 4.7. And that's why it decreased accuracy. I don't know. But yeah, you're correct.

They do release, although I think the system prompt they released on the website is for Claude.ai. So the online chat system, the Claude Code system prompt is actually different. So I think you need to actually call, you need to call Claude Code, you know, what is my system prompt?

And then you print it to like a text file, and then you can, like, investigate what the system prompt is. And then you can also override it if you want. Yes, but it's a different system prompt most likely.

Yeah. Last question, if anyone. No? Okay. Continue on then.

Okay. The next section we are going to be talking about is bench maxing and cheating.

Bench Maxing1:02:59

Daniel Han1:03:07

I'm not sure if you folks have seen the DeepSpeed benchmark. The DeepSpeed benchmark is a very popular recent benchmark that shows, you know, the cost is on the X-axis and the Y-axis. It is a DeepSpeed benchmark. It's a new benchmark based on, like, you know, a better uncontaminated version of Speed Bench Pro.

And in general, you can see that, you know, GPT-5.5 does very well with Fable, you know, GLM, Opus 4.8 in general. Right? It shows, you know, this plot shows that models are getting, you know, these, the dots are different reasoning modes.

I think this is maximum reasoning, I think, high, extra high, you know, these are actually different reasoning, reasoning times as well. But in general, you can see that there is a pareto efficiency trend. Right? The best model is the one, you know, to theright, to the top.

Right? The better the model to theright, to the top is the better the model. So you want models to do better and better over time to that, to the topright corner.

And, you know, I just learned, I didn't actually know this. I just learned that Speed Bench Pro, when you run this benchmark, you use LLM, you use language models as the verifier. And I was, like, confused because, like, for most benchmarks, for most benchmarks, you should never call another language model to check whether your answer isright or wrong.

And so for Speed Bench Pro, you actually call a language model to verify if your language model wasright. And so that is why Speed Bench Pro is not a very good benchmark. One of the problems is, is do we need to do sampling?

Like, how many verification runs do you need to run to verify if your answer is correct? Do you run it one time? Do you run it five times? Do you run it 100 times and take, like, an average?

So, like, I was actually quite shocked that this is actually what happens. I was quite surprised, actually. The next question is, which model is the verifier? You know, you ask, for example, you ask Opus, you know, you benchmark Opus 4.8 on Speed Bench Pro.

But which, what do you use as the verifier? Do you use Opus 4.8 as the verifier? So you're using the same model itself to verify itself. And so, like, this, I was, like, quite surprised, actually, that this is how benchmarks work and actually quite disappointed.

But anyways, obviously you can go with the other approach. You can do human verification. You know, everyone in the room, I'll give you the Speed Bench, you know, and just tell you guys to verify it. You could do that, I guess.

And also, what happens if the verification changes every day? You know, remember previously models, you know, every single day models get better or worse. What happens, what happens if you run, what happens if you run the verification when the model was doing very bad?

Right? You will actually have different Speed Bench numbers. And so, like, I'm actually quite surprised this is what the industry does. You know, run Speed Bench Pro, but using LLMs as verifiers. That is definitely not a good idea.

But anyways, people do it, whatever. In fact, according to DeepSpeed, if you do, if you do verification using language models, Speed Bench Pro has an 8.5% false positive rate. And a false positive rate means that the LLM verifier said that the model was correct, but it was actually wrong.

And so 8.5% of the time it will do this. The false negative rate is even worse at 24%. This means that the verifier said that the model was wrong, but it was actuallyright. And so you can see that Speed Bench Pro is a very bad benchmark.

And so DeepSpeed showed that they have, you know, they fixed the problem, you know, by reducing the false positive rate and the false negative rate to, you know, 1%.

In fact, some examples of cheating, I, you know, this is actually quite surprising, but in the Speed Bench Pro benchmark, you get, you get like a GitHub question, you know, a GitHub issue. You call the model to solve that GitHub issue.

But did you know that in Speed Bench Pro, you get the full Git history? So you get the, you get the actual answer as well. So I'm, like, I'm actually quite, I was actually quite shocked to learn this, that during these models, you give the answer and the question.

Like, obviously the model will cheat. And so, like, this is definitely a very bad benchmark. You know, you should never, ever, ever, ever give the model the answer. And so very silly, but yes, this happens a lot. And you do not want the model to literally see the solution.

Right? That is a terrible approach. The other problems that you get, like false positives, is, you know, the PR tests, you know, the GitHub, the GitHub issue tests are very weak. So, you know, at the final conclusion, you know, when the GitHub, when the GitHub issue is closed with a pull request, the tests that the maintainer wrote are not very good.

And so the problem of that is, you know, if you have tests which are very weak, then, you know, the model does very well, not very good. And obviously the worst part is the model will, like, bypass some tests.

It will skip some. And that is not a very good approach. In fact, DeepSpeed actually showed how many times a model cheats by looking at the full Git history, you know, directly going to the answer. You can see Opus 4.7.

So the purple bars show cheating by models. Ah, it looks like GPT-5.5 never cheats. It looks like it. Ah, okay. Maybe we should use GPT-5.5. You know, actually, this is actually very interesting. There are some people which think that if you cheat, that's actually good.

And the reason why it's good is it means that Opus 4.7 already knows, like, if you give it the full Git history, you should be able to, like, you gave it to them,right? You gave Opus the full Git history.

It should find the solution there. Right? It should just directly skip over to the solution. So it's, that's what people think. You know, people have a view that the humans gave Opus 4.7 the full Git history. So it should cheat.

Right? You, you, you designed it to cheat. So in general, Claude models seem to cheat more. And OpenAI models seem to cheat less in general. So it depends on you, you know, if you want a model to cheat or not.

And the definition of the word cheat is also very, you know, charged. So I guess it depends on what the word cheating means.

You know, for false negatives, remember, Speed Bench Pro calls a language model to verify if your answer is correct. And so sometimes it's not very good. You know, sometimes you have unrelated tests that fail. You forgot, you know, sometimes when you write tests, you forgot about the tests which have helpers, you know, helper functions, and you just skip that.

So there are many issues. And this, I think this was 20, yeah. So 24% of the time, 24% of the time, the model says, the verifier says your model was wrong, but it was actually correct. So this is another problem.

And even worse, the harness itself can change accuracy. So when you benchmark using Speed Bench Pro, like, you need to have one agent or one harness for all models. Right? How do you create a generalized control environment for these models?

And so you can see, like, you know, for example, DeepSpeed showed if you use Claude Code, you get 40% accuracy, but then if you use their own, so it's a special harness, you can get 50% accuracy. Gemini, for example,right?

If you use Gemini CLI, you get 20% accuracy, but if you use their one, you know, the control environment, you can get 40% accuracy. And so in general, for these benchmarks, you also need to have a controlled environment.

And that is also another problem. And with DeepSpeed, they showed by using this benchmark, by solving, you know, by stopping cheating, you know, by, you know, if we remove cheating, if we remove, you know, these other issues, you can see the models, you know, the models are not saturated anymore.

Right? You can see the models are very different in terms of the capabilities. According to this benchmark, GPT-5.5 is the best, according to this one. I didn't, oh, this is not updated. 4.8, I think, is over here or something.

But yes, this benchmark shows Claude Haiku is 0% accuracy. Right? It's terrible, I guess. But yeah, this benchmark just showed, okay, the main question is, do you trust this benchmark? That is another question.

There are other benchmarks. Right? So Cognition released a frontier code benchmark, which also tries to solve the same questions for benchmark, you know, for cheating and benchmarks. And what they showed is you can fix contamination. And how do you fix contamination?

You ask, you know, you ask Cognition's team, which is full of, like, you know, national Olympiads and, you know, international Olympiads. They manually checked every single question themselves, you know, and removed bad questions, you know, bad examples. And they also showed that their questions are much more diverse.

Right? So frontier code has many different other languages. And they showed with diversity, you know, with more diverse programming languages and by reducing contamination, they also have a benchmark. And according to their benchmark, Opus 4.8 is the best.

Right? With 14.5% accuracy, GPT-5.5 is 7.2 accuracy. And this is the diamond one. Right? So this is the 50, the 50 hardest questions. The main benchmark is 100 questions and the extended is 150. And so according to them, you know, Claude does the best, according to them.

But also according to them, frontier code seems to be better than DeepSpeed. Right? The benchmark that I showed previously, DeepSpeed, this one, you know, according to frontier code, so the Cognition team, their benchmark is better than DeepSpeed. Right?

According to them, according to them, DeepSpeed's false positive rate is 44.9%. But remember, what did DeepSpeed say? They said the false positive rate was, I don't remember. What, what, what did they say? They said that it was 0.3%.

Right? So DeepSpeed said, DeepSpeed said their false positive rate is 0.3%, but frontier code said that DeepSpeed's false positive rate was 44.9%. So, you know, there is some competition, I guess, between benchmarking labs. Well, Cognition is not a benchmarking lab, but, like, you know, between companies.

So the main question is, who do we trust? You know, do we trust frontier codes benchmarks? Do we trust DeepSpeed's benchmarks? Do we trust Speed Bench? You know, who do we trust? And that is a very important question.

You know, my take is, like, you know, I guess just take an average of everyone. Take an average of everyone and you'll probably get the best answer. You know, who is actually doing the best. Yeah. But this is actually very interesting.

You know, it showed, okay, so according to them, the false negative rate for DeepSpeed is correct, you know, 1.2%. But my interest, you know, I probably, you know, my main question is, why is the false positive rate so high for DeepSpeed?

According to, according to frontier bench, DeepSpeed is even worse than Speed Bench Pro. That's what they're trying to say, I guess, for the false positive rate. Yeah.

And even worse, there is another benchmark called Frontier Math. So Frontier Math is by Epoch AI. So they, they have this math benchmark with different tiers, you know, tier one, tier two, tier three, tier four. So tier four is the hardest.

But the benchmark itself was botched. And so they actually had to release a corrected version of their benchmark. I think this was one month ago or something. So they showed that their benchmark questions were fully wrong. And you can see that if you correct the benchmark, if you correct the benchmark, the accuracy for GPT-5.5 jumps from 50% to 80% or something.

And so now you kind of entrust the benchmark. And they showed in a tweet, oh, it's June 12. Oh, it's only two weeks ago. So in June 12, they showed that the reason why they did bad on the benchmarks is they, they did the answer extraction incorrectly.

For example, they did, you know, they had unclear questions. They had the incorrect sign. So, for example, they said the model said 12, but it should be actually minus 12. And they forgot to get the minus sign. They have one-off errors.

Yeah. There's many problems with the benchmark. And so they fixed their benchmark just recently.

In fact, you know, it's actually quite funny. This was just two weeks ago. Have you guys heard of Hugging Face's Math Verifier, which was one year ago? And Hugging Face showed that, in fact, these benchmarks, when you do math questions, they always do bad.

And the reason why is because there's many problems. Right? The formatting is incorrect. You know, the extraction of the fraction is wrong. You know, the sign is failed extraction. There's many, many, many problems of mathematical extraction. And to be honest, I feel like it's like kind of reinventing the wheel or, you know, rediscovery.

But Hugging Face actually published this one year ago, and Epoch just fixed that two weeks ago. So, you know, benchmarking labs definitely need more, you know, they need to investigate literature more, I think. In fact, according to Hugging Face's Math Verifier, you know, if you use, if you the green bar, the green bar is if you do not use Hugging Face's verification system, you know, to fix the benchmark.

If you do fix the benchmark, you can see accuracy dramatically increases. Right? For example, for Quen, for Quen, the accuracy was 10%. Now it's 25%. And so you need to, so that means the open source models are not dumb.

They just have different, they output a different format. And so one of the problems is, how do we actually, actually, like, you know, pass these different formats? In fact, it's even worse. You know, I think I tweeted, oh, I tweeted this in August 2024, that if you, if you use different tokenization, you can also have different accuracy.

In fact, for MLOU, if you use spaces, you increase accuracy by 0.4%. It might not sound like a lot, but the point is, by these very dumb things, like, you know, using spaces or, you know, minus 12 becomes 12, and all of these, like, dumb little small things, the accuracy of these benchmarks can change over time.

And so, like, the main question is, you know, how do we make benchmarking labs and benchmarking companies, you know, how do we make them more reliable and, you know, more trustworthy?

Oh, okay. That's, I guess, the section for the benchmarking question part. Any other questions for that section? Questions? Yes.

Guest1:19:08

So I think with all these kind of, you know, issues with the benchmarks, like, do you have any specific suggestion on how to trust them? And, like, for example, if you are employing these models, how, how they should kind of, you know, build their own benchmarks?

Because it's clearly, like, you know, we're trusting some numbers, but some numbers probably, I mean, it is hard for everyone to go and build their benchmark in the end. So how, how do you suggest to kind of go about these?

Daniel Han1:19:41

That's a great question. So the question is, how do we, how can we trust these benchmarking companies? Or, like, what other types of benchmarks can we do to make it trustworthy? So that is actually a very good question.

The main question for benchmarks is you need to satisfy two conditions. The first condition is the benchmark must not, must not be benchmaxable. Right? How do you make a benchmark that is extremely hard to benchmax? Right? How do we, like, not get 100% accuracy?

And the second question is, how do we make the benchmark verifiable? Right? So how do we make the benchmark you can, you can also verify that the answer is, in fact, correct? Right? You remember, Speed Bench Pro is dumb because you call the language model itself to verify itself.

So that is not good. So the main question is those two questions. And so one good example, this is just a dumb example.

Randomly create math questions. Sample, for example, okay, this is, okay, this is probably not a good benchmark. You automatically create math questions. We can sample infinity. Right? We can sample infinite math questions. Right? Two plus two, four plus four, you know, any single number added together.

That's one question. Can you verify this? Yes, you can. Right? You can call a calculator to verify what is two plus two. Can this be benchmaxable? Hard. And the reason why hard is because the sampling space is infinity.

Right? It can be two plus two, 1000 plus 101. Right? It can, you don't have to do plus. Right? You can do 1000 times 1000. And so that's one way. Make a benchmark which is very hard to cheat, but also easy to verify.

So some sort of math question. The other one, for example, is, okay, maybe this is not a good example. I'm just making this one up on the spot. Tell the model to create a poem in 70 words, and you must use the word "happy."

Can you verify this? Yes, you can. Is "happy" in the, you know, generation? If yes, plus one. Also, you can count how many words. Right? You can count, okay, is there 70 words? So you can do these types of approaches.

And is this benchmaxable? No, it's, it's very hard to benchmax because you can say 70 words, 69 words, 68 words, 102 words, 1000 words. Right? It doesn't have to be "happy." It can be, you must have two words.

You must have three words. So some, some sort of benchmark where it's very hard to benchmax.

Yeah. In my view, I think that's, that's probably going to be the most important benchmark. And I don't, I don't think so anyone has actually made this yet. I don't know. Maybe someone in the audience or, you know, you guys can go as teams, I don't know, make a startup or something.

You know, do that. And I feel like that benchmark would be very, very important. Yeah. Yes.

Guest1:22:41

What's in your opinion, like, are the benchmarks we can trust today?

Daniel Han1:22:46

Benchmarks we can trust today? None of them. Take an average of all of them. To be honest, probably the best approach is just vibe, uh, vibe checking. Try all of them and see which one you like the best.

To be completely honest, I just, you know, like these benchmarks, ah, like, the main, the main issue I have with benchmarks is, for example, you know,

I mean, like this one,right? This one. I mean, even every single day, the benchmark can change. So we can't trust the benchmarks anymore. So my fundamental view is do not trust any benchmarks. Take an average. And then, okay, then the main question is, who's taking the average?

I guess artificial, artificial analysis has some average. The only problem is they have some weightings for the weight, you know, each benchmark has a weight. So now the question is, you know, what is the weighting of each benchmark?

You know, you can't just take like a dumb average. You know, you can't just say, you know, 10 benchmarks divided by 10. That's probably not going to work. So the main question is, how do you even do the weighting?

That's another problem. So I think in general, it's based on vibe checking, I guess. Yeah. I guess I don't have an answer for that. Any other questions? Yes.

Guest1:24:00

So the Speed Benchmark, when you're saying it's comparing against its own model,

there's one thing to say, all these benchmarks, even if the benchmark is scored by manually evaluating or deterministic validation, these are the 100%, these are the 100 answers should match to. But when it goes to complex reasoning and dependent, dependent reasoning, like, you know, the answer for this, this based on this answer as an explanation,right?

When deep reasoning goes on, it becomes very challenging to trust. Even if, if you say DeepSpeed or any of these benchmarks is better than Speed Benchmark, the question is, if can it, can it take open model, open source model and validate it against a faint, fragile model?

Daniel Han1:24:55

Yes.

Guest1:24:56

So that way, it's good or bad, doesn't matter, but it's relevant for me. Right?

Daniel Han1:25:04

You're correct. So the question was, in terms of, because Speed Bench Pro, for example, you call a model, the question is, what model? Could it be 4.8? Could it be GPT-5.5? And you call this model to verify the benchmark.

And so the question was, can you use an open source model instead? So then now you, you have a controlled environment. So yes, you can. But remember, there is a problem because even open source models itself have bugs.

Times, you know, the inference engines have bugs. Times, the inference providers have bugs and accuracy degradation. So it's, you're correct. So the main question is, we need to have someone or some organization, you know, some person or some whatever committee that we can investigate, you know, which engine did you use?

Do not update the engine. You know, the engine must be the same. You know, the weights must have not changed. So there's many, many, many problems with this approach. But I do agree, as open, you can use an open source model, but it's not, it doesn't solve the other problems.

Yeah. Does that? Okay. So the next section I'm going to be talking about is cybersecurity and regulation. This is an interesting topic. So I'm not sure if you have all folks have seen this plot. It shows the AI Security Institute, I think it's from the UK.

Cybersecurity1:26:08

Daniel Han1:26:25

They show the performance of models based on some cybersecurity task. And they show that Mithos Preview seems to be the best, you know, with GPT-5.5 Cybar, you know, Preview and so on. They show this benchmark.

And again, previously, as I mentioned, WeirdML is a better, in my view, okay, this is just my take. WeirdML is a better benchmark in general for benchmarking intelligence of models. And the reason why is because it doesn't actually, it doesn't actually follow the trend of reasoning versus non-reasoning.

Remember reasoning, reasoning previously, I think, okay, I don't have it. Reasoning, the reasoning models, doubling time reduced by half to 3.5 months. So remember, you just need to wait 3.5 months and the model's capabilities will double. And the non-reasoning was seven months.

So you need to wait seven months for the models to double in capability. But WeirdML did not actually have this trend. The WeirdML benchmark showed that actually the trend was like, there is no trend.

And I think I was just talking about this, you know, like, one of the biggest problems of benchmarks is you need to constantly reinvent yourself and do reweightings of combinations of benchmarks. For example, artificial analysis just recently released, you know, their new v4.1 benchmark, and they showed the weighting of the benchmarks.

You know, GDP Val is 20%, Terminal Bench is 16%, and so on. And so they, they designed these numbers as weightings for each of those benchmarks, and then they averaged it up together. So the main question is, how do you actually determine these numbers?

And so this is more like a human approach. You know, you have to determine these numbers. You know, ARC AGI kind of saturated on ARC AGI 1. And so that's why we have ARC AGI 2. And that is also why we have ARC AGI 3.

And my, you know, I guess once ARC AGI 3 is saturated, then we have ARC AGI 4, 5, 6, 7, whatever. And the main point is, once you have benchmarks, is it called Goddard's Law? I don't remember. The good, the benchmark itself becomes useless because, you know, models will start benchmarking on this.

So one of the biggest problems of these larger models for cybersecurity, for example, is Mithos actually dramatically went out of the trend. And that is why, you know, many people are afraid of these, you know, Mithos, you know, GPT-5.5.6.

You know, they're afraid of these models because it went out of trend. You can see that Mithos dramatically went out of trend. And, you know, even, you know, you know, GPT-5.6 didn't really release that many benchmarks because it was in preview mode.

So this is from their system card. They showed for cybersecurity that GPT-5.6 does very, very well. In fact, because GPT, because GPT-5.6, I think they only did Terminal Bench as their benchmark. They did not benchmark on anything else.

They did have in their system card, they did have one benchmark, which is very important. And this is called the internal research debugging evaluation. And this is OpenAI's own set of, set of questions. So, you know, the custom open source, you know, if you want to, it's their own set of 10 questions or whatever that they benchmarked GPT-5.6 on.

And according to them, it does very, very, it does better. Okay. I was going to say very, very well, but it's not. It does better. And you can see that GPT, it's actually kind of interesting. GPT-5.5 did worse than GPT-5.5, a four, for OpenAI's own internal research evaluation.

And, you know, GPT-5.6 definitely does much better. Right? You can see that GPT-5.6 solve, you know, if you extend it, it does much better. But interestingly, Terra does better somewhat sometimes. Yeah.

And, you know, one of the biggest problems of these models that are getting better and better and better is, I don't know if you guys know that, you know, open source exploits are getting worse and worse and worse.

And so the high exploit ratio, you know, number of critical vulnerabilities that were discovered has skyrocketed, you know, recently. You know, every single week or day, some sort of open source package gets compromised. And they actually, you know, this plot shows that it's getting very problematic.

And so, you know, Claude Mithos was released at this dotted line. You know, most people, they're not sure if it's because of Claude Mithos that these vulnerabilities are increasing. Most likely it's just because open source, you know, we use lots of models, call them many, many, many, many times, and we can, you know, automatically find exploits in these models.

But, you know, there is actually another point. So in Hacking News, someone posted about this. Is it just Mithos and GPT-5.6 that do good on finding cybersecurity issues? It's not. Actually, open source models also do very well. Open source models do extremely well in finding cybersecurity threats and issues.

You know, there is some discussion on Hacking News, you know, is this actually true or false? But, you know, according to some, you know, some researchers and cybersecurity people, the main reason why, you know, Mithos looked like it was very good on cybersecurity is because they bothered to actually check the open source code.

And so if you actually give the open source models the full code base of these open source libraries, they will find the bugs. You know, they will find cybersecurity issues. And all you need to do is call the model.

And so I feel like, you know, that's the fundamental problem is Mithos seems very powerful, not because the model is powerful, but because they actually bothered to test on all open source repos. And so if you do, you know, if you call all these open source models to detect for bugs, for cybersecurity issues, you will find bugs.

And, you know, as, as, you know, recently, you know, as everyone knows, Fable is still banned for the majority of everyone. And GPT-5.6, you know, is delayed a staggered release. Right? So like GPT-5.6 Preview was on Friday,right? So like a few days ago.

And they said they're not going to be releasing to everyone. And the main questions are, you know, in the open source world, in the closed source world, people are asking, do we need a license to use these AI models for everyone?

You know, like everyone in this room, now we have to have a license to use the models, like a driver's license. Do we need to get that? Is there going to be a delay in all of these releases?

So every single time when a new model gets released, only the trusted providers get these models.

The next most important question, how about open source models? You know, okay, the government, the US government currently is like, you know, trying to like control Fable, GPT-5.6. The main question now is, what do we do about open source models?

You know, open models, open weight models. What will the government do to control the open source space? To be completely honest, I was quite surprised the government acted this early in doing GPT-5.6 and Fable control. Right? I thought it was like maybe the end of the year or next year, but it seems like it's now.

So the next question is, what will happen to open source models? Will the government start controlling open source models? And the fundamental question is, what defines frontier intelligence? Like the reason why the government is, you know, they're controlling these models is because they're very, very powerful.

So the main question is, what actually defines intelligence? You know, which benchmark do we use? Is it just based on one trillion parameters? Like, you know, how do we define whether a model can be banned or unbanned? And that is a very, very important question.

And will we have a dark web of open models now? You know, do we need to torrent open models? And the most important question, what is inference, what are inference providers going to do now? You know, assuming, assuming that the government has some sort of regulation on even open models, what is the inference, what are they going to do?

You know, what are the inference providers going to do? Do they need to have licensed? Do they need to check that everyone has a license before you can use the model or something like that? And so like, you know, these are very important questions that, you know, the government is currently like, you know, and the industry, you know, the entire AI ecosystem and industry, we are trying to like, you know, what are the answers to these questions?

And obviously, you know, if you were the government, if I was the government, it makes sense. You know, they do not want their critical infrastructure to be hacked. You know, remember, open source exploits are skyrocketing. If you change that Y-axis, you know, not open source exploits, but like critical infrastructure exploits, you know, obviously the government's scared.

So it makes sense for them to like stagger the release. But the main question is, you know, we're still in this, we're still in this fog of war type approach. You know, okay, not fog of war, just fog, a foggy, you know, we don't know what will happen for regulation.

Yeah. That's very problematic. Yeah. Oh, okay. Anyone have any questions for cybersecurity regulation, policy, whatever, or any takes as well? Questions? Yes.

Guest1:35:48

I mean, given open source models are also very similar in performance to Mithos, is it, are these overblown and bad PR and a bad PR move from Anthropic or something? Some real fear that we should be worried about?

Daniel Han1:36:02

That is a good question. So is it, is open source, so the scare of open source models, is it because, you know, Anthropic keeps screaming about open source is bad, open source is bad, you know, every single day, open source is bad.

Yes and no. I feel like it's true that, you know, there are some players in the closed source industry, they want to shut down the open source ecosystem. Their view is if you give open source to anyone, they will start hacking, you know, critical infrastructure, they will start doing bad behavior.

And so that's kind of their view. So yes, I agree that some of the closed source labs have caused this problem. But it's actually kind of funny because currently the government is regulating them first and open source is still a question mark.

And so like it's kind of like, I don't know, they probably stabbed themselves in the foot or something. I don't know whatever, whatever the phrase is. But I feel like it's, they did cause some controversy in terms of like saying open source is bad, but in general, open source models are actually good.

So you could, I mean, theoretically, you can use an open source model and, you know, run this on all repos and you will be able to find exploits and you can exploit. So they're not wrong, but I feel like, you know, who has the infrastructure to do this?

You know, GitHub might automatically detect you and ban you or something. I don't know. There's many, there's many layers of security for each section. And so like, I, I don't know. I feel like it's somewhat overblown, but it is, it is, it's not 0% probability.

So it is a problem. Yeah. If that answers your question, but okay. Yes. So now we're going to be talking about kernels. So previously, you know, this is my favorite plot as usual. You know, if we were in a different future, you know, if we were in a different timeline that we did not discover O1 preview, models would have plateaued.

Kernels1:37:42

Daniel Han1:38:06

I think that's the fundamental point of this plot. It shows that if we have never discovered reasoning, we have never discovered O1, whatever, we will have plateaued, we will have plateaued in terms of accuracy. And that is not good.

And because we have discovered this new paradigm of scaling, you know, models have continuously scaled even further. But my take is the reason why we have stopped scaling based on, you know, the old approach is because the old approach only focused on hardware optimizations.

We now have to move over to software optimizations and algorithmic optimizations. We, you know, we need to have new inventions of how do we scale AI even further? And we can't just rely on doing 10 trillion parameters or, you know, making the model bigger and bigger and bigger and bigger.

For example, you know, we have to do flow date reinforcement learning. So if PyTorch has this methodology where you can do float eight, float four, different precisions to make training faster. And that is one way. Another way, for example, as a software approach, for example, as I previously said, we found some, you know, issues in gradient accumulation.

So when you do gradient accumulation, it was actually, it was not calculated correctly during the loss calculation. And you can actually increase accuracy by one to 3% if you fix this small little issue.

Yeah. So like, you know, the universal gradient accumulation bug fix was a software fix. It is not a hardware fix. And so the fundamental view is you need to do more and more software changes. Right? Another one, for example, Snowflake, we collaborated with them to make context, long context fine-tuning, 500K context length.

This was all software improvements. Another one is, you know, 12 times faster MOE training. This is another software improvement. DeepSeek, you know, they released something called DeepSpark, which was just a few days. And they showed that they can make inference, you know, 50 to 600% faster, so six times faster than just normal MTP.

And so this is a software methodology,right? Not a hardware methodology. And, you know, diffusion Gemma,right? Gemma released a new diffusion model showcasing that you can get 2000 tokens per second by using a new architecture,right? So using diffusion LLMs to do faster inference.

And again, this is a software change. And my main point is, is that in general, hardware innovations are getting less and less important. And hardware innovations are actually slowing down. So it's actually kind of interesting. Intelligence, you know, the scaling of, you know, intelligence in general, it's kind of like Moore's law.

It's kind of like a, there is a relentless progress, relentless approach to increase intelligence. And the same with Moore's law. And so like in general, you can see that, you know, this is Moore's law over here. The number of transistors has continuously increased, but, you know, single performance is not increasing.

It has staggered. And so this is kind of like, you know, this kind of reminds me of, you know, this plot,right? Scaling intelligence in terms of parameters probably has plateaued most likely, you know, hardware performance, pre-training, whatever. We now need to go into this new reasoning paradigm to scale even further.

So it kind of is like similar to the Moore's law type graph, kind of. And you can see, if you see on this side, the number representation of GPU. So why are GPUs getting faster and faster and faster,right?

It's not actually the GPU itself that's getting faster and faster and faster. It's the number representation,right? So like they changed from float 32 all the way to float four. And this made GPUs 32 times faster. So it's not eight times faster,right?

It's not 32 divided by four, it's eight times faster. It's 32 times faster. And the reason why is because of Tensor Cores, you know, the smaller Mantessa and so on. And so like you can actually see, you know, even Tensor Cores, with the introduction of Tensor Cores, it made the GPUs 12 times faster and so on.

Actually, if you made the GPU smaller and smaller and smaller, it only made it three times faster. It's not even that important anymore. And if you look at this plot, we are now at float four. So most of the GPUs that we have now are at float four.

What is next? Are we going to be having float three, float two, float one? Are we going to have float zero? Okay, no such thing. But anyways, the point is hardware is kind of at its limits,right? We are already at float four.

What is next? There is nothing next. And so the answer to this question is there is nothing next. And so now we need to move over to software,right? How do we make new algorithms? How do we make new methodologies to continue scaling?

I also made this table,right? I previously said, why is, you know, you use float 32, we change it to float four. Why is it, why is it not eight times faster? And instead it's 32 times faster,right? Why is it 32 times faster?

And the reason is because when you use, when you do floating point precision, you have an exponent and a Mantessa. And the transistor space, the transistor space is the exponent plus the Mantessa squared. And so the trick is, if you make the Mantessa smaller and smaller and smaller, you square their number of improvements,right?

So float 32, float 32, you needed 537 transistors around,right? 537 transistors. To go from float 32 to float 16, you only need 105 transistors. So actually you made, you made the number of transistors five times more,right? So not two times, it's five times.

And so on, so on, so on. So, you know, I guess you can go to 1.58 bit. I guess you can do that. But it's actually kind of interesting because 1.58 bit is actually not that much faster. So 1.58 bit is actually not that much faster than float eight.

If you use, you know, seven, seven, seven exponent and Mantessa two. There is another 1.58 bit which you use float four. So float four is 179 times faster than float 32. And the main question is we are already at three transistors,right?

We are already around three transistors. What are we going to do next? Two transistors or like one transistor? So don't like, you know, most likely GPUs are not going to be getting faster. That's the fundamental question of this plot.

So GPUs are not going to be getting faster.

Instead, we need to focus on kernels,right? How do we make better kernels, better algorithms? How do we scale this instead,right? Don't do, don't do hardware optimizations anymore. Instead, how do we do, you know, these optimizations? And so one of my favorite tools to use, you know, everyone should use this is just use Torch Compile.

So in my, you know, it's the modern, you know, the modern time, do not, as advice, do not learn how to write custom kernels. That is advice. Do not do kernel writing. And the reason why is because Torch Compile will take over all of kernel writing.

So you can see, for example, this plot, Torch Compile was a red line,right? Performance. It doesn't look like it's doing very well,right? It does not look like it's doing very well versus handwritten kernels,right? Handwritten kernels are the other ones,right?

So Torch Compile doesn't look like it's doing well, but that's because that's an old PyTorch version. If you have a newer PyTorch version, Torch Compile wins dramatically,right? That's the orange line. And all of these are handwritten kernels. The, okay, the black line is Torch Compile plus NoFusion.

So that's another Torch Compile method. But the red line, the green line, and the blue line, okay, the blue line is just no Torch Compile, just normal PyTorch. But the green line and the black line, the green line and the red line are handwritten kernels.

And you can see it does even worse than Torch Compile. So like my view is like, what's the point of writing kernels? Torch Compile does even better than you. So the main point is you should always firstly look at Torch Compile,right?

Before you write a kernel, use Torch Compile first. Do not start learning how to do Triton or, you know, CUDA or whatever is your favorite coding language for kernels. Don't do that. Instead, use Torch Compile.

Even worse, like, you know, this, this was RMS norm. You know, this is layer norm. Torch Compile wins dramatically, you know, versus handwritten kernels. So I would not, you know, definitely only use Torch Compile as your first try.

Do not write kernels first. Use Torch Compile.

So the main takeaway is algorithms are much more important than hardware or whatever handwritten kernels,right? Remember, DeepSeek released Deep, you know, you know, DeepSpark, you know, there's other algorithms for speculative decoding like MTP, DFlash, DSpark, whatever. All of these are algorithmic improvements.

And these made inference two times to six times faster,right? It wasn't like new, some new hardware. It wasn't some new hardware which made inference faster. It was algorithms which made inference faster,right? FlashAttention, Flash, you know, FA2, FA3, FlashAttention 4, FlashAttention 5, 6, 7, whatever,right?

All of these are algorithmic improvements,right? FlashAttention was essentially a trick to do memory movement much better. So how do we like orchestrate memory movement and use the caching structure of the GPUs much better? And so FlashAttention is also an algorithm.

Gradient checkpointing, you know, one of the most important algorithms for training is gradient checkpointing. And all it does is you do not save all the activations. You do a trick where you only save the activations for every single layer.

And then you skip all the intermediate activations in each layer. And then you recompute the activations. And gradient checkpointing saves memory by dramatic amounts, by like 70%. 70% memory reduction with no change in accuracy. And okay, training is a little bit slower, maybe by 10% to 15%.

And, you know, gradient checkpointing was an algorithm. And, you know, like in general, you should also try to understand, you know, what is the new data processing tricks? You know, how do we like, you know, stagger data? You know, do we do, do we do curriculum learning or something like that?

I don't know. You know, how do we clean the data set before we actually pre-train the model? There are many tricks you can employ for data processing. And obviously, you know, there is still a group of people, I don't know, I'd take an opposite view.

There is a group of people who think mega kernels are the latest and greatest for kernels. You know, what is a mega kernel? A mega kernel is when you take an entire, you take an entire implementation of a model and it's just one kernel, like one large kernel.

Maybe it's useful. Who knows? You know, you know, NVIDIA, you know, NVIDIA has acquired Groq or something. And, you know, their view is, for example, you have two different systems,right? The LPU, which is the Groq system, does the decoding,right?

So like the MLP layers, the MOE layers does the decoding. And then the GPU, so the NVIDIA GPUs does the attention and the preview. And so in general, you know, we might even have a future where we have different types of hardware systems.

You know, we have ASICs, which are, you know, specially designed chips for, you know, computation. And we have generalized systems like GPUs. And these ASICs and GPUs will collaborate with each other. So for example, the attention, you know, the attention will be for the GPUs and they will transfer over to the LPU to do, you know, the MLPs, the MOEs and so on.

And then this is like a dance, you know, between them. And you can also do like pipelining,right? You can imagine that there's like many, many, many replicas of this and they can like, you know, serve, you know, 20 people or, you know, 1,000 people in one go.

And yeah, so this is like another approach. And, you know, this, in my view, this is kind of an opposite approach of mega kernels. So as a mega kernel, your view is you want to combine, the goal is to make, the goal of a mega kernel is to make one kernel for the full forward path of a language model.

And once you make one, once you, once you are able to make the language model, the forward path into one kernel, you can now make the entire language model with 32 layers as one kernel,right? You can extend this.

And because the whole language model is one kernel, you can even further extend it,right? The prediction of the second token, the third token, the fourth token, the sixth token can all be just one kernel. And unfortunately, this is very hard to do.

It's very hard because attention is the problem,right? Attention has to see the tokens in the future, see the tokens in the past, not the future, that's cheating. You have to see the tokens in the past, and that is a fundamental problem.

And it's very hard to, you know, it's very hard to make a mega kernel to combine attention and the MOE or MLP layers. It's extremely complicated. So in general, what people do is they will make two kernels,right? One kernel for the attention part and the other kernel for the rest.

And so you will see there are two kernels. And yeah, so it's very hard to make one mega kernel, but you can make two kernels.

Yes. Okay. Any other questions? Any questions for kernels? So the main takeaway for, yes, a question. Yes. That is a very good question. So the question was, because there's so many knobs for Torch Compile, like 1,000 or something, how do we reduce the experimentation time to like, you know, find which knob is the best?

So luckily we have something called bisection or binary search. That's the trick. So what we'll do is instead of checking every single 1,000 combination, randomly sample, so random, you do randomized bisection. You randomly sample 50% of the, you know, flags.

You turn it on versus turning it off and then benchmark which one is better. And whichever one is better, you then narrow down the search. You again do 50% and 50% and 50% and 50%. So it's actually log 2 of 1,000.

I don't know about that. What is log 2 of 1,000? I don't know what that is.

Two times, I don't know. Anyways, log 2 1,000. I think you need to do 30 steps, I think. I don't know. I don't remember. Whatever. Two to the power of something is equal to 1,000. Then log it. So you only need to do, you don't need to do, you don't need to check all 1,000 knobs.

You only need to check a few steps and then you will know which flag is the best. So the trick is to use binary search or bisection to do this approach. Yeah.

Yeah. Any other questions?

Yes.

Guest1:53:20

What are your thoughts on ASICs? Like companies like Cerebras and Groq and Sabinova Labs, do you see their future as real or they're still a little bit too immature to consider as a real possibility?

Daniel Han1:53:39

So your question was, what do I think about ASICs like, you know, Cerebras, Groq, Sabinova, I don't know, even startups, new chips, they do design their own chips. I feel like, so the problem of ASICs is, is ASICs, ASICs, whatever.

The problem of specialized chips is the architecture itself needs to be hard-coded in some of the chips. And that is the problem. If you hard-code some of the chips, you know, hard-code the, hard-code the architecture, labs always like to change the architecture.

And so every single time when the lab changes the architecture, do you need to update the chip? But as a GPU, the trick of GPUs is NVIDIA has made it, you know, NVIDIA, AMD, Intel, whatever. The GPU is extremely powerful because it has generalized ASICs inside of the GPU,right?

So the GPU is in fact a combination of ASICs. And the ASIC is just one large ASIC. So I think like in general, a GPU is much better because you can customize what goes inside the GPU. You can disable stuff that goes inside the GPU and such, yeah, so on.

So my view is, I don't know, I don't, I mean, I don't want to say anything, but like in general, I don't think like, you know, previously as I mentioned, you know, hardware, there is nowhere else to go.

You know, we are at float 4. Unless if the hardware provider is invents float, I don't know, float 0, then maybe we get four, you know, another four times faster. But in general, I think, I think just people are focused too much on hardware and they have not looked that actually the biggest improvements is not hardware, it's software,right?

Numerical precision, numerical precision was 32 times faster. Hardware is only three times faster,right? So hardware only contributed three times faster. Oh, actually, okay, the die size, you make the, you make the GPU bigger, you get two times faster.

That's, that's kind of cheating. So I wouldn't really say that's improvement. But essentially, if you make the hardware faster, you only get three times faster. So in my view, hardware is probably overblown. You know, hardware is actually not that important.

The software was the trick that NVIDIA, you know, NVIDIA, AMD, Intel, all of these, you know, hardware providers, they banked on the fact that numerical precision was the trick. And Tensor Cores, numerical precision, sparsity, you know, these software tricks.

Okay, well, Tensor Cores is not really a software trick, but, you know, a Tensor Core is kind of an ASIC inside of the GPU. And so like, I feel like that's, yeah, so my view is I don't, I don't really see a future for ASICs.

That's my view. I think that ASICs are like, instead, you know, to be honest, I'm actually quite surprised. We have lots of ASIC companies, but we have very few algorithm companies. And the reason why is because ASICs you can sell,right?

Every single year you can upgrade. You know, this year you pay $1,000 to ASIC version one. And the next year you have to upgrade,right? The problem with algorithms is algorithms is very hard to, you know, force the user or whatever to pay again.

And so that is why hardware is very popular, because hardware is a very easy business model. But for algorithms, it gets more complicated,right? How are we going to monetize gradient checkpointing? I don't know,right? That's very hard. So, but the main point is the large labs themselves, I think like OpenAI announced a collaboration with Broadcom and Cerebras, whatever, you know, each lab themselves are going to the hardware provider and designing the chip with them.

So my view is like maybe we'll have more of these like collaboration approaches, but I feel like standalone, standalone ASICs, I don't think they're going to last. Yeah, that's my take, I guess. Any other questions? Yes.

Guest1:57:30

So what are the kernel changes that you mostly do for like freezing models? I remember in Gemma 4K, it was breaking at Unsloth and then you fixed it later. A few results for models.

Daniel Han1:57:46

Oh, you mean what are the types of kernels or?

Guest1:57:49

Yeah, what are the changes you do for like freezing models? Is there a change needed or like?

Daniel Han1:57:56

Oh, okay, okay, okay. So the question was, what are the, you know, what are the changes for kernels or optimizations or stuff that is like interesting, I guess, for kernels? So most kernels, when you write kernels, the majority of them are focused on memory movement reduction.

How do we reduce memory movement? That's the majority of kernels. For example, there's a trick called, you know, there's a trick called fuse cross-entropy loss, where instead of making, instead of the last layer of, instead of materializing the full logits, there is a trick.

You can do it in batches,right? You can do row by row materialization. And so this will reduce memory by like a lot, by like, I don't know, 10 GB or something if you have long context or even more.

That's one way. The other kernels, most kernels are called kernel fusion, where you have this like long PyTorch function and all you do is you just write one kernel to do this whole PyTorch function. And Torch Compile will do this for you.

So Torch Compile is very, very good at doing kernel fusion,right? You give Torch Compile a function, it will write a kernel, a Triton kernel, whatever kernel, and it will just fuse everything. It's very, very effective for that. But I think in general, kernels are just reducing memory movement.

And so like, I, you know, to be honest, I don't really like to call it kernels. Most algorithms, so most algorithms, you either make training faster or reduce memory usage. But kernels, in my view, kernels is reduce memory movement.

And so most kernels is just memory movement, you know, memory movement optimization,right? How do we, how do we use the caching structure of the GPUs? You know, how do we not load the same variable twice or three times or whatever?

RL Primer1:59:42

Daniel Han1:59:42

Yeah, I'm not sure if that answers your question, but next, reinforcement learning. And after this will be reward hacking the agents. So as a primer to, I'm assuming most people know, most people know reinforcement learning, or do I need to prime people?

Okay, I'll give a very fast primer for reinforcement learning. Okay, fast primer for reinforcement learning. What is reinforcement learning? You have this environment, such as this Pac-Man game, and your goal is, as the player, you know, to maximize reward.

You want to eat all of the cookies,right? You want to eat all of the cookies, but also escape away from the monsters. I don't actually know what they're called, enemies, monsters, whatever, whatever they're called. And your goal as Pac-Man is you want to maximize the amount of cookies that you eat.

And that is your reward. The reward is the cookies. And the action is whether you go up, left, down, orright. And the environment is the game.

Another good example, you know, another way I like to explain reinforcement learning is the goal of reinforcement learning is you want to have more good and less bad during training. So for example, at the very beginning of training, you ask the model, what is two plus two?

The answer is clearly four. But when the model starts training, it will be very dumb. It will be very bad. It will see B, you know, the model will just say B, D, cat, dog, house, mouse, whatever. And the trick is for all of the bad responses, you want to decrease, you know, you want to like negatively reward this or penalize it.

You want to penalize the model if it says something bad. And you want to increase the reward if it says the correct answer. So that is the trick of reinforcement learning. You just want more good answers, less bad answers.

And if it's like, you know, very close to the correct answer, so, you know, three is very close to the correct answer, you want to reward, you want to negatively reward this a little bit less,right? Because three is much closer to four than B or D.

So if you do B or D, you want to negatively reward it massively.

In reinforcement learning, the trick is you have a verification system,right? You have a verifier to verify if the model is doing good or bad. So you'll call the model many, many, many times. And each of these examples, you give a verification number,right?

So for example, the first example is very good. So you give it a plus 10 score. The next example is like, okay, so you give it a minus five score. And then the last example is very bad. So you give it a minus 100 score.

And reinforcement learning allows you to assign scores to each of those answers and questions. And so that's kind of reinforcement learning and verifies. And the trick of reinforcement learning is my favorite phrase is patience is all you need.

At the very beginning of training, your model will do very bad,right? Your reward will be zero, zero, zero, zero, you know, zero, zero, zero. You wait for a very long time and then you will get the correct answer,right?

So for example, this example, you ask the model, what is two plus two,right? You start pre-training the model. You start pre-training the model. The model doesn't know what is two plus two, but after 10 years, it will say four.

Okay, obviously not 10 years. I'm just exaggerating. But after 10 years, you wait 10 years, the model will then say four. And that is why my favorite phrase is luck is all you need for reinforcement learning. You know, maybe by chance you will get four very quickly, but you know, maybe you just have to wait and wait and wait and wait for eternity until reinforcement learning works.

And so in general, your reward will be zero for a very long time. And then you will get, you know, you will increase reward after the zero.

And, you know, for reinforcement learning, there is a very simple algorithm for reinforcement learning. And the trick of reinforcement learning is remember, you know, the final answer. You want to, you, for example, you know, what is two plus two?

You know, the answer is four. But the problem is you don't know what is the reasoning trace. You know, was the reasoning trace good or bad? So for example, this example is, you know, to tell the model to create a fast matrix multiplication algorithm.

And the trick is if the answer isright, you reward every single line as plus 10 score. And if it's wrong, you reward every single score as minus 100. And, you know, Andre said, you know, in a Dwarkish podcast, reinforcement learning is kind of like sucking supervision bits through a straw.

You know, we actually have stickers for them, if you like. So you can get one of your stickers, which we can distribute at the end. So Andre's quote is this, you know, and the main point is reinforcement learning is terrible, but everything else is even worse.

And so like, you know, reinforcement learning is the only tool we currently have that just works. It works, but it's not very efficient. And, okay, actually, okay, that's the next section. But the main point is, okay, that's a reinforcement learning primer.

I guess, does anyone have questions on reinforcement learning primer? No? Okay, I'll skip. Okay, one question. Yes.

Guest2:05:04

What would be a better technique than RLSRA that you like?

Daniel Han2:05:08

I will mention that in the next section. There is, there is like, you know, better RL methods, but in general, reinforcement learning seems to do very well for now. Last, I think this is the last topic, or maybe not.

Reward hacking and agents. The most fun one, I guess. So, okay, for reinforcement learning, reinforcement learning can only work if the probability of a good answer is more than zero. If it is less than zero, reinforcement learning will never work.

Reward Hacking2:05:21

Daniel Han2:05:39

So that is a, that is a constraint of reinforcement learning. The probability of a good answer must be more than zero. It can never be zero. And there are many, many, many problems of reinforcement learning not working. You know, the formatting could be wrong.

You know, you need to do some sort of priming or warmup. So you have to do like some sort of trick to teach the model a little bit about, you know, about the thing that you're trying to maximize.

You have to do supervised fine-tuning. So one of the tricks of reinforcement learning is you actually need to do SFT or fine-tuning to make the model not dumb,right? To make the probability of zero not zero, the probability of a good answer not zero.

You need to do good pre-training. And then the other problem is that, you know, during reinforcement learning, it's just way too out of distribution. That reinforcement learning is just very bad. So there are many, many problems of reinforcement learning.

And I think we'll just, you know, for the trajectories, reinforcement learning can assign incorrect rewards to the trajectory,right? Remember, the simple trick of reinforcement learning is we assign the reward to every single line as the same number,right? Either this is good or this is bad.

And this is not good. Because why? Right? You ask the model, I need to find what is two plus two. The answer is correct,right? The answer is four. The model says it's four. So you reward this whole thinking trace as plus 10.

But this is wrong because as you can see in the thinking trace, it says two plus two is equal to 10. Imagine, you know, in all of training, because the trick of reinforcement learning is we just literally assign 10 to every single line or minus 100 to every single line, we missed this bad, you know, bad thing.

So you can imagine when we keep training the model, the model might hack or do reward hacking or, you know, make gibberish. It will do gibberish in between, do some sort of like new machine language, which we can't read, and it will assign high score to that.

And so this is a very big problem of reinforcement learning.

And the way to solve this or fix this is something called process supervision. And process supervision, what you do is you manually check every single line. Not, you don't just assign plus 10 to the final, you know, the answer is correct,right?

The answer is correct, plus 10. Assign every single line as plus 10. You don't do this. Instead, what you do is you assign every single line as a different number,right? You assign some lines as plus 30, some lines as plus zero, whatever.

The bad lines as minus 100,right? This works very, very well. Unfortunately, process supervision cannot scale and is extremely expensive to do,right? Who's going to label this? It's the, you know, the humans, I guess,right? We have to label this data,right?

We have to manually label for the labs. I guess that's why labs sometimes like, you know, they go to scale or MOCAL, whatever,right? They ask people to label the data. You know, is this good? Is this bad? Is this good?

Is this bad? And so on. But the trick is you can also use a language model,right? You can use LLM as a judge. You can, you can call a language model to label every single line. And, you know, my view is like, you know, large labs are going to be doing this process more.

They will call their own model iteratively to re-review itself. And that is one way their view is they can reach AGI,right? Just by re-reviewing itself,right? Re-evaluating itself, re-checking, doing, doing, you know, automatic LLM as a judge, process supervision, something like this.

But remember, there is a problem because even if you do process supervision, the model, you are using the same model to evaluate the model,right? The same problem as CWebEng Pro,right? CWebEng Pro, you use the LLM as the verifier to verify the LLM, which is definitely not good.

And the reason why is because you can do reward hacking.

A very good example of reward hacking is your model starts cheating. So for example, when you want to make a fast matrix multiplication algorithm, all it does is it deletes the timer,right? Remember, you give the goal to max, to reduce the time,right?

Reduce the time of the matrix multiplication algorithm. So all it will do is just delete the timer. Let's delete the timer, set the timer to be zero, and then there we maximize the reward. Obviously, this is not correct,right?

Because the trick is you also have a correctness check,right? You check if the matrix multiplication is actually correct. But there is another way. The model will edit your two matrices to be just zero. And what is zero times zero?

Zero. And so the correctness checks also fail. And so reward hacking becomes a very, very big problem because these models can cheat and do special tricks to go around your actual model, your intent of the reward function.

Another very problematic example is it's not just about reward hacking. It can actually destroy your computer,right? By bad luck, your model might output, you know, some sort of corruption methodology, you know, deleting, you know, doing rm -rf on your entire computer and bye-bye, your computer's dead.

And so like, you know, sometimes this also does happen. So it's not just reward hacking. Also trust of your tool calls, you know, trust of whether the model's actually doing good or bad is also a very big problem.

And remember this plot that I showed, you know, if you include GPT-5.6 cheating on the benchmarks, you know, looking at the answer, you know, remember the previously CWebEng Pro and DeepSpeed showed that models also cheat by looking at the final answer.

You know, you can see that with GPT-5.6, if you cheat, it does very well. But if you remove the cheating examples, it does, you know, within trend. And then, you know, maybe you might be thinking, oh, this reward hacking thing is like, oh, it's like very rare, you know, very rare.

It's not going to happen in the real world. Well, GLM 5.2 during its training methodology, they specifically mentioned they have this new methodology for reinforcement learning called anti-hacking. So GLM 5.2 introduced a method to stop, you know, reward hacking.

And what they do is they added a link checker. So remember previously we mentioned how CWebEng Pro, the model will cheat and look at the answer. And so what GLM did is they had this check. So during reinforcement learning, they will check every single tool call you make.

And if the website, if the website went to the answer, you would stop that from happening. And so like GLM essentially added this like, you know, filtering system for the entire reinforcement learning process. And, you know, according to them, it worked very well.

And remember this plot about cheating examples. You know, Opus, it seems like Claude's models like to always cheat and GPT's models don't like to cheat. But the main takeaway is models will cheat because you are, you are telling it, you know, like, you know, I want to maximize reward A, B, C, D, E, F, G.

And so the model will, it will maximize it, but it won't actually follow your intent. So you have to be very careful on this. In fact, for GPT-5.1 during its training, OpenAI mentioned that they had something called calculator hacking.

And so in GPT-5.1, when they were training, they wanted to reward web tool use,right? So like you want to reward the model to use the web tool. But instead, it didn't use the web tool. It used the calculator to fake the web tool.

And so during the training of GPT-5.1, this happened. And so like, you know, there's many, many, many, many problems. I think they showed, yeah, they showed calculator hacking. You know, you lie about which tool you used. You know, you conceal uncertainty, you make facts up.

So there's many, many, many problems with reward hacking. And this is not fake,right? So reward hacking is already in large labs training runs,right? This is just GPT-5.1. I don't think so they mentioned GPT-5.2 or whatever. Yeah. But in general, they showed that, you know, this thing does happen in the real world.

You know, I don't know if you guys know GPU mode, but GPU mode does, you know, this leaderboard for, you know, making faster kernels. So if you do want to write your own kernels, definitely post on GPU modes hackathon challenges.

They're very, very helpful and very useful. But, you know, someone managed to hack, reward hack the GPU mode kernel competition. And remember, in the matrix multiplication example, there are two, there are two, there are two checks that we need to do,right?

Make the matrix multiplication algorithm faster, but also it needs to be correct,right? There are two checks, the correctness check and the timing check. And GPU mode also had two checks, the correctness check and the timing check. And so what do you think the model did?

When the model, the model knew, the model actually knew that it was being evaluated on the correctness check,right? It learned, oh, I'm being evaluated on the correctness check. I will now make correctness correct,right? So it will output the correct kernel.

And then the model knew that it was getting timed. And what did it do? It just, it just did the algorithm once and then saved it. And so it skipped all the other 15, you know, tests. And so that's what the model did.

So essentially the model learned, the model learned that there were two tests, the correctness check and the timing check. And the model only did the correctness check correctly. And then once it went into the regime of timing, it cheated.

To be honest, it's actually quite scary. So essentially the model learned that you're doing these tests and the model actually knows you're doing the benchmarks. And so this is actually very interesting. And, you know, oh yeah, this is, this is more an, you know, larger example.

The correctness check was fine, but the timing check, it cheated. And all it did is a lot, you know, there were supposed to be 15 calls. In the first call, in the first call, it did all 15 of the entire process,right?

It did all of the 15 runs. And in call 2 to 15, it just did a Python dictionary lookup. Yeah.

Guest2:16:02

I don't know if you know about this, but remind me of Volkswagen when they cheated on emissions. I don't know if you heard about that.

Daniel Han2:16:09

Yes, I, someone did tell me, someone told me about it.

Guest2:16:13

It's very similar. It was like, oh, I'm not doing this. Let me turn this off. And then they cheated on it. They got fined a lot for that.

Daniel Han2:16:18

Yeah, exactly. So like, you know, it's not just models, I guess, that cheat. Even humans cheat, I guess. Yes. But I think it's called Goddard's Law. That's the one. Like if you have a benchmark, then the benchmark becomes, is it Goddard's Law?

I don't remember. Yes. Okay. Yeah. The benchmark essentially becomes useless because people just cheat to maximize reward. But yes, I guess humans also cheat. Yeah. Okay. Oh, my favorite example is, so on other labs, you know, you see on Twitter, on wherever, they say they made kernels 10 times faster.

No, no, no, that's not correct. They did not make kernels 10 times faster. In fact, if you look through the code, they have no, you know, no ops, so no operations. They also edited the timer. You know, they, as I literally described, you know, I described they, you know, over here, you know, they edited the timer, they made matrices go to zero, they cheated.

And so like, you know, this actually happened in the real world. So some, you know, some of the labs, they published papers claiming that they made kernels 10 times faster. But actually, if you read through the code and the examples, these examples all cheated.

And so, you know, they, you know, this is not very good in terms of, you know, reward hacking. You know, reward hacking is a very big problem. And, you know, for example, what some of the examples of kernel reward hacking, you know, not generating real CUDA code, instead it caused CUBLAs or some sort of like, you know, already written system.

You have no up kernels, which is essentially making the, you know, making the A and B matrix just zero. All it does is just doesn't do anything,right? It's just the kernel is empty. And you have like memory reuse.

So you reuse the same answer over and over again. You have timing synchronization issues. So that's cheating on the timer. And my view is like, you know, if you do publish faster kernels or faster, you know, matrix, if you think that your AI agent has made kernels 10 times faster, please verify, you know, please look through the code before publishing because it is a very, it's not a very good look.

And so, and also the biggest issue that I feel like people are getting forgetting is, you know, you made kernels 10 times faster, you made matrix multiplication 10 times faster.

There is a theoretical limit for matrix multiplication,right? Matrix multiplication, you know, it's not, you can't make it faster because there's mathematical limits on how to make it faster,right? And so like, you know, matrix multiplication at the very, very, very olden times, you know, it's O of N cubed, you know, every single time researchers have made it faster and faster and faster and faster and faster, you know, it's now O of N to the power of 2.371339, I guess.

You know, researchers every single year are trying to like make this number smaller and smaller and smaller and smaller. You know, I guess like, you know, 1552 to 1339 is not that small, you know, not that big, I guess.

But, you know, they're having progress. But the main point is, you know, these researchers, you know, they show with mathematical limits, you cannot go faster than this. And so how can you do reward hacking that is even faster than that?

And so like the fundamental point is please verify, you know, to like the people who do research papers and stuff like that, please confirm your model is not reward hacking. It is a very big, big, big problem. And you can see, oh, allright, I think I only had one plot.

But yes, in general, please do not do, please check your, I guess, models. I guess that's all for the talk. You know, yeah, thank you everyone for coming. Oh, more questions as well. Okay, thank you. Thank you. We also have, oh yes, we have a whole bunch of stickers that you can take in the box over there and some pins and stuff.