AIAI EngineerJan 19, 2026· 1:15:52

How METR measures Long Tasks and Experienced Open Source Dev Productivity - Joel Becker, METR

Joel Becker of METR argues that AI benchmark scores are soaring while real-world developer productivity barely budges, reconciling this gap by presenting METR's time horizon measurements and a 16-developer RCT showing no significant speedup from AI on mature open-source projects. He cites reliability issues, task distribution mismatches, and the J-curve effect from Meta's developer data, noting that experienced users still fail to accelerate. The episode explores why AI struggles with messy enterprise data (e.g., LinkedIn's 5,000 'impressions' tables), the failure of computer-use agents in 'Agent Village,' and the possibility that hardware constraints may slow progress. Becker also previews new metrics like 'watched versus unwatched' time horizons and future studies on greenfield coding and data science.

  1. 0:00Compute Horizon
  2. 4:23Eval Limit
  3. 6:32J-Curve
  4. 14:10METR Study
  5. 38:29Speed-Up Math
  6. 46:30Watched Horizon
  7. 52:52Beyond Benchmarks
  8. 56:12Agent Village
  9. 1:06:31Physical Limits

Powered by PodHood

Transcript

Compute Horizon0:00

Joel Becker0:21

Here's the very simple argument. If you look at the sub-notion of compute over time, you know, this could be like R&D spending on compute, this could be experimental compute, it could be training compute, you know, whatever that some particular lab is using.

It goes like this, no surprise. If you have another chart of like, you know, log time horizon, let's say this meets a measure from this figure that many of you would have seen on Twitter. Over time, it looks like that.

You know, let's say that this was like not merely a coincidence, but these things were causally proportional. In the sense that if

compute growth were to half, then time horizon growth were to half. So, you know, for the sake of argument, let's say that, you know, starting from 28 or so, the compute curve begins to bend like that, where this would be no growth and this would be the original growth, something like half.

Then, you know, if they were causally related, and in particular they were causally proportional to one another, then you'd expect this to go like that. And then for some milestone that you care about, let's say here we've got one work horizon up there, one month.

Then the delay implied in AI capabilities is potentially enormous. Now, like why, you know, lots of people have speculated that there might be some slowdown in compute growth. I'm not an expert in those forecasts, but I think the prior reasons do seem like somewhat strong to me.

One is like physical constraints that we might hit, power constraints as you mentioned, or there are various other ones that's EPOC have reported on that they consider, all of which seem to not spike through 2030, but, you know, potentially could buy sometime after 2030.

I think the more likely one is just like dollars is a constraint. Like you can't, you know, large tech companies can only spend so much at a certain point, like large nation-states can only spend so much, like you can't.

I guess there are some scenarios in which you can continue going, but that seems to kind of naturally imply this slowing down. And then the, you know, additional point that this paper is trying to make is that under a very contestable but standard assumption from economics, you should in fact expect these two to be causally proportional.

I think in particular you should expect them to be causally proportional to the extent that, or for the period that, software-only singularity is not possible. And that's a whole other discussion that we can talk about. But at least in this kind of somewhat business-as-usual

scenario, or sort of until that scenario no longer applies, I think this is maybe a reasonable model and does imply some slowing of AI capabilities in the near future.

I have no plan for this session whatsoever.

Host3:24

That also assumes that we don't have a technological advance that dramatically improves capabilities relative to compute. Like an unpredictable technological advance,right?

Joel Becker3:34

Yeah, yeah. I mean, all predictions, you know, assume no unpredictable. Yeah, I'm like,

you know, time horizon or like in general in AI, kind of straight lines on log linear plots have been a, I think, you know, a very highly underrated forecasting tool. They've done extremely well over now many orders of magnitude.

You know, I think it's reasonable to have the default expectation that the log linear lines continue through like approximately the same number of orders of magnitude, except maybe if there's, you know, some significant break in the inputs. Yeah, of course, on the upside there could be something quite dramatic.

Software-only singularity is the first thing that comes to my mind, but, you know, another transformer-style moment seems like another candidate, actually.

Host4:23

Of course, also one of the problems with testing this will be that like, I think most of the tasks that you have, a test will equal eclipse the maximum possible amount of time that those tasks could take at some point in the evaluation set.

Eval Limit4:23

Joel Becker4:36

Yeah, so I think, you know, there are some ways around this that we're working on. I'd be excited to talk about that. They all feel pretty early. But yeah, you know, I think it'sright. That if time horizons are doubling, you know, eventually you, you know, the doubling time is such that you can't possibly make long enough tasks in the relevant period.

Host5:00

It's possible also that like we actually hit a place where time horizon is no longer a useful measure because actually you now want time, now you want total time to decrease. Like what you want is you want the same results at a lower time.

Joel Becker5:11

Oh.

One.

Host5:14

You want higher reliability at a lower time horizon.

Joel Becker5:17

One thing to say about time horizon is there's like two notions of time here. Like a human time axis thing, like calendar time axis. The time that the model's working for, I think you should like kind of approximate as zero.

It's not actually zero. They are taking actions, but they largely do their successful work pretty early on to the extent they're going to be successful on tasks. So my guess would be that it will continue to be the case that there's not sort of so much extra juice on that margin of making the model's complete tasks more quickly, although reliability very much so, obviously.

Host5:57

So most of it's like the human, like the iteration loop? Most of the time is spent in like the human-machine iteration loop?

Joel Becker6:04

The humans are working without AIs and the AIs are working without humans. So for the humans, I guess it's all humans.

Host6:10

That's what I thought when you looked at the paper. So like it wasn't slightly different.

Joel Becker6:13

Yeah, yeah, yeah. Yep.

Cool. Any questions on Meetwork? I can go through some like upcoming things that we're excited about if people are excited about those things.

Host6:26

Yeah. I did have one question about the perceived one, like the time perception.

Joel Becker6:31

Yeah.

Host6:32

One of the open source.

J-Curve6:32

Joel Becker6:33

Yeah, yeah. Yep.

Host6:36

One thing I thought, you brought it up a little bit in the paper, which is, you know, whether or not familiarity is a confounding factor. Although one of the things.

Joel Becker6:45

With tools, you're thinking.

Host6:46

Yeah, tool familiarity is a confounding factor. And of course also, like you also brought up that like tool capability has dramatically changed. But there was an interesting presentation from Meta at the Development Activity Engineering Summit this year. And they had done a, they have probably the best infrastructure for quantitative measurement of like developer experience in the world of any company.

And they're able to tell you basically how long it actually takes to make a PR, basically. They call it the asset meta, but like how much actual effort, like human time effort, it took to make a PR. And what they saw was they saw a J-Curve when they gave people agents.

And that J-Curve was, I don't remember how long it was, like three months or six months. And so one of the things that I also wonder is like if it would be interesting if there's a cut-off of how much familiarity the person has.

Like have they been using this as their full-time daily driver for a period of months? And if there's like an interesting cut-off that occurs once there's like a certain level of familiarity occurs.

Joel Becker7:43

Yeah, I'm totally on board with like, not just in this case, but in many economically relevant outside of software engineering cases, you know, J-Curve, like explanations being a real thing. And like, yeah, you know, developers, not just developers, experiment with tools.

You tend to be slower the first time that you're experimenting with tools. But, you know, if you're doing this so that you have some investment benefits, you know, later on you might be more proficient at the tools. Or in the case of AI, maybe you just sort of expect the models will get better.

And so even if you don't become more proficient, it will be like the kind of thing that you want to do. You know, those explanations broadly make sense to me. I can give you some reasons why I'm skeptical.

I think, so one thing to say is, you know, we're, what are some things to say?

As background, you know, we're continuing with this work and we'll see. You know, another thing to say is just like quantitatively, you know, difference between this and this, very large.

I'm like, how much is J-Curve explaining? I think it's not explaining that much.

Host8:54

It does explain that because we see this over and over actually in software engineering studies, that the one question you can't ask people in a survey is how long did a task take? Like you can ask people how much more productive did you feel and they will give you an accurate response that correlates with your quantitative data.

If you ask anybody the amount of time that something takes, they are almost always wrong. So that, I was like, when I shared this with my colleagues, I was like, okay, I'm not surprised about that at all. But what is interesting is how much is the slowdown aspect.

That was what was interesting.

Joel Becker9:24

Yeah, yeah, yeah. Yeah, point well taken. That makes a lot of sense. I do, I do. So I think we, despite this, were interested in time estimates because, you know, we're interested in providing.

Host9:39

Yeah, I mean, the perceptual, like I didn't think that's relevant too also because like the perceptual aspect is also the hype aspect. Right? Like so developers will tell you that they were faster when they weren't. And I think that is worth knowing.

Joel Becker9:52

And, you know, to the extent that we're interested in

measuring the, you know, possibility timing nature of

capabilities explosions or sort of R&D being automated, one commonly proposed measure to do this is just like ask developers or researchers how much they're being sped up. And for exactly the reasons they're pointing at, I don't put a lot of faith in those estimates.

So nice to see it like this. Yeah, some more J-Curve things.

So the forecasters who are not predicting time to complete,right? They are just predicting this effect size. The non-developers, the expert forecasters. They are told the degree of experience these developers have. And some of the forecasters are in thinking about how this population might be different to other populations, pointing out various facts about the study.

Like they're more experienced, I expect experienced people to get less speed up. Or, you know, the repositories are larger, I think AIs are less capable at working on large repositories, I expect less speed up. They never, never mention familiarity with tools.

My sense is that, yeah, they share the sense that I had ahead of time, which was like most of the action is in understanding what's AIs, the kind of things that AIs are good at or bad at in the first place.

And all of these developers have experience with LLMs and their core development workflow. It's just cursive that they're quite, that three quarters of them are totally unfamiliar with at the start of the study. So I just, I wasn't seeing much margin.

Yeah, I don't know. I think it is an open question. I also, you know, we watched so many hours of screen recordings of these developers working. And I just do not see, I think they're like prompting very reasonably, you know, in some cases worse than me and my colleagues, in some cases better.

I'm not seeing these like advanced workflows that they're not accessing.

Host11:46

Yeah, and my experience is not that far off from this is that there are times when like I am dramatically slowed down and there's times when I am accelerated. And although as my familiarity with the tool increases, I definitely don't improve a lot because I learn over time what I can tell it to do and what I can't tell it to do.

In addition to like it's just getting better with it, like understanding like, okay, now I need to plan, now blah, blah, blah, blah. But that's why, so the thing I mean is like.

Joel Becker12:13

Before you make a like high-level architectural decision that, you know, 10 conversations, 10 conversations turns down is going to blow up in your face. You're like really trying to think about it.

Host12:23

Yeah, yeah, yeah, exactly. And also like scope it down to like a smaller problem. Like at first I would try problems that were too large and like it can't handle that. But just, I mean, just for the future if you ever do, I mean, I think it's obviously really hard with the sample, with the 16-person sample size, but.

Joel Becker12:41

That is changing.

Host12:42

That's great, great. Because in the future what I think having a cut-off, like trying to figure out if there is a cut-off of familiarity where the number changes would be interesting to see if that meta result generalizes outside of Meta.

Joel Becker12:55

We are on it. I think the AIs have been getting better during this period, which is going to compound a lot of what's going on, obviously. But yeah, yeah.

Host13:05

Another thing is the projects themselves are very optimized for people coming onto new projects and figuring out how to, you know, they're already, the ones that struggle to be organized well for humans to come on board and be able to navigate them quickly don't survive very long in the open source ecosystem.

And these are fairly mature open source projects. They're a little bit different from like in enterprise settings where things survive because they make money even if they're a pain to develop on,right? So the context is a bit different.

Joel Becker13:36

Yeah.

Host13:39

Yeah, that is a really interesting point. Because like actually some of the repos that I was helped the most with were ones that I was completely unfamiliar with and which had no decent documentation of any kind. And where like I had to come in on this legacy code base that had existed for years and like make a change.

And like the developer who owned it was like only partially available to answer questions to me. And so in that case, like Cloud Code was a huge help.

Guest14:02

Yeah, legacy code bases don't exist because they work well. It's because they make money.

Joel Becker14:06

Yeah.

Host14:09

Interesting point.

Guest14:10

Yeah, the question I had was, like did all the developers have the same level of AI familiarity with Cursor? Was there some variance? And was that like, is there a plot of like each of their familiarity?

METR Study14:10

Joel Becker14:24

There's always a plot.

Guest14:25

Individual users.

Joel Becker14:29

There's always a plot.

Guest14:30

Can we just like kind of like dig into like the question of is there a J-Curve?

Joel Becker14:34

Yeah, so here's some evidence. So, okay, you know, I've shown you some plots. I think the sample size is just small enough that like you shouldn't really believe any of the, I mean, I think the plots aren't going to show much, but then I don't want to say that's like strong evidence.

This is not something that's going on. I just think the evidence is kind of weak. The thing that really convinces me is like, I watch the videos. I watch the videos and I'm working. And, you know, often they're better at using Cursor than me.

And I'm like, wow, you know, I'm working on this project using Cursor. But here are some graphs.

So this is by whether they have various types of AI experience coming into the study. And, you know, basically you see no movement in point estimates. People for whom Cursor was primary ID before, yeah, not a huge amount of difference versus people for whom it was not.

Then the next one is, you know, you might think maybe you have a view that, you know, some J-Curve cut-off comes off at this point. But still, you know, within the study, there's some variation in how experienced people are with AI because they have multiple issues.

You know, after the first AI issue, they're slightly more AI exposed than after the second AI issue. So you might try sort of excluding those data points over time and seeing what pops up. And, you know, they don't seem to get better at using AI over time.

Host15:52

Although I think there's probably a stat-seg issue with that.

Joel Becker15:55

You think there's probably what, sorry?

Host15:56

There's probably a stat-seg issue with that plotright there. Like those bars are very, very wide.

Joel Becker16:02

Oh, I mean, I think, yeah, I think like all of the plots outside of the main plots, all of these subset things, you should like not put a lot of stock in. Yeah, I totally agree. Okay, and then lots have been made, so this graph is the reason we put it in unclear evidence because we're like, ah, things point in different directions.

A lot has been made of this plot suggesting, you know, something J-shaped. In particular, that, you know, at the end, once people have more experience, they do experience some speed up. Here are some issues. You know, first, like the other plots don't.

I think that's important to include. Second, these hours are coded very conservatively. So, for instance, someone in the 30 to 50 hours bucket

had Cursor as their primary IDE in 2024. They had recorded themselves on their time tracking software as having spent 140 hours using Cursor. They conservatively estimated that they'd spent 50 hours using Cursor. And so they end up in our 30 to 50 hours bin.

This is someone whose primary IDE was Cursor last year. And, you know, people have been commenting about this. They've been using Cursor for less than a week. I think that's not a very fair assessment. If you were to move that developer over from the penultimate bar into the, again, you shouldn't believe this because of statistics, but if you were to move that developer from the penultimate effect size estimate to the last one, then you see some balancing out where you get back to essentially zero in the last bucket.

Yeah, again, so like penultimative, I think J-Curve explanations, you know, still like very unmistakable.

Guest17:41

Is it not likely though that the 50-hour group also is similarly underestimating their time they've spent using Cursor and that actually if you just had a longer scale that you would still see a J?

Joel Becker17:52

Oh, that is an interesting point.

That seems plausible to me. And then I guess I want to, I'm not sure it's underestimated because we're using this like very conservative.

Host18:05

Yeah, okay.

Joel Becker18:06

Totally. Yeah, yeah, I think that seems plausible to me. And then for this not to be strong evidence, I'd retreat back to, I think you shouldn't really believe in any of these plots.

Guest18:16

Yeah, yeah, no, I think the biggest thing is it's a small sample size and there's also a lot of bias in the data set effectively,right? Like it's a certain kind of data set.

Joel Becker18:25

You mean like the kinds of developers?

Guest18:26

Yeah, open source developers and also working on open source projects that are pretty mature.

Joel Becker18:30

Yep.

Guest18:31

You know, those two things are, if you're working with open source developers on projects that are pretty mature, this is probably reasonably indicative, maybe, but the sample size is pretty small. But outside of that, it gets a little harder.

Joel Becker18:44

Yeah, in talks about this, I'm like, I think, yeah, this group is really weird. It's really interesting. It's like interesting for the same reason it's weird,right? Yeah, we were interested in, you know, again, studying possible effects of AI for R&D speed up or automation.

There, if any types of developers are not being greatly sped up, it implies the whole thing isn't being sped up. So it is kind of curious to see even like particular weird populations. You might imagine like large, you know, sort of production inference code bases maybe have a bit more of this shape than scrappy experiment scripts.

Host19:22

Yeah, yeah, yeah.

Joel Becker19:23

But yeah, totally.

Guest19:25

No, I think it's very interesting. It's just, it's hard to generalize. We just don't know.

Joel Becker19:29

Yeah. Although we are doing this large study and I think, you know, I think unfortunately after the large study, which includes more greenfield projects, I think it's still going to be hard to generalize for not so similar reasons, yeah.

Host19:42

Although I don't feel like your results are particularly contradictory with any actual independent research that's been conducted. The only research that I've seen that I would say is contradictory to yours is research that has been funded by model shops or agent shops.

Joel Becker19:58

What can I say about that? I do think that most of the research that's put out

is associated with large tech companies. And I think there are other methodological concerns that I have.

Host20:16

I have methodological concerns with that as well. I know people who work at some of those places who have methodological concerns with the work that was output, so.

Joel Becker20:23

I mean, you know, I think there are concerns about ours as well.

Host20:26

Sure, sure. But I actually feel that like I remember somebody sent me your paper and when I saw the headline I was like, no way.

Joel Becker20:33

Well, me too.

Host20:36

I was like, that sounds like BS.

Joel Becker20:37

Yeah, yeah.

Host20:38

I read the paper and I was like, oh, this doesn't suck at all.

Joel Becker20:42

Well, a little bit.

Host20:44

Well, no. I feel like at least your high-level conclusion both is intuitive, like from a person who's read a lot of software engineering research, and also is well justified. I think people, I have had people argue with me about the 16 developer thing, but I don't think that actually matters in that particular case because I think they're actually a fairly good control set, more or less,right, for an experiment because they remove a lot of validity concerns by being experts.

So yeah, it's true that they don't represent certain, like the broad aspects of developers, but they also remove a lot of variance in what you would expect from the population. And they allow you to have like a sort of an epistemological function of like, hey, let's isolate that factor away and then let's see what happens with that.

And that's, I like that. And then I thought the way that the study was conducted was completely sufficient to draw the conclusion, the high-level conclusion that it draws.

Joel Becker21:38

Thank you very much. Here's a curiosity. So we did, we haven't published this because of organizational reasons that I won't go into. But

we did conduct this, you know, people would throw sort of their various explanations for what's going on here. You know, many of which have lots of merit, some of which more skeptical of. You know, a natural one is brownfield versus greenfield projects.

So we ran this kind of enormous hackathon where we randomized half of teams to use AI versus not, kind of, you know, maximally greenfield or something. And then we'd have a bunch of judges score them, you know, many judge scores per project or something to try and even out that noise.

And we'll see, you know, is it the case that like the bottom 50% are all the AI disallowed group and the top are all the AI allowed groups or something like that? Now, unfortunately, it was sort of even smaller.

That's like part of the reason we're not publishing it. I think the evidence is really quite weak. The degree of overlap is enormous. Like the point estimate that we, I'm a bit nervous about saying this because, you know, it hasn't gone through the kind of review processes that something like this goes through.

So maybe I've messed something up. But I think the point estimate is something like four percentage points higher on a, sorry, four percentile points higher if AI is allowed versus if it's not, off the controlling for everything else.

That is like, you know, extremely noisy and you shouldn't draw any conclusions. But seemingly maybe kind of small effects hackathons from allowing AI. Yeah.

Guest23:17

So one question I had, I guess this is related to the study, but also related to other research that you guys have done. So have you found a similar pattern? Or I guess first, have you explored like the effect of AI in other domains and specifically software engineering?

And if so, have you also found this kind of surprising result that maybe it's not as much of a speed up?

Joel Becker23:41

No, no, no, no. I mean, note new directions. Stuff that we have not done.

Yeah, I, yeah, you know, we're interested in understanding the possibility of accelerating R&D. You know, coding is not the only kind of thing that happens at major AI companies. Much more conceptual work happens. You know, I'd be very excited about, you know, working with math PhD students or very different types of software developers or, you know, running these kind of studies inside of major AI companies or large tech companies or something like that.

I think we are very interested in, you know, not necessarily directly, but some somewhat close analogy to the large AI company case. So to the extent that something really deviates from that, that's probably less interested.

Guest24:35

Interesting. So yeah, so I guess it sounds like you're interested in measuring capabilities for like math research and some other research.

Joel Becker24:45

Yeah, I'd say I'm interested in like what the hell is going on in AI. And, you know, how am I going to learn the most about what the hell is going on in AI? You know, I think something a bit more conceptual, something where, you know, fewer humans are currently working on it, so it's less appearing in training data will help me better sort of triangulate the truth about what's going on in AI.

Even if I don't care about math research in particular, it'll still sort of draw helpful qualitative lessons is kind of the sense I have.

Host25:15

Yeah. I mean, if I was going to pick the areas that I think it's most successful in or the areas where I would expect it to be more successful, but where I think it is being less successful, I would pick probably data science as an interesting one.

Like how does data science, how much are data scientists helped by AI today?

Joel Becker25:32

Say more about what you expect it to be less successful.

Host25:34

So in a real, so let me give you an example.

Joel Becker25:38

Yeah.

Host25:40

And at LinkedIn, there are 5,000 tables with the name impressions in the table,right? So if an analyst wants to understand how many impressions happened on a page, where the hell did they go?

Joel Becker25:49

Yeah.

Host25:49

We can't figure that out.

Joel Becker25:50

Yeah.

Host25:51

Today, there is no existing AI system that we have that could be hooked into like a corporate environment like that and processed through, I mean, there's trillions of rows in those tables. So like, so what a data scientist needs to do is they need to be like, I need to like, you know, analyze a bunch of data and come to a conclusion,right?

And I hear lots of like thoughts about building systems. You know, people talk, talked about ML for SQL. The models are much better at writing SQL than they used to be. But I believe that the state of underlying data is so bad that the actual data scientist is going to get way less value out of the AI than the software engineers are also mixing in.

Joel Becker26:34

That is very curious.

So one view that some more bearish people have looking at the future of AI is, you know, so much, there's so much tacit knowledge around, there's so much knowledge that's sort of embedded inside of companies that you're not going to pick up from, you know, these like RL training environment startups or something, something, something.

You know, maybe it's not sort of the state of nature that there needs to be many specialized AIs. Indeed, like much of the lesson of the past few years is that one big general AI seems to be more performance.

But, you know, at some point in the future when data is like locked up inside of companies, you know, we will have more of this proliferation of many more specialist models as I have, you know, GPT-N fine-tuned on LinkedIn data in particular, something, something, something.

I have one reaction that's kind of like that.

Host27:21

Yeah, I don't know.

Joel Becker27:22

I do have a disbelief-like reaction. I'm like, ah, science, you know.

Host27:26

No, but also like, but also like, so contradictory facts. So the problem is all these data sets contain contradictory facts. Like the name of the field will be like, you know, date started, or like it'll be time started,right?

And then it will contain only a date, except for it will only contain the date up until like November of last year. And then after that, it will contain only the month. But then after that, it'll contain maybe the seconds that the thing finished.

And in order to actually successfully query the data set, you, the data analyst or the data scientist have to know what those cutoff dates were. It's not written anywhere. Although what you could do theoretically is import a bunch of the SQL that other analysts have written to try to figure out like how they triangulated these things and work backwards from those reports.

But today, so I think today, for example.

Joel Becker28:17

Sorry, I've just like, I haven't worked with a large company. People don't fix this source.

Host28:22

Oh, no. So.

Joel Becker28:24

I feel like the lesson I learned over and over again, this data specs really matter.

Host28:29

Really matter. No, I've also been working in data analysis and developer research.

Joel Becker28:35

Yeah, yeah.

Host28:36

And so, yeah, so the problem is like their job is like produce this report for this executive,right? Not go make infrastructure to produce this report for this executive.

Joel Becker28:46

Yeah. I'm like, if I, okay.

Host28:50

I'm with you. I live that dream every day. Yeah. You just have, you end up having to,right? You have to build out infrastructure for it. That has to be part of the job description. And the other part is you have to fix the problem at the source.

Like you're really, I still remember having a conversation where someone said, it's too difficult to fix it at the source because there's too much complexity of all the systems that end up at the source. And I said, okay, wait a minute.

You're saying it's too complicated to solve at the source. Downstream somehow a problem that is too big for the entire organization to solve. It's easier to solve there. Come on. That doesn't make any sense.

Joel Becker29:27

I just think there's so much potential here and I have not seen a lot of studies done on like how people who are working in that data space are experiencing AI. And what's fascinating about that is real ML is mostly data work.

Like ML, especially outside of LLMs, ML outside of LLMs, the majority of ML engineers spend most of their time doing like feature curation rather than they spend actual direct model training and like trying to clean up bad data for feature curation.

So like theoretically, the potential even for the improvement of ML by enabling ML to be a better data scientist is huge. And I suspect that if you, my hypothesis is, if you went into this space, you would discover it is great at telling me how to write SQL or how to like write pandas and/or Polars or whatever you're using.

It is okay at doing very trivial things and it fails at all complex tasks. It fails completely at all complex tasks. I don't even know, I haven't even seen a benchmark on it. Can you give me an example of a complex task?

Host30:27

Sure.

Let's say a complex task is determine the time between, give me the p90 of time between deployments for all deployments that happened to Capital One. It struggled with that?

Joel Becker30:44

Yeah, that does seem surprising to me.

Host30:46

That seems surprising,right?

Joel Becker30:48

Yeah. So. And like, you know, if it has sort of reasonable context about where it would find this.

Host30:53

So we find that data,right? Sure, sure, makes sense. And then, so, okay. So fine. So give me that number. And then also, make sure that you can break that down, you know, by team hierarchy. So give me that in a table so I can break it down by team hierarchy.

Where is the team hierarchy data? Like how, here's a funny thing. What PRs were in those? So how do I know, how would I actually determine what the time deployment started and ended was? Because it turns out that's not clear in the base telemetry.

And you have to like know magic to figure out when the deployment started and ended.

Oh, and also tell me, you know, for my ability to analyze it, tell me how many PRs were in each of those deployments and which PRs went into each of those deployments. Well, guess what? The deployment system only, this isn't being recorded,right?

Joel Becker31:41

I think it is being recorded. Yes, but before you.

Host31:47

So then, you know, imagine the deployment system doesn't contain sufficient information about that data,right?

Then like, where do I get that data? Well, that data doesn't exist in any other system. So what I, well, maybe I have to go like, I have to go to GitHub and I have to call the GitHub API.

And like the chance of the LLM or any agent figuring that out today is pretty minimal.

Joel Becker32:13

I do still, you know, relative to my colleagues, I'm pretty bearish on AI progress. I do still have some reaction that's like, ah, like can't you spend a day getting this into a Cursor rules file? You know, like where the hierarchy exists.

Host32:30

I would go, I think that's why I think it's interesting. I think what you were saying, I have not seen any real comprehensive study on the experience that data scientists have.

Joel Becker32:40

If you have any ins

to us running studies at large tech companies, then I am all ears.

Host32:48

There is a fellow at OpenAI that I was talking to, who was one of the speakers who does evals, internal evals. And he has mentioned that he's done some work with data scientists. So he might know some people who have that data.

But it's all been internal between him and like Cursor, between him and like, you know, Anthropic or whatever,right?

Yeah, and I also think one of the ones that I'm curious about too is lawyers.

Curious about like more traditional, like older lawyers, doctors, and I think mathematicians are all really interesting to me. Just because, you know, both lawyers and doctors are so constrained by a legacy history of like the constraints around them and how they work.

Joel Becker33:35

Yeah, legal issues, I'm imagining that needs to be a significant barrier.

Host33:39

Yeah. And there's stodging. Like I'm also interested in like what's the, how are the stodging?

Joel Becker33:46

Stodging I feel like is a, I think I'm less bought into as a long-term explanation for, I'm like the legal restrictions, they sort of continue to be the case through time. The stodging is I can like set up a new law firm that's less stodgy and then take the previous law firm.

Host34:02

I agree. I agree. I don't think it's persistent. I just think it's interesting to see. One thing that would be interesting to see is like if that affects the mental model that they have today. Like if they're, like how they've been talked to about it or how their trust in it affects how they use it.

It'd be interesting to know. To me, I don't know, it's a worthwhile study. It's more of like one of those things that I wonder about idly. You take a lawyer who just got out of college and sort of, you know, has spent a lot more time using ChatGPT and you take a lawyer who's been in the business for 50 years and, you know, has a giant folder full of Word docs that contain like all the briefs that all their junior associates have written for decades and decades.

And he just opens up those briefs and like changes a few words in them and then sends them out to the judge. And he like, you know, has known those judges for like 30 years, 40 years. He knows exactly what they want.

And like, you know, is he getting any, is he going to get any value? But is there a value he should get? Is there something that, is there some way that like he would be helped by AI? I certainly know discovery.

Discovery in AI is like in law is like a huge, huge problem. And I know that like there's Harvey. I don't know anything about how much success they've had.

Guest35:17

I know a lot of people working in that space specifically, like that's an ongoing thing,right? There's always technology for it, but it's kind of, the adoption of it is a very different thing from.

Host35:31

That's the thing,right? Because one of the first things that I thought of, because I have a little bit of a legal background, and one of the first things that I thought of the first time, like when ChatGPT-3 came out, I was like, oh, this could totally change discovery.

Like this could, because discovery is like the most painful and most difficult and most expensive. Like you could have serious social consequences by making discovery less expensive. Like that is the expensive part, having a lawsuit. And so like you could actually have significant impact on a society if you could make discovery cheap and instantaneous and reliable.

Guest36:09

Yeah.

Host36:09

I have a question about your graph.

Joel Becker36:12

Yeah.

Host36:13

Cursor.

You missed it.

Here. Whoop, sorry. It was a scatter plot,right? It was what? Cursor in 50 hours. And according.

Joel Becker36:33

Sorry, I see. Yep. Yep.

Host36:37

It'sright there.

Joel Becker36:40

I say it's this one.

Host36:42

Yes, that one.

So you're saying that people, there was no difference. Cursors, we're talking about the vibe coding and they use it for 50 hours. I was very intrigued by that because everyone talks about vibe coding and how Cursor is instrumental.

Why did you get to, how did you get to 50 hours? I was curious.

Joel Becker37:09

So this is including time.

Host37:12

The vibe at 50 hours is.

Joel Becker37:13

This is including time in the experiments that developers have spent in the experiments plus their past experience. So for some developers working on some issues as part of the experiment, some of them have gotten to more than 50 hours of Cursor experience.

And that's used coded up in that bucket at the end.

Host37:35

Was it the same task for each group?

Joel Becker37:37

No, these are kind of their natural tasks that pop up on the GitHub repositories, which I mentioned that are kind of, I don't want to, I'm a little bit nervous about saying they're weird because it implies that I want to say it's very interesting and it's very weird.

And it's interesting for the same reasons it's weird. These are repositories in which they have, these are projects in which they have an enormous amount of mental context built up that the AIs might not have, that they've worked on for many, many years, that they can, I'm not sure this is always the case, but you know, I imagine it in my head that they basically know how to execute on the particular task they have before they even, you know, go about attempting it because they're so expert in the project.

Host38:24

Can you mean when Pa is a speed up? Is it like 5%? Like what do you mean by Pa? How do you quantify Pa's speed up?

Speed-Up Math38:29

Joel Becker38:33

So you might think about, let's go to this one instead. So on the left-hand side, we have the averages for what the developers say will happen in terms of their time to complete if their issue or their task gets assigned to the AI disallowed or the AI allowed group.

You know, they think that if AI is disallowed, it'll take them a bit more time, closer to two hours and I guess more like an hour and a half or a little bit less if AI is allowed. But then, you know, we randomize this particular task to allow AI or not allow AI.

And it turns out, you know, if we randomize to AI allowed, then the times are more like a bit above two hours rather than a bit below two hours. And then you can think of the change in time estimate as sort of being one divided by the other here.

It's not quite that for reasons I can go into, but it's, you know, it's effectively what exactly is the transformation? You know, it's something like AI disallowed over the AI allowed minus one.

So to draw that out, I'm like, you know, I might be like, what's the speed up? You know, is it like 1.1x that, you know, these developers are going 1.1 times faster when we're actually on a time to complete scale, not a speed scale, but ignoring that detail.

You know, is it 1.5x? Is it 0.5x that they're actually going sort of twice as slow? How would we get that information? Well, we'd do something like take the AI disallowed times divided by the allowed AI times. You know, if this was 1.1, let's say times as long as the allowed times, then we'd get to 1.1x speed up.

Something like that that's going on.

And in fact, you know, we find that slowed down.

Host40:34

I just read a fascinating article last coming up, I don't remember, but basically journalists were allowed to

use vibe coding,right? Do a pull request, meaning there was some feature and AI was used to assist with building out the requirements. And she practically, according to the article, just kind of did a little couple of tweaks and then just signed off on it.

It was just really fascinating. That was the whole vibe coding thing.

Joel Becker41:13

Yeah, I've.

Host41:13

You needed coding. Like that was the whole thing. It was like you didn't have any software development background. That was the whole thing. I was just curious. You've tried to do a study on that.

Joel Becker41:27

So I definitely do, I definitely do the share the search, but you know, if you've got like no idea what's going on, then probably these are going to be some significant speed up. You know, I will say, I guess number one, it's not, you know, it's not a priori obvious.

You know, in fact, we went out and did this hackathon with, you know, very experienced people and much less experienced people and tried to see what happened. And what we found is, you know, the scores, the judge scores are extremely noisy and I think you shouldn't believe it.

But, you know, the judge scores were not that much higher when AI was allowed versus when it was not. The people aren't actually making that much more progress. And then another thing to say is, I think there's going to be more expertise in this room than I have.

My understanding from, you know, sitting with these open source developers for a while and not being a very capable developer myself

is that the quality bar on the repositories in this study is just very high.

Host42:28

Yes.

Joel Becker42:28

Typically. And so I would be very surprised if journalists, you know, even frankly, if like a good software engineer without lots of experience on the repository, but certainly, you know, someone who wasn't a software engineer was able to get up a clean PR on these repositories first time.

In fact, I think that's a lot of the story for what's going on here is that the AIs, you know, they actually kind of do make progress in theright direction, some good fraction of the time. But for, you know, for various reasons, sometimes for reasons of correctness, but sometimes for reasons of like, you know, how they've tried to solve the problem and, you know, whether that's the typical way of solving the problem or like how various parts of the project speak to one another, these kind of considerations, you know, they haven't properly accounted for that.

And so, you know, the humans not only need to spend expensive time verifying, but also like clean up, clean up all the stuff. And my sense is that someone who didn't have all that experience like basically wouldn't know how to do that step and so wouldn't be able to submit a clean PR to these repositories.

You know, that said, like I, relative to these people at least, I suck at software development and I'm getting up, you know, PRs internally all the time. And I think their work quality and, you know, and they're getting over time, they're getting better over time.

You know, I do believe that people are coding when they wouldn't be able to code. They are submitting, you know, PRs at a lower quality standard when they wouldn't be able to do that at all. But getting up these expert level PRs, I do feel kind of skeptical.

Host43:57

And that's actually part of what I was getting at is they often get, PRs often get rejected by more novice folks on these bigger quality projects for no other reason other than the developer ergonomics impact of the PR,right?

So the fact that it makes it harder for me to future maintain, because for an open source project, almost all the incentive is biased towards making it easier for me to maintain the project,right? So every time a PR comes in, if it doesn't make it easier for me to maintain the project, I have a tendency to reject it.

Joel Becker44:28

Yeah.

Host44:29

If it does make it easier to maintain the project, then yay, I'm into it. That is unlike what you have in a typical business context,right? Where the most important thing actually is to get something done,right? Because you know, the fact that someone's going to have to spend a lot of time maintaining it's almost job security,right?

But for open source, it's the opposite. It's actually what causes people to leave projects is when it's difficult to maintain,right? So it is a different bias on what you accept for pull requests.

Joel Becker44:57

Can you remind me the name of the English gentleman who maintains the Haskell compiler?

Simon something?

Host45:05

Yeah, no, I can't remember it.

Joel Becker45:07

No? Okay.

Host45:07

I'm not sure what name we call it.

Joel Becker45:08

So here's one story that might be relevant. You know, a bunch of repositories in the study, they all have, you know, broadly these characteristics. One of them is Haskell compiler. Famously on the Haskell compiler, there's like some chance, I don't know if it's 50% or 30% or whatever, but there's some chance that if you submit a PR, the, I'm being recorded.

Host45:28

Simon.

Joel Becker45:31

The.

Host45:31

Simon.

Joel Becker45:31

The Simon something.

Host45:31

Marlowe maybe?

Joel Becker45:32

I'm not sure. The creator of the Haskell compiler will come into the comments and argue with you for many, many hours, much longer than you spent working on the pull request until the PR hits exactly your specifications. Combine that fact with the remarkable fact, I think, that the median PR in this study, the time they spend working on the code post-review is zero minutes.

That is the median PR is like perfect first time around because the professional incentives of these developers are like that. Now, there's a very long tail on one of them. On one of them, I think literally Simon, this gentleman pops up and argues in the comments for many hours and that one's a lot longer.

But yeah, they are maintaining this extremely high bar.

Host46:22

I'm interested in your other upcoming stuff that you had in your doc.

Joel Becker46:24

Yeah, let's do it.

So, yeah, so, you know, so one thing I, what to say? I guess let's go in order. As I think you mentioned, you know, if capabilities are measured by time horizon, keep doubling, it does seem very, very challenging to keep up with that.

Watched Horizon46:30

Joel Becker46:47

In the short term, we have a number of directions for getting on top of that, but, and I think that will last like through the year, but through two years, you know, that seems challenging. I think still possible.

Through three years, I think still seems possible. You know, it starts to get harder and harder. Anyway, in the short term, building these much longer tasks and ways in which we might get around the problem entirely. For instance, here's one thing that might be somewhat.

Host47:15

You could also raise the accuracy bar.

Joel Becker47:18

You could raise the accuracy bar, although, you know, the reason we're interested in this in the first place is we're like, you know, is GPT-5 exponentially dangerous? Okay, and the answer is no, I think.

Host47:31

Yeah.

Joel Becker47:32

But like why do we think the answer's no? Okay, at least I think there are multiple reasons, but at least we can say, you know, GPT-5 is just like not that good at stuff. Like you're trying to get it to do, you know, data science on these columns that have very similar names and it's like not clear exactly the logic that's led to the columns.

It's not doing that sort of thing. I'm like, you need to do that sort of thing to be not exactly that, but you know, be capable of stuff to be exponentially dangerous and it's not capable of stuff. So, you know, I see that the AIs are failing at these hard tasks.

I'm like, you know, great, whatever. But if, you know, if it's succeeding 90% of the time and not 99% of the time on like very challenging long tasks, then, you know, I think there are, you know, other things that we can retreat to and I'm prepared to retreat to those things.

But at least the naive, the reason it's not exponentially dangerous is that it's not capable of stuff. I think it at least starts to get pushed on as an explanation. So here's another possibility for something we could do.

So what if instead of time horizon here, we had time horizon whilst

being watched

very closely?

And then, and then, you know, maybe the time horizon whilst not being watched very closely is like this, but the time horizon whilst being watched very closely is like this. That is like if,

you know, I hope that major AI companies have monitoring setups going where, you know, for example, AIs are looking over transcripts of AIs performing work to see if anything suspicious is going on. And, you know, it's not, it won't be impossible perhaps, a conjecture, to fool those monitors, but, you know, be challenging to fool those monitors.

So, you know, maybe you can do, you know, 40-hour tasks without, 40-hour, you know, dangerous capability relevant tasks without any monitoring, but only like two-hour long tasks if you're being watched very, very closely.

Host49:41

Being watched like if there's another model evaluating the output.

Joel Becker49:43

That's the kind of thing I'm thinking about. Maybe you escalate to humans and.

Host49:46

With and without safety controls, you mean?

Joel Becker49:47

Yeah, yeah, yeah. Roughly, yeah. And then, you know, now we've bought a load more doublings and we can sort of keep doing the capability extrapolation thing. That's one of the kinds of things I'm thinking about in addition to creating ever longer tasks.

Host50:00

Yeah, I mean, I imagine some of the model stuffs do have like evaluations of capability with and without safety because I'm sure that they're like there's an argument between their researchers and their safety teams.

Joel Becker50:10

Yeah, yeah, yeah. Yep.

Host50:15

Seem like I have seen something about this, but not a lot.

Joel Becker50:19

Yeah, yeah. Yep.

Yeah, I guess I think that

this might be sort of like an especially quantitatively important consideration or I expect that it will reduce the effective time horizon by like maybe an order of magnitude or two. Yeah, I agree that there are some important senses in which there's not really a difference in kind.

Host50:46

Yeah. Of course, then I would also worry that like publishing that encourages people to like focus less on safety or to like try to argue against safety because of how it impacts capability.

Joel Becker50:55

Yeah, I think there are lots of landmines

in all sorts of safety work, not just in AI.

Host51:02

Oh, of course.

Joel Becker51:05

Okay, next thing. You know, we have this trend. I spoke about this at the beginning, but you know, we have this trend. Is it going to continue forever? Is this a fact of the universe or does it, you know, somehow depend on inputs or what you think about intelligence exposures or something like that?

Trying to think about that. Where's this line actually going is a pretty active area of work. Also, you know, the ways in which this line or the particular points don't quite correspond to the thing I care about. So one obvious way is that, you know, these models are being judged according to, you know, I think the algorithmic scoring that we use on meets tasks is importantly sort of more robust or more covering the relevant concerns than might be the case in just sort of three benches and the unit tests, but it still sort of, it still has a lot of the same character.

There are, you know, considerations like being able to build on this work in future outside of the immediate problem facing you that aren't being captured by meter scoring. And maybe if you did capture that, you know, you'd get something a little bit like going from 50% success to 80% success.

You know, you could do hour-long tasks if it doesn't matter whether you can build on the work, but, you know, only 30-minute tasks if it does matter whether you can build on the work. But bringing these numbers again to something I care about a little bit more.

And then, yeah, projecting out both if there are compute slowdowns, if we are going to end some regime where AIs are building AIs and that leads to some sort of steepening of the curve, these kind of considerations. That's another thing I'm thinking about.

Beyond Benchmarks52:52

Joel Becker52:52

Oh, and then capabilities measurement from new angles. So here's, you know, here's one history of meter that I think is not the accepted history and also probably not a very accurate history. Certainly not the most accurate history, but here's one possible telling.

You know, near the beginning, meter has early access to when I wasn't there and I have sort of no internal knowledge of this. When meter has early access to GPT-4 and there are just sort of Q&A datasets going on everywhere or like LSAT datasets or something.

You're like, you know, can GPT-4, it like seems so smart relative to stuff that went before. Can it do stuff? You know, so you like you tried out some tasks, can it do stuff? And the answer is, you know, it can do some stuff and it can't do other stuff.

And people are like, oh, that's cool. You know, you've tried this, you've tried this, neat new kind of thing, getting models to do stuff instead of answering questions. And then later you're like, well, different models, you know, they come out over time.

You know, this model comes out in January, this model comes out in February. Can they do different kinds of stuff if we test them on the same, if we test them on the same stuff? Then we'll try and think of kind of the most obvious in some ways, summary statistic of whether they can do stuff.

There's like single data points or a number that reflects whether they can do stuff, the time horizon plus it over time and see what happens. You're like, oh, that's kind of interesting. And then you're like, well, what's the next sort of, in some sense, kind of dumbest or like most obvious thing you can do?

Well, we'll run kind of the most obvious RCT design or like allow AI or not allow AI and then we'll see what happens and we'll try and, you know, it'll be messy. There's lots of, there are lots of methodological problems that people point out as there are with this work, but there are different kinds of problems.

You know, there are different pros and different cons and maybe with these sort of two different things, they give two different answers and have two different sets of pros and cons we can kind of triangulate the truth from that.

And then now I'm like, well, can we pull that rabbit out of that one more time? Are there, or multiple more times, are there other sources of evidence that have, you know, different pros and cons that I won't believe in fully, but they're different pros and cons and they might give different answers and so on and so forth.

Here are two suggestions for things I'm curious about at the moment. The first is in the wild transcripts. So, you know, agents in Cursor and Claude Code and whatever other products or services, they leave behind traces and traces of their diffs that they've contributed to code or diffs of their actions and their recent gains and so on and so forth.

The traces that they leave in the wild are, you know, importantly different from this where it's more kind of contained and, you know, the tasks are sort of neatly packaged and stuff. This is going to be, you know, like the example with the many different columns that are very confusing.

This is going to be like whatever real crap shows up in the wild, how do models handle that? There are important reasons why you shouldn't believe that kind of information. It's like not very experimental. It's like hard to know exactly what to make of it.

But it does have these important pros that it's like, it's more real. It's, you know, the data's enormous. Perhaps the data on transcripts is enormous. You know, perhaps there's a lot you can learn there. That's one thing. And then here's another one.

There's this group which you guys should check out called Agent Village. AI Village, sorry. Where they have a lot of different models or agents kind of living in this village, occasionally talking to humans, trying to accomplish fuzzy goals that are set to them, basically using compute use.

Agent Village56:12

Joel Becker56:27

They try and do stuff like, you know, organize this event in the park or run a human subjects experiment or run this merch store, you know, stuff like that that's not so clearly specified. And basically all the time, they find that the models fall on their faces and suck.

And there are lots of reasons not to believe this evidence. You know, here are some of the reasons. Number one, it is using compute use and I think compute use is just way worse than CLI-based. Compute use capabilities are considerably worse than CLI-based stuff at the moment or text-based things in general at the moment.

And maybe you care more about text-based things because that's more relevant to various types of things you care about. And also lots of GUI, GUI-based things can be converted into text things. It's, you know, there's all these different models hanging around in the village.

I'm like, why are there so many models? Like, why is there a village instead of just like some big agent orchestration set up? I don't really understand what's going on there. Anyway, lots of reasons not to believe it.

But on the other hand, it is models doing stuff in the world. It's not benchmark style tasks. It's like trying to accomplish some goal and they can't accomplish even sort of, you know, very basic subsets of the goal.

And I feel like that's extremely interesting. And I wonder if you could get rid of some of the most obvious cons, you know, make this only text-based, give them some relevant text-based tools, work a bunch on the elicitation to make these models sort of more performant, get rid of the less performant models in the village, so on and so forth.

But then try and get them to do these fuzzy goals and, you know, just observe like where do they mess up. Like, you know, they went about step one, it went great, but then they sort of, they became incoherent or they, you know, went into a strange psychological basin with one of the other models or, you know, they weren't able to interact with external services in the appropriate way or figure out their resource use.

And I'd be very interested just kind of qualitatively in what goes on when you do that. Again, keeping in mind that we're interested in the ability of, at least at the moment, I'm most interested in the ability of AIs to automate R&D and speaking to why that's not the case at the moment and why that might not be the case in the near future.

Some things shaped like this seems like it might be kind of, it might curiously point to why that's not the case. Not sure exactly what's there, but yeah.

Host58:48

And my observation is that they are effectively neurodivergent individuals,right? And none of our world was not built for that. Everything that we have that are defined for a human to do, they're shaped and sized to humans. Just like, you know, the military, like, you know, how big are packs?

Well, it's based on how much they think a person can reasonably carry,right? And how much we expect someone to handle for their taxes, that's based on what we think a human can do. And

if you think about neurodivergent individuals, they struggle with challenges with the way the world's expectations don't align with them. And compared to a neurodivergent individual, these intelligences are really, really different,right? And so all of the rough edges where they don't align with our world, that's why they need an assistant, essentially a human assistant in order to accomplish anything real in our world.

It's just too hard for them currently.

I think someday it'll change, but for now, they're just hopeless,right? I. Have to get really, really good or our world will have to change. One of those two things.

Joel Becker59:54

You know, I agree. I like so strongly share this sense, but, you know, but if you ask me to really pin down like why exactly is that the case again when they're like, you know, beating all the GPQA experts on these extremely hard science questions and they, you know, blah, blah, blah.

Like exactly why are they not able to accomplish things in the world?

Host1:00:14

Have you ever met a neurodivergent individual who wasn't terribly good at something that's completely useless at getting through life?

Joel Becker1:00:20

Yeah, yeah. They're all very good at reading books.

Host1:00:24

There's a lot of those people in the world. It's not that surprising. Although my only feeling about AI village is it's like, well, today is the 200th day my car didn't rocket off the Earth and escape velocity and fly to the moon.

Like that's because you didn't build the rocket yet.

Joel Becker1:00:42

Yeah, I think there was a lot of talk a year ago about, you know, maybe I'm mischaracterizing, but I thought there was a lot of talk a year ago about compute use capabilities being impressive today.

Host1:00:53

There was. There was a lot of talk about it. And yet I have talked to almost nobody who has used them for any practical.

Joel Becker1:01:01

Totally, totally, totally. Yeah, but if we move this to text only and it seems reasonable to complete text only, you know, would you still have the rocket concern?

Host1:01:11

No, I wouldn't have it. I wouldn't really. Well, it depends on what the task was.

Joel Becker1:01:15

Sure. Yeah. Yeah, the kind of thing that you could, that a human could do over CLI.

Host1:01:22

So I think this relates to the Anthropic talk that earlier today where they talked about how, you know, one way to

use perfectly is to give them, if you have a task, like figure out a way to present the task or transform the task to something that is indistinguishable, you know, for the model. And I feel like this conversation kind of, you know, ties in on that.

Like, you know, interacting with Chrome is less indistribution than a CLI. So I think that could be an interesting area of research is like, you know, okay, so if you're interested in exploring like how well can it perform these really open-ended tasks, like first, I guess, creating harnesses and creating an interface that is much more indistribution for them.

So that way that's, you know, less of a concern.

Joel Becker1:02:13

Yeah, I mean, I think also it speaks to the point about quote-unquote neurodivergent models. You know, there's some, it's not so different from management scale or something, giving, you know, giving appropriately scoped tasks to your very talented interns or very talented neurodivergent interns, something like that.

I do think that'sright. From the, sorry to be a, you know,

sorry to be so passive, from the perspective of capability explosions and automating R&D, you know, I think maybe the models will get extremely good at scoping tasks for themselves such that it's benchmark style or something like that. But, you know, if they can't do that, I'm like, well, there's a lot of things that aren't, that don't look like benchmarks that crop up in the real world.

And you do need to be able to kind of flexibly work with that if you want to do something as complicated as automate a major AI company.

And, you know, so I do think it's, yeah, I think it can both be the case that the AIs are incredibly performant on some particular type of problem or if you make other types of problems more similar in scope or shape to the type of problem that they're best at.

And also that they, you know, can't flexibly substitute for human workers because that requires, you know, yourself setting up the problem in a way that's appropriate or not having those constraints yourself.

Host1:03:35

Yeah, it is interesting though, just to your point about new capabilities, is thinking of all these other axes on the graph that you have. Because I think there's not just, I wonder if there's not just a time horizon issue, but there's a task category or a type of work category.

Like as your example of compute, like computer use is one of those examples,right? Like if we think about the capability of computer use versus, or capability that would require computer use versus the capability that could become, can be accomplished entirely in text.

Yeah. So yeah, sure. But like a lot of these are like, like almost all these benchmarks are basically text.

Joel Becker1:04:12

Yes, yes, yes. And indeed, you know, the ones that aren't, the ones that require sort of vision capabilities are notably lacking parts. Yeah, I'm not sure exactly what to make of this graph. I think one thing I make is that, yeah, one thing I make of it is that,

you know, there probably is maybe not so much variation in sort of slope or doubling time across task distributions. I think there's only weak evidence for that, but, you know, in intercepts or, you know, the base of where we are now, yeah, there's possibly a great deal of variety, especially on this sort of

image-like capabilities versus not to mention, but physical abilities even more, you know.

Host1:04:55

Yeah,right. So there's exactly like, so I mean, you could even go through senses,right? Like you could go through like a tactile, like today, like they would all score zero. Nothing has tactile. So like it can't tell you anything about anything tactile.

Joel Becker1:05:10

Well, you know, in producing this graph, we, you know, we try and make the models as performant as possible on some held out set. So we, you know, we try and give them some tactile stuff. I'm not sure they perform zero.

Host1:05:22

Sure, sure, sure. I mean, space, we do have some examples.

Joel Becker1:05:26

Yeah.

Host1:05:27

Yeah, yeah. Space judgments, spatial judgments, things like that.

Joel Becker1:05:31

Yeah.

Host1:05:32

You know, we've obviously seen from figure fine control and stuff like that with other robotics.

It's just, I haven't even, I don't even know if anybody, maybe somebody has listed out what all of the capabilities that we would expect in the future. Like if we actually wanted AGI, what is the entire list of capabilities?

That's a way to start a debate that doesn't end.

Joel Becker1:05:55

I think it's. Basil Halperin and Arjun Ramani, hopefully have a paper on this in a small number of months.

Host1:06:02

Yeah. And then we should think about where are we at and do all the capabilities follow the same? All the capabilities that we currently measure, do they follow the same log?

Joel Becker1:06:12

Yeah. It does seem like a reasonable null hypothesis to you as well as me, I think. Not certainty. I mean, who knows? Yeah, yeah.

Oh, there was something I wanted to add there.

Oh, yeah. Here's another thing I'm thinking about, not super in a research capacity, although kind of.

Physical Limits1:06:31

Joel Becker1:06:38

So, you know, some people like me are sort of skeptical of software-only singularity. That is the idea that you could automate AI research without also automating

chip design and maybe also chip production as well, that you'd quickly get bottlenecks by compute. So there are only, for fixed hardware, there are only sort of so many experiments that you can run that will be sufficiently productive to boom progress upwards.

But, you know, even for people like me who are skeptical of that,

you know, you might think that in fact, like chip production is going to get automated. You know, the robots, like they're coming. They can do the stuff that humans do. And then maybe you really do have a fully self-sustaining

robot plus AI economy. And so, you know, and so you have some slow trend from compute slowing down, but then you have sort of a booming back up once the whole thing is in a tight loop. One interesting debate that I heard about recently and would like to think more is, you know, I think there's, in the public discussion, there's some sense that, you know, why are robotics capabilities lagging?

Lagging LLM-like capabilities so much? Well, it's through with training data or something like that. Or maybe it's through with hardware constraints. I'm curious if it's not through with hardware constraints. What exactly are these hardware constraints? If we put superintelligence inside, hypothetical superintelligence, inside of, you know, hardware parts that existed today, could it build chip production facilities?

And I have no idea because I'm, you know, I'm beyond, beyond, beyond novice. But it's not obvious to me what the answer is. I think it's kind of plausible. I'm not sure you need this, like, yeah, I'm not sure you need this, like, very flexible fine motor control in order to do it.

Also, I think maybe the fine motor control is there, subject to having superintelligence controlling it.

Host1:08:38

I mean, to be fair, like the key aspects of chip production are done by robots.

Joel Becker1:08:46

Oh, but I'm also thinking like building the robots and the whole, you know.

Host1:08:50

And that's where I'll tell you, I have a friend who spent most of his career doing software development, but during COVID started working on manufacturing things like Peppers and things like that to help people. And he found out how hard the manufacturing world is and how slow the iteration process is.

And it is really, like he put it, like he knew it was going to be worse. He didn't understand that it was like next level, like an order of magnitude worse. And I think that probably, like, you know, from our perspective, people who don't do it, it seems like, oh, how bad can it be,right?

The feedback I've had from everybody who actually works in that space is it's way, way different.

Joel Becker1:09:29

That's what I've heard as well. I've only talked a little bit with, like, people who work in fabs and stuff, but I was surprised when I did talk to them of the level of human expertise required in order to work at the fabs.

Like a lot of those jobs are like fairly high-paying engineering jobs in order to, like, successfully.

Host1:09:45

Also, the rate of improvement is actually glacial,right? Compared to software,right?

Joel Becker1:09:50

I think also because it costs a billion dollars to build a fab.

Host1:09:53

That's what I'm saying. Like each iteration is a huge cost of time, money. It's brutal. So it's, I think that's why it's been hard to get it all the way there is just like, give them a couple more centuries, maybe they can get it done.

Joel Becker1:10:07

Is that really your view? Centuries, centuries.

Host1:10:09

I do. I do think, I'm skeptical like you about how easy some of these tasks are. We think they're easy, but in my experience, like, I remember when the self-driving thing came out, when people were like pushing that and it was, I actually worked in that space for a while and it was like, I get that we can get really close to it, but getting all the way to something that is acceptable is extremely difficult,right?

And we underestimate how much work is involved in getting that last little bit done. The first 90%, I knew we could do it with computers like, you know, 10 years ago, pretty much, but getting the last bit that everyone's happy with it, you need a lot of work.

Joel Becker1:10:50

I feel this myself, you know, I didn't get a driver's license when I got something because I expected self-driving cars to come. Yeah, I think, I think totally, but it hasn't been that long, you know? And they're expanding to the entire bay.

Host1:11:06

They're going to get there. I don't think it's going to take a hell of a lot.

Joel Becker1:11:08

Is the robot economy building the chip production going to take centuries?

Host1:11:12

I don't know about, well, I can see that it might take, so part of the trick with self-driving is the economic incentive is moving it along faster,right? And probably the robot building robots kind of thing would also, but like, you know, where we're atright now is like ripwreck is kind of as far along as we've got of robots building robots,right?

Which is.

Joel Becker1:11:35

Oh, but I feel like, you know, is that paying sufficient attention to the charts? GPT-2, 2019. It's so recent, you know, I have some, this is so

nonsensical, but I'm like, maybe we're in a sort of GPT-2 moment.

Host1:11:56

Yeah, no, it's a fair point. I could be wrong. It's just my guess is it's going to take a lot longer than we think. At least to be able to do like real mass production

at a scale that causes the kind of global impact that you're talking about. I think they can already do a great job building one-offs,right? Robots are very good at doing one-off builds at a small scale, but it's totally impractical for doing it at a large scale.

Joel Becker1:12:26

There is, um,

one facts that I think is kind of remarkable. Is this, maybe it's this, is that the rate of, is it this? Yeah, yeah, yeah. The rate of compute put to robotics models lags behind, sorry, is about the same, but the levels two orders of magnitude difference.

I am kind of curious if that gap closed,

what we'd see. It does seem like at least sort of more capable robots are, in some sense, very on the table as something that could be the case very soon if this, I'm not saying all the way, I'm certainly not saying chip production.

It just does seem like there's some sort of overhang.

Host1:13:23

Yeah, yeah.

Joel Becker1:13:25

Intuitively.

Host1:13:26

That's interesting.

Joel Becker1:13:29

Also thinking some sort of, some like you don't just need to be scaling data, you can also scale parameters, use the same amount of data, you know, flexible ways to use compute to close some gap.

Host1:13:43

Interesting. So you wanted to just give me a very interesting overview of where AI is going into fabrication at the very time.

Joel Becker1:13:52

Dan was the same.

Host1:13:54

So it says there's a lot of ways whereright now it's going to help probably pretty dramatically in the near future and a lot of it's in computational aspects. There's a lot of computational aspects that are extremely expensive designing like a mass, basically the hole that you're using for the laser to get the transistors and like calculating that, how to build it and ensuring that it conforms to the spec that you've written basically is extremely computationally expensive.

And there's a lot of opportunity for AI to help there. And there's also theoretically the possibility for, so like chip, obviously chip manufacturers are extremely precise, but also fragile. And the opportunity for an AI to detect parameters that are basically out of whack and leading to failure, potential failure in like imaging a wafer

is, could theoretically dramatically improve yield and yield is a big problem in chip manufacturing. Like the reason that you get different speeds out of your CPUs is because they actually just have the one line that produces all those CPUs and some of them come out better and some of them come out worse.

And that's why the higher gigahertz models are more expensive than the lower gigahertz. Like if you have like your NVIDIA, like your home GPUs, your 5040, your 5050, your 5060, your 5070, your 5080, your 5090 are all the same chip.

That just had different quality levels.

Joel Becker1:15:20

Different levels of functionalities essentially.

Host1:15:22

Yeah. But the problem is that.

Joel Becker1:15:28

Sorry, I'm going to cut the recording. They're going to kick us out soon, but feel free to continue the discussion.

Host1:15:32

Yeah, cool.

Joel Becker1:15:32

You can also hang on, but I'm just going to.