Coding Origins0:00
Well, hello, hello. Is the mic working for you?
Check, check, check. One, two, three. Allright, first hard technology problem of the day down.
Yeah, yeah. Well, the Wi-Fi is the other one. Everyone here knows. So, Greg, welcome to AI Engineer. Thank you so much for taking the time.
Thank you for having me.
We're going to go a little bit chronologically, and a lot of people are sending questions, and I've sort of grouped them up for you. So we're going to just getright into it. So, you know, I did some deep research on you.
You literally did deep research?
With deep research. I call it peep research because we're researching a person. You actually did theater growing up, and chemistry, and math, and you wrote a calendar scheduling app, and that's what got you into coding. But, like, what really inspired your love for coding?
Like, why are you the coding guy?
Well, the funny thing is, I thought I was going to be a mathematician when I grew up.
Yeah.
You know, I'd read about people like Galois and Gauss. You know, we were working on these, like, 100, 200, 300-year time horizons. And I was like, that's what I want to do. If anything that I come up with is ever used while I'm still alive, it wasn't long-term enough.
It wasn't abstract enough. And I was writing this chemistry textbook after high school, sent it to one of my friends who'd done something similar in math, and he said, no one is going to publish this. You can either self-publish I was like, oh, that sounds like a lot of work, a lot of capital or you could make a website.
And I was like, guess I'm going to learn how to make a website. And so I literally went on W3Schools and did their PHP tutorial. How many people here remember W3Schools? Yeah, a decent number of hands. And I remember the very first thing I built was a table sorting widget,right?
I had this picture in my head of what it would be, and I remember the moment that I clicked the column and it sorted according to that column, which was exactly the thing that I wanted. And I was like, that was magic,right?
And I was like, this is so cool. Because the thing about math is that you think hard about a problem, you understand it, you write it down in an obscure way you call it proof, and then, like, three people ever care,right?
But in programming, you write it down in an obscure way we call it program, and then maybe only three people ever read that program and care about the code, but everyone gets the benefit. No one has to understand the details.
That thing that was in your head, it's real, it's in the world. And I was like, that, that's the thing I want to do. Forget about that 100-year time horizon. I just want to build.
You do just want to build. It's so you were so good at it that somehow, somewhere, you got cold email by Stripe while you were still in college.
That'sright.
Joining Stripe2:57
What's the story of that? First of all, how did they find you, and what was it that convinced you to drop out to join them?
Well, so I had mutual friends with all the people at Stripe, the, you know, giant company of, like, three people at the time. And they'd asked, you know, it's the usual thing where they'd ask someone at Harvard who the, you know, people around campus to talk to, who they might recruit where.
My name came up. They asked the same for the people at MIT because I actually had dropped, I'd been at Harvard and actually dropped out to go to MIT. So I had the advantage of, I guess, you know, getting upvotes on both sides.
But I remember when I met Patrick and it was, you know, I'd just flown in, it was, like, late at night, you know, it was storming, and I showed up and we just started talking about code,right? And it was just, like, one of those moments where, like, this, this is the kind of person that I've wanted to work with and been looking for.
And so I ended up dropping out of MIT and, you know, flew out and been out here ever since.
Yeah, yeah. We have a special we have some guest questions sprinkled along the way, as you know. So a guest question from someone named Matthew Brockman.
I've heard of him.
CTO of Julius AI. When do you think our parents will give up on the dream of you finishing your degree? Maybe, maybe Harvard or UND will take you back.
Yes. Well, never. It was definitely, you know, I think it was no matter where you're going, if you tell your parents you're leaving Harvard, it's going to be hard. You tell your parents you're leaving school altogether, it's going to be difficult.
And I think that, you know, it was actually, to their credit, you know, I think even though it was difficult, they were like, you know, we trust you. Like, you must see something and understand something from where you sit that's hard for us to see from halfway across the country.
But, yeah, I think that as, you know, I did Stripe and had a good time and actually learned things and turned out it was a real company and not just, you know, just dropping out and doing nothing, I think that they really were, you know, have warmed up to it.
And so.
I think they're very proud of you.
Yes, absolutely.
So you were with Stripe from 40 to 250 people as the first CTO, eventually. One thing I found recently that Hacker News maybe doesn't know is apparently the call list and installation only happened, like, a handful of times.
It wasn't, like, a thing at Stripe. Was that?
That's, I think that's true. Yeah, it is the thing that, you know, it's, like, survived the.
It's an urban legend because it's, like, so cool. It's, like, you get so customer obsessed. Anyway, so what else do people get wrong about early Stripe? Like, why do we want to clear the air?
Yeah. Well, I think people don't understand how hard it was,right? It was just, like, I remember, you know, first of all, the kind of thing that we did a lot of is that we added all of our customers on Gchat.
And so it was very much the case that we were in constant contact with them. And so even if you're not literally sitting over their shoulder, you're doing the next best thing. But I remember, like, one, you know, one day we realized that, you know, the payment backend that we were on, it just wasn't going to scale.
We absolutely needed to be on Wells Fargo, and we got sort of the deal done, but now we need to do a technical integration. And they said, well, this technical integration is going to take, like, nine months because that's how long it takes.
And we were, like, that's crazy. Like, you're a startup. Like, we can't sit around waiting nine months to get this thing done. And so actually, in 24 hours, we completed it by just basically treating it like a college problem set.
And it was, you know, I was implementing everything. John was working from the top of this test script and testing everything and being, like, this is broken. Dara was starting from the bottom and working his way up. And in the morning, we got on with the certifying person, and we sent some test messages, and there was an error.
And the person was, like, allright, I'll see you next week. Because that's how all their customers operate,right? There's an error, like, you know, clear. You need to send it to your dev team. And we were, like, no, no, no, there must just be, like, some sort of glitch in the system.
Like, and Patrick was just, like, talking to keep her on the line and frantically, like, I was there editing the code. And so we got, like, five turns in, and we actually failed. But fortunately, she was nice enough to reschedule two hours later, and then we passed.
And so you realize that was, like, six weeks' worth of normal dev work that you got done in that moment because you didn't just accept the, like, arbitrary constraints of how other organizations would work.
Yeah, yeah. I think there's a do you think there's a lot more opportunity like that in most jobs? Like, how do you advise other people to be that, I guess, fast or, like, to cut that many cycles?
Yes. I mean, I think that the way I think about it is that if you think from first principles, you can find where things need to be slow or done the way that they're normally done or whatever those things are.
Those exist,right? The general principle of, ah, just don't worry about the constraints and just do the thing, I think that that is not 100% true. I think it's really about mapping to where is there unnecessary overhead that's there for constraints that are no longer applicable, that don't apply to your specific circumstance.
And I think this is especially true in this world that we're in now with AI that's accelerating productivity so much.
Yeah. Just fire off a Codex. Why not,right?
Yeah.
Independent Study8:00
One thing, one last thing about your sort of pre-OpenAI life was independent study. I just, I found that just it's a recurrent theme from high school. You did RecurCenter?
I did.
And you're sabbatical as well. Just, like, you've just done it repeatedly. What makes independent study effective? Like, I think there's a lot of people who don't do a good job of it and kind of waste a year. What do you do that makes it so effective?
Well, I think it was a key part of how I grew up. You know, in sixth grade, my dad taught me algebra, and in seventh grade, we showed up at the high school is the first time that you track into advanced math pre-algebra.
And we went to the teacher, like, can he skip this and go directly to the eighth grade course? And the teacher looked at my mom and me very condescendingly and was like, every parent believes that their child is special.
And after, like, a month of being in this teacher's class and, you know, I was paying no attention and just doing, you know, calculator games in the back, and she'd try to trip me up and, you know, call on me to answer questions from the whiteboard, and I would just get them allright.
She was like, allright, like, fair enough. Your child should be in the next year. But then when I was in eighth grade, there was no more math left in my middle school. I didn't have a car, so I had to do online courses.
And in that one year, I ended up doing three years' worth of high school math. And so I think that, for me, a lot of it is about suddenly these, if you're excited about something independently, something you want to do, that you can break the constraints there as well.
You can do three years of math in one year, and then it compounds because the next year I was out of my high school, finished math there, and then all through 10th, 11th, and 12th grade, I had, you know, no more math, so I did have a car, and I was able to go to the University of North Dakota, take whatever classes I wanted there.
And so I think that that kind of compounded, compounded, compounded to learning programming. And then I think that the way I learned programming was very much self-study, just building things and experiencing things out in the world. And so I think that the thing I would just advise is, like, if you have an opportunity to explore and you have a passion, you're actually enjoying it, just go deep,right?
And by the way, it's not always fun,right? I think that it is very easy to get kind of, you know, sort of feel like, ah, I got kind of bored. But if you just push through those hurdles, then I think that the reward is worth it.
Yeah. You self-studied machine learning too.
Machine Learning10:18
That'sright.
Like, that was a whole period of your life. Any particular highlights from there? It sounds like you talked to Jeff Hinton at one time.
I did talk to Jeff Hinton.
Yeah.
Yes.
And, like, what was, you know, did that help or what was the most helpful thing? Like, you became a machine learning practitioner.
Well, so when I started out, so, you know, I'd been at Stripe. I was reading Hacker News posts about deep learning. And, yeah, it was like, you know, there's a deep learning for X, like, every day it felt like.
And this was, you know, 2013, 2014. And I was like, what is deep learning? And I knew, like, one person in the field, and so I talked to them. They introduced me to some more people, and then they introduced me to more people.
And the thing that surprised me was I kept getting introduced to a bunch of my smartest friends from college. And I was like, that's interesting. All of these people ended up in this field? Like, what's going on? And I started to realize that there was something real that was building,right, that was being developed, that people were really making these systems do material new things that computers were not able to do before.
And I was like, that, that is the thing. And so after I left Stripe, you know, I knew I wanted to do something in AI, start an AI company, but I didn't really know how to contribute what my skills would be useful for.
And so I was in New York, and I was like, you know what? I'll build a GPU rig and see if I can do some Kaggle competitions. And so I went on Newegg and just, like, you know, bought some Titan X cards, and it was really cool, you know, physically assembling this machine.
And you can find some tweet from 2015. When I powered it on, you see all this, like, green and all the fans going. And I was like, this, this is what computers are meant to be.
I think many folks in the audience have that exact experience as well. Awesome. OK, so what convinced you that AGI was possible? Like, you had a point where you were sort of disillusioned with it. You wrote, you tried to write a chatbot.
It didn't work. But what made you go all in on it?
Yeah. Well, so, you know, part of the journey for me was reading Alan Turing's 1950 paper, Computing Machinery and Intelligence. This is the Turing test paper. How many people have read it?
Read it.
Fewer hands than W3 schools, but equally as important, worth reading. The thing that is so fascinating to me is he lays out in the beginning, OK, Turing test, this idea of just does a machine think? Is it intelligent?
And you can say it's intelligent if, you know, a human can't tell the difference between talking to it and talking to a human. Fine. But the thing that was, that has not really become as embedded in the pop culture, but to me was so astounding, was he said, well, how are you going to program an answer to this?
You will never be able to write down all the rules. But what if you could build a child machine that learns like a human child, and then you just apply rewards and punishments, and boom, it's going to be able to pass the test.
And I was like, that, that is the kind of technology that we have to build because as a programmer, you have to understand everything. You have to understand the rules of how to solve the problem. But what if the machine can understand things and solve problems that you yourself cannot understand?
Like, that feels fundamental,right? That feels like how you actually solve problems that are important to humanity. And this was, you know, 2008.
Message that I read that he was a NLP professor, and I asked if I could do some research with him. And he said, yeah, here are some parse trees. And I was like, OK, this is not what Turing was talking about.
Yeah.
This is like WordNet and.
The whole thing. Exactly. So it's like, you know, definitely a little bit of Trough of Sorrow there. But with deep learning, the thing about deep learning that's magic is that, you know, it really started in to show promising results in 2012 with AlexNet,right?
And that it just blew everyone out of the water in the ImageNet competition. And so suddenly you have this, like, general learning machine. You know, it's got a little bit of a prior in there of convolutions, but it's better than 40 years' worth of computer vision research,right?
People trying to write down all the rules as well as possible. And then people are like, well, OK, it works in vision, but it's never going to work in my field. It's never going to work in machine translation, never going to work in, you know, in NLP, never going to work in this or that.
And suddenly it starts being the best in all of those areas. And suddenly the walls between these departments are being torn down. And you're like, that, that is what Turing was talking about. And so I think for me, just seeing the type signature of this technology and by the way, this technology is not new,right?
Neural nets were really, like, if you go back and read the McCullough-Pitt's neuron paper from, like, 1943 or so.
I told people, I told him to give homework to peoples.
OK. Yeah, there you go.
Yes. Classes assigned.
The images in there, they look just like the kinds of images that you see now of just, like, you know, layers of neurons and things like that. And so you just realize there's something deeply fundamental about what we're doing.
And you can find these, you can find this paper from the 1990s talking about what caused the deep learning winters and that it was these NeuralNet people. They have no new ideas. They just want to build bigger computers.
And I'm like, yes, that's what we need to do. And so I think that all of this together just feels like we are, to some extent, continuing this wave of 70-year history. And in many ways, you know, the whole computing industry has been really trying to build up to the point that you can have machines that are able to perform the kinds of tasks that we're just starting to scratch the surface to solve new problems that humans cannot, to be assistive to us in our daily lives, to not have to, you know, be typing with our, you know, meat sticks, but instead to have something that you can interact with just like a person where the machine comes much closer to you rather than you closer to it and having to learn assembly language or,
you know, whatever it is. And so to me, it felt like all of the factors were lined up, and now we just need to build.
Eng & Research16:08
Yeah. I like that consistent theme that you keep coming back to, we just need to build. So in 2022, you wrote that it's time to be an ML engineer. Actually, I have a personal friend who read that post and called the email to you and joined OpenAI and all that.
You said that great engineers are able to contribute at the same level as great researchers to future progress. Is that still true today? You know, I think a lot of engineers look at the researchers who are making millions of dollars, and they're like, how do I contribute as much, you know?
I think it's absolutely, if not even more true. I think that, like, if you look at the phases of deep learning research since 2012, I think at the beginning it really was, and this is kind of what I expected when we started OpenAI, you know, just like research scientists who had gotten a PhD who would go and kind of come up with ideas and test them out.
And, you know, there's engineering to be done. If you actually look at AlexNet itself, you know, it's fundamentally the engineering of, let's get fast convolutional kernels on a GPU. And fun fact is people who were in the lab with Alex Krusevski at the time actually felt very bad for him because they were like, he has some fast conf kernels for, you know, some image data set that doesn't really matter.
But, you know, Ilya was like, well, clearly we just need to apply this to ImageNet. It's going to be great,right? So it's like the combination of great engineering together with the idea of what to do with it,right? That that's what makes the magic work.
And the thing that I think is still true and even more true is, OK, so the engineering required, it's now not just let's build some kernels, but let's build a system where let's, like, actually scale to 100,000 GPUs.
Let's actually, you know, sort of do this crazy RL system that orchestrates things in all sorts of ways. So the idea, if you don't have the idea, you're dead in the water. There's nothing to do. But if you don't have the engineering, that idea is not going to live and see the light of day.
And so you need to have both of these coming together harmoniously.
Yeah. I think that Ilya-Alex's relationship is really emblematic of, like, the research-engineering partnership that now is the philosophy at OpenAI.
That'sright. Yeah. And if you look at how OpenAI operates, like, I think from the very beginning, we had this ethos of engineering and research be valued and work together as partners. And I think that that is something that we, you know, it's like something that we really work at every day.
Yeah. It's my explicit goal to try to throw curveballs in this stuff. So in terms of the relationship between engineering and research, what did OpenAI do wrong in the early days that you do well now?
Well, I think that the relationship between engineering and research, the way I think about it, is you never fully solve it,right? You just sort of solve the current level of problem, and then you move on to the next level of sophistication.
And I noticed that actually the kinds of problems that we ran into were basically the same problems that had been run into at every other lab. And it was just like, you know, either we would be further along or there would be a slightly different variant of it.
And so I think there's something deeply fundamental about this. So at the very beginning, I could really see people who came from the engineering world, people came from the research world, just sort of thinking about system constraints very differently.
And so as an engineer, you're like, hey, if I've got an interface, you should not care what's behind that interface. We agreed on the interface. I can implement it however I want. Whereas if you're a researcher, you're like, if there's a bug anywhere in the system, all I'm going to get is just slightly degraded performance.
Not going to get an exception, not going to get indications of where. And so I am responsible for understanding everything. The interfaces, they don't matter. Unless they're, like, truly rock solid and I can just, like, never think about it, which is a pretty high bar, then I am actually responsible for this code.
And that causes friction,right? Because then how do you actually work together? And I saw a project very early on where the, you know, the people from the engineering background would write the code, and then there'd be this big debate over every single line.
And I was just like, this is never going to move. It's going to be so slow. And instead, the way that we ended up proceeding was, so I actually worked on that directly, and I'd come up with, like, five ideas at a time.
Someone from the research side would say these four are bad. I'd be like, great. That's all I wanted,right? And so the value that I think we've really realized is critical and that I tell people from the engineering world coming into OpenAI is technical humility,right?
It's like you're coming in because you have skills that are important, but it's a totally different environment from, you know, something like a traditional web startup. And figuring out when those intuitions apply and figuring out, like, when to leave them at the door is super hard.
And so the most important thing is to, like, come in, really, really listen and kind of assume that there's something that you're missing until you deeply understand the why. And then at that point, great, make the change, like, change the architecture, change the abstractions.
But I think that that kind of approach of just really, really read and listen and understand with that humility, that that is, I think, a really key determiner.
Yeah. Awesome. We're going to tell some stories from recent launches of OpenAI, the greatest hits. So one of the things that is kind of interesting is just scaling in general. Everything breaks at different orders of magnitude. So when ChatGPT launched, you got a million users in five days.
Scaling Launches21:04
This year, when 4.0 ImageGen launched, you got 100 million users in five days. How do those two periods compare?
They echo very similarly in a lot of ways. You know, the thing about ChatGPT, it was supposed to be a low-key research preview, and we put it out very, you know, sort of chilly. And then suddenly everything was down.
And we, you know, we kind of anticipated that ChatGPT would be a very popular thing, but we thought that GPT-4 would be necessary to get it there.
Had it internally as well. So you just weren't impressed by.
Exactly,right? It's like that's the other thing about this field is you update so quickly.
Yes.
Right? It's like you see magic and you're like, this is the most amazing thing I've ever seen. And then you're like, well, why can't it, like, you know, why can't it, like, merge, you know, 10 PRs for me?
Exactly. And the ImageGen moment was very similar in terms of it was just so, so loved and so popular. And it just went viral in ways that, you know, just like the numbers were just off the charts. And so internally, we actually did something that we really, really try not to do, which is we pulled a bunch of compute from research for both of these launches, actually, because that's mortgaging the future to make the system work.
But if you can actually deliver and keep up with demand, then, of course, people get to experience the magic. And I think that that's something that is really worthwhile, and it's really important to sort of, you know, maximize those moments.
So I think that we really have that same ethos of really serving the user, really trying to push for the technology and just do things that are materially new that no one's ever seen before, and then whatever it takes to get those out into the world and make those successful, that that's what we do.
Amazing. Well, I mean, incredible job. GPT-4 launch. So I am told. Wife drew the joke website.
That's true. Yeah. Fun easter egg. My handwriting was so bad that even our AI couldn't tell what to do with it.
So like, apparently, did you improvise some of this? I heard.
I.
Great vibes.
Yeah, definitely. Definitely. Like, you know, usually when I do these kinds of demos, like, I've tested the general shape of them ahead of time, but I've always had to, like, it's very easy in this field to have ones that are just like, if you slightly typo a character or something, then the demo will not work.
I don't like doing those. I like to have some robustness to it. So there's always variation in terms of what actually ends up being shown.
To me, this was the first time I think the world ever saw Vibe Coding. It is now a thing. What are your thoughts on Vibe Coding?
Vibe Coding23:43
Well, I think that Vibe Coding is amazing as an empowerment mechanism,right? And I think it's sort of a representation of what is to come. And I think that the specifics of what Vibe Coding is, I think that's going to change over time,right?
I think that you look at even things like Codex, like, to some extent, I think our vision is that as you start to have agents that really work, that you can have not just one copy, not just 10 copies, but you can have 100 or 1,000 or 10,000 or 100,000 of these things running.
You're going to want to treat them much more like a coworker,right? That you're going to want them off in the cloud doing stuff, being able to hook up to all sorts of things. Your sleeper laptops closed, it should still be working.
I think that the, you know, current conception of Vibe Coding in an interactive loop, you know, that that's something that I think is like, you know, it's, OK, so my prediction of what will happen is, like, I think there's going to be more and more of that happening, but I think that the agentic stuff is going to also really intercept and overtake.
And I think that all of this is just going to result in just way more systems being built. And the thing that I think is also very interesting is that a lot of the Vibe Coding kind of demos and the cool, flashy stuff, for example, making the joke website, it's making an app from scratch.
But the thing that I think will really be new and transformative and is starting to really happen is being able to transform existing applications and go deeper and be able to, you know, like, I think so many companies are sitting on legacy code bases and doing migrations and updating libraries and changing your COBOL language to something else is so hard.
And it's actually just not very fun for humans. And I think we're starting to get AI that are able to really tackle those problems. And so the thing that I love about where Vibe Coding started has really been like with the most, like, just like make cool apps kind of thing.
And it's starting to become much more, like, serious software engineering. And I think that going even deeper to just, like, making it possible to just move so much faster as a company, that's, I think, where we're headed.
Yeah. Speaking of Codex, I've heard that you've just, it's kind of your baby a little bit. And you've started, I think on the livestream, you were talking a lot about just make things modular and well-documented and all that good stuff.
Codex26:05
Like, how do you think Codex changes the way that we code?
Well, I definitely think that it's an overstatement to say it's my baby. Like, I think that there's a really incredible team and that, you know, I've been trying to support them and their vision. But I think that the direction is something that is, like, just so compelling and incredible to me.
The way that, and sorry, could you repeat the.
How does Codex change the way that we code?
I see. Yeah. The thing that has been most interesting to see has been when you realize that the way you structure your code base determines how much you can get out of Codex,right? That if you match the strength of, like, basically all of our existing code bases are kind of matched to the strengths of humans.
But if you match instead to the strengths of the models, which are sort of very lopsided,right? Models are able to handle way more, like, diversity of stuff, but also are not able to, like, sort of necessarily connect deep ideas as much as humans areright now.
And so what you kind of want to do is make smaller modules that are well-tested, that have tests that can be run very quickly, and then fill in the details. The model will just do that,right? And it'll run the test itself.
And the connections between these different components, kind of the architecture diagram, like, that's actually pretty easy to do. And then it's the, like, filling out all the details that is often very difficult. And if you actually do that, you know, what I described also sounds a lot like good software engineering practice.
But it's just, like, sometimes because humans are capable of holding more of this, like, conceptual abstraction in our head, we just don't do it,right? That, like, yeah, it's like, you know, it's a lot of work to write these tests and to, you know, to flesh them out.
And that, you know, the model's going to run, like, these tests, like, 100 times or 1,000 times more than you will. And so it's going to care, like, way, way more. So in some ways, the direction we want to go is build our code bases for more junior developers in order to actually get the most out of these models.
Now, it'll be very interesting to see as we increase the model capability, does this particular way of structuring code bases remain constant? And I kind of think that it's a pretty good idea because, again, it starts to match what you should be doing for maintainability for humans.
But yeah, I think that to me, that the really sort of exciting thing to think about for the future of software engineering is what of our practices that we've kind of just cut corners for, do we actually really need to bring back in order to get the most out of our systems?
Yeah. Can you put numbers on, like, ballpark numbers on the amount of productivity you guys are seeing with Codex internally?
Yeah. I don't know what the latest numbers are. I mean, there's definitely double-digit percent of our PRs are written, low double-digit, written entirely by Codex. And that's super cool to see. But it's also, like, you know, that it's not the only system that we use internally.
And I think that to me, it's still in the very, very early days. It's been exciting to see some of the external metrics. Like, I think we had 24,000 PRs that were merged in, like, the last day in public GitHub repositories.
And so it's just like, yeah, the stuff is all just getting started.
Scaling Bottlenecks29:20
Yeah. It's doing a lot of work. Guest question from Dylan Patel on scaling and reliability. So we're doing more tasks that take longer and utilize more GPUs. They're also just unreliable. They fail a lot,right? And this is just well known.
So this causes training to fail as well. So, like, but, like, you know, you've mentioned that you can sort of just restart a run and that's OK. Like, how do you deal with this when you have to train long-horizon agents,right?
Because you can't really restart something that has a trajectory that's kind of halfway that is maybe non-deterministic.
Yeah. I mean, I think that there's a bunch of problems that you kind of solve and then you make the models more capable and then you have to resolve them. And so, yeah, when the rollouts are short, you know, 30 seconds, you kind of don't care that much about this problem.
If they're going to be days, now you really care about this problem.
Yep.
And you have to start thinking about how to snapshot state and a bunch of things like that. The short answer is that I think that there's this, like, ladder of complexity that you keep climbing with these training systems.
And it goes from, you know, like, a couple of years ago, all that we cared about was just doing good old-fashioned free training,right? And that's, like, very check-pointable. And even there, it's not trivial,right? It's like, you know, if you go from checkpointing once in a while to, like, you want to checkpoint every single step, now you need to think really hard about how you're going to avoid copies and blocking and all these things.
Then for something like these more complicated RL systems, there's still checkpoint ability in terms of, you know, maybe you care about, you know, checkpointing your cache so you don't have to recompute everything. And the nice thing about our systems is that, you know, language models, their state is very explicit,right?
And it's something that actually can be stored, something that you actually can handle. Whereas if you have tools that you're hooked up to that are themselves stateful, maybe those are not something you can restart and recover from. And so I think that if you consider the whole system end-to-end, thinking about what checkpoint ability looks like, and there's also a question of maybe it just doesn't matter,right?
Maybe it's fine that you restart the system and you get some little wiggle in your graph, but these models are smart,right? That they can handle it.
One thing we're looking at tomorrow that's launching is maybe you can sort of take over the VM and checkpoint the VM state and restart it.
Yep.
I think we have a dial-in call-in question from Paris. If someone can play the video from the special guest.
Oh. I wish I could be there to ask you in person. One of the questions that I have is, in this new world, the workloads in the data center and the AI infrastructure is going to be incredibly diverse.
On the one hand, agents that are doing deep research and they're thinking, they're reasoning, they're planning, and they're working with other agents, and they're, you know, working on a lot of memory to have large context. On the one hand, some of it you also want to think as fast as possible.
So, you know, how do you create an AI infrastructure that is optimized for workloads that have a lot of prefill, a lot of decode, a lot of something in between on the one hand? And on the other hand, the type of workloads that I'm super excited about, these multimodal vision and speech AIs that are essentially your R2D2, your companions on all the time.
It's instantly available to you. And so these two workloads, one of them super compute-intensive and might take a long time and, you know, test time scaling and all that, on the other hand, wants to be very low latency.
So what does a future AI infrastructure look like that's as flexible as possible, as performant as possible, low latency, high throughput? You know, all of that is just incredibly complex. So how do you think through that and what kind of an AI infrastructure would you think would be ideal going forward?
Well, lots of GPUs, of course.
So if I were to summarize, Jensen wants you to tell him what to build.
What would be your dream? But also, like, there's just two needs. There's two kinds of infra. There's long compute and there's real-time now, now, now.
Yes. Yes. I mean, it is hard,right? Because, I mean, this code design problem, it is a mind-boggling one. And so, you know, I'm a software person by background and that, you know, we think we're off here just, like, writing the software for AGI and then you realize you have to do, like, these massive infrastructure projects,right?
Like, that's not how we set out, but it actually kind of makes sense in the end,right? If we're going to build something that's going to be transformative to the world, like, yeah, probably is going to require some, you know, maybe the biggest physical machines that humanity has ever created, like, kind of type checks.
So I think that there's two answers. Like, the naive answer is, OK, yeah, you want two kinds of accelerators. You want one that's really compute-optimized, one that's very latency-optimized. Throw, like, tons of HBM on one of those and, you know, tons of compute on the other.
You're all good. Now, one thing that's really difficult is predicting the ratios,right? Now you have a new problem you have to think about and if you get the balance wrong, suddenly you're going to have a whole part of your fleet that's just useless.
Yep.
And that sounds really scary. Now, the thing is, because the way that these things work is there's no requirements in this field. There's no constraints in this field. There's just sort of this linear program that people are optimizing.
And so, yeah, if you give our engineers some sort of misbalance of resources, like, we will find ways to utilize it, maybe at great pain,right? But an example of this is, you know, you've seen the whole field move towards a mixture of experts.
And to some extent, what a mixture of experts is is saying, well, we have all this DRAM sitting around that isn't being used for anything because the balance is wrong. Fine. Well, let's fill it up with parameters and we'll actually not cost any compute and we'll just get extra ML compute efficiency out of it.
Like, boom, there you go. And so I think that there is some of that where if you get the balance wrong, it's actually not the end of the world. Homogeneity of accelerators is, like, a very nice default to start.
But I think that ending up with purpose-built accelerators is also not super crazy. And the more that we move to these worlds where it's the just dollars of CapEx for this infrastructure starts to become so eye-watering, then starting to hyper-optimize for some of these workloads is pretty reasonable.
But I think the jury is a little bit out because if you think about it, that the research is just moving so fast and to some extent that dominates everything else.
OK. I wasn't planning to ask this, but you just brought up the research stuff. Can you rank current scaling bottlenecks for GPT-6?
Research Priorities35:21
Ah.
Compute, data, algorithms, power, money.
Yes.
Which one's, like, the, you know, number one and two? Which one are you, like, most rate-limited on?
I mean, look, I think we are in a world where basic research is back. I think that is really amazing,right? There was this period, yeah, basic research. There was a period where it felt like, allright, we got a transformer, let's just scale it, you know?
And I find those problems very exciting. I have a lot of fun. Just, like, you've got a very well-defined hard problem. You want to just move the number up and to theright. But it also is a little intellectually dissatisfying in some ways.
It's like that it feels like there's more to life than just, you know, attention is all you need, paper, you know, in vanilla form. And so I think that what we've started to see is that we're operating at a scale now where we've pushed the compute, we've pushed the data so far that you can start to get, you start to have algorithms is, like, again, just back as an important and really almost a long pole in terms of future progress.
And so all of these things, they're all important poles of the tent and, you know, on any one day, it might look a little opposite one way or another. But yeah, fundamentally, I think it's like you want to keep these all in balance.
And it's really exciting to see things like the RL paradigm. That's something that we invested in very deliberately for multiple years. It was like when we trained GPT-4, the very first thing, like, the thing that was really interesting was when we talked to GPT-4 for the first time, we were like, is this an AGI?
Like, it's clearly not an AGI, but it's really hard to say why,right? It's like there's something about it. It's so fluid and smooth, but somehow it falls off the rails. And it's like, well, we got to solve that reliability problem.
And you're like, well, it has never actually experienced the world,right? It's like someone who's just read all the books or, you know, sort of read, you know, sort of observed the world, has observed the world and never experienced it itself,right?
It's like, you know, sort of just, you know, watching it through a pane of glass or something. And that to me is, you know, was something where you're just like, OK, clearly we need a different paradigm and we just pushed on it until we made it really work.
And I think that that remains true today that there's other very clear missing capabilities that we just need to keep pushing and we will get there.
Awesome. Broadening out just from just broad OpenAI things. Well, honestly, I'm just going to let, so we asked Jensen for one question. He's an overachiever, so he sent in two. So let's play a second video.
Future Workflow38:00
The AI Native engineers in the audience, they are probably thinking in the coming years, your OpenAI will have AGIs and they will be building domain-specific agents on top of the AGIs from OpenAI. And so some of the questions that I would have on my mind would be how you think their development workflow would change as OpenAI's AGIs become much more capable and yet they would still have plumbing workflows, pipelines that they would create, flywheels that they would create for their domain-specific agents.
These agents would, of course, be able to reason plan, use tools, have memory, short-term, long-term memory. And there'll be amazing, amazing agents. But how does it change the development process in the coming years?
Yeah. I think that this is a really fascinating question,right? I think you can find a wide spectrum of very strongly held opinion that is all mutually contradictory. I think my perspective is that, first of all, it's all on the table,right?
Maybe we reach a world where it's just like the AIs are so capable that, you know, we all, you know, just let them write all the code. Maybe there's a world where that you have, like, one AI in the sky.
Maybe it's that you actually have a bunch of domain-specific agents that require a bunch of specific work in order to make it happen. I think the evidence has really been shifting towards this, like, menagerie of different models. And I think that's actually really exciting,right?
That there's different inference costs, just even from a systems perspective, that there's different trade-offs, like distillation works so well. So there's actually a lot of power to be had by models that are actually able to use other models.
And so I think that that is going to open up just a ton of opportunity because, you know, we're heading to a world where the economy is fundamentally powered by AI. We're not there yet, but you can see itright on the horizon.
They're working on it all.
Exactly. I mean, that's what people in this room are building. That is what you are doing. And the economy is a very big thing. There's a lot of diversity in it. And it's also not static,right? That I think when people think about what AI can do for us, it's very easy to only look at, well, what are we doing now and how does AI slot in and, you know, the percentage of human versus AI.
But that's not the point,right? The point is how do we get 10X more activity, 10X more economic output, 10X more benefit to everyone? And I think that the direction we're heading is one where the models will get much more capable.
There'll be much better fundamental technology and there's just going to be, like, way more things we want to do with it and the barrier to entry will be lower than ever. And so things like healthcare that you can't just, you know, it requires a responsibility to go in and think about how to do itright.
Things like education where there's multiple stakeholders, you know, the parent, the teacher, the student. Each of these requires domain expertise, requires careful thought, requires a lot of work. And so I think that there is going to be just, like, so much opportunity for people to build.
And so I'm just so excited to see everyone in this room because that's theright kind of energy.
Thank you for encouraging us and being an inspiration. Thank you so much.
Thank you.
Greg Brockman, everybody.
Thank you.





