1-GHz Homage0:00
What we really wanted to pay homage to today is—actually, you know, just 25 years ago we crossed the 1-GHz speed barrier in microprocessors. What's really crazy is, when we started thinking about this talk, I actually thought it happened a lot before 1999, and I just kind of remember my own arc of getting involved with computers.
But really, it was 1999; I had to kind of double and triple-check it. This is the exact press release when Intel broke the 1-GHz speed barrier. And obviously, that was interesting, you know, for a couple of perspectives. One, it was this, you know, really big number and moment.
But two, it was really after this that, you know, Intel started to change about how they think about processors will be used, and they went for, I guess, you know, multi-cores and things like that. And it's really something that we need to think about in terms of what's going to happen with LLMs.
And really, if you go back to the rate of increase, it only took, you know, about two decades to get three orders of magnitude speed improvement in microprocessors. And so if we take a step now and look at where we are with LLMs, and we think about anywhere close to the speed of innovation, and in fact, you know, what we hear a lot of people talk about, you know, including Jensen, is that we're beyond the sort of curve of Moore's Law, so we're actually innovating even faster than that in LLMs today.
You know, just to look at what we've been able to do at Groq, just in a short amount of time, you know, this is between April and June of this year, you know, we were able to increase the speed of LLaMA 3 8B by over 50%.
And so the improvements that are happening in this area are really, really quick and super exciting, and we're really kind of keen to kind of dive into what could happen here. And so let's think about, like, the state of the art,right?
Speed Leap2:08
And so, you know, there's models today that, you know, we can process and others can process at, say, huge inputs. So you have the equivalent of, you know, 10,000 input tokens per second, which gets you down to, say, a third of a second across, you know, processing all of those.
And when you do that, you actually end up with these capabilities from a, you know, speed perspective that far exceed human capabilities for both integrating and analyzing information, and it's happening, you know, really, really fast. The example I like to talk about here, and I don't know if you've used this, but I highly recommend it, it's this, you know, really cool service called Globe.Engineer.
And what it does is you give it a task, or, you know, so I say something here, and I think the example I use here: "Help me plan a trip to New York to try, you know, the best pizza" or something like that.
And what it will do is, and, you know, I couldn't even capture the whole screen here, but it'll basically figure out all the different elements that have to happen, and it's doing this live online, it's connected to the internet.
So everything from the flights to the taxi options to the hotel options and then the food options and then itinerary and how I can do it. And it, you know, it does it all in, you know, maybe less than five seconds.
And if you think about what's really happening there, and I like to, you know, think about when I plan for trips myself, I end up basically opening, you know, tens to sometimes even hundreds of tabs, and those tabs each have like a research stream happening for me.
And now all of that is solved in, like, you know, a simple interface, you know, really enabled by these LLMs being able to, one, input process tokens, input tokens faster, and then ultimately output tokens faster. And it's really giving us a huge edge up in how we operate as humans.
And, you know, where does this all go? Like, if we start thinking about, you know, human superintelligence and optimizing and accelerating models, it really takes us to, like, interesting paradigms here. And, you know, we'll talk about this more in a second, but, like, you know, the high-level way to think about it is: what if an LLM, you know, really becomes either like an operating system or like the core of, you know, how we think about compute today, and we think about it completely differently than any of the approaches that we've had before?
Industrial Shift3:55
You know, the way we program these things, the way our expectations are about how they analyze things. And so we're really, you know, that's interesting in terms of where this is going in terms of superintelligence and staying away from AGI, but more about changing the paradigm from where we are today.
And, you know, the thing that crosses my mind here is what happened in the Industrial Revolution. You know, if we think about three industries, let's think about making food, making cars, and making clothes. All of those before the Industrial Revolution were bespoke,right?
So you'd have, you know, people that would make one or two cars a day. You'd have people work on farms that could, you know, maybe farm for less than a city, even a small village. Or someone that was making sweaters could, you know, make them, you know, one a day or maybe even one a week.
And when we had the Industrial Revolution show up, we basically had this ability to make hundreds or thousands of cars a day, food, farming at a scale that could be national, clothing that could be made at national scale.
And we're really, you know, we haven't had that in technology. The arc of technology has been, and this isn't my own framework, it comes from Paul Maritz, you know, who was a long-time Microsoft guy and then VMware and then Pivotal, where he and I met.
You know, he said the first era of computing was just taking paper processes and making them digital. And he goes, that's evident in the way if you think about how the operating system is structured: files, folders, inbox, outbox.
Those are all paper processes that got turned into, you know, digital processes. The next era for us was basically making those things connected,right? That's the internet era. And what we've been through now, you know, maybe in the last 15 years, is form factor changes,right?
Either pushing things into the cloud for scale or mobile so you can do it on your phone. But finally, with AI, we're starting to get to a place where we have the industrialization in the same way we saw for those, you know, manufacturing and physical industries, we see that for technology.
So, you know, 18 or maybe 24 months ago, if you needed to have a Photoshop made of some kind of artifact that you were going to put in a presentation, you'd go to your designer, and maybe that designer would make one or two a day for you.
Now you can go to Midjourney and get 1,000 made in the next minute if you want to. So we're going through that same kind of industrialization for technology. And if we just dive in deeper here into, you know, where we go as we can get into, like, 10,000 complex decisions per second just by getting this down to, you know, 0.1 milliseconds.
And then if we really, really kind of start increasing that, it does become viable to think about the core of our computing becoming an LLM. And I think this is a real challenge for a lot of people because we, you know, obviously we have existing paradigms that we're really, really locked into.
But this paradigm shift is fundamentally different in terms of how software will be built, how software will run, and how software will scale. And we don't think about it too much today because we think about the speed associated with, you know, running LLMs and their capabilities.
Core & Possibilities7:14
But if we can imagine the same growth that we saw in CPUs happen in this era, we can imagine that the core of these devices change to become, you know, something. And this is, again, Hattipta Karpathy. This is a diagram that he drew.
But we can imagine an LLM being a core at, you know, whether what happens in video and audio. We're starting to see that today. What happens in our browsers, how we interact with other LLMs, how we interact with, you know, code interpreters, and even our file systems and how we interact with those type of things.
And so what is the art of possible if we start doing this? And so I'll just kind of rattle off some things here that, you know, crossed our minds as we were putting this presentation together. You know, we really don't spend a lot of time thinking about it, but many responses today in LLMs are sort of near real-time.
They're at sort of reading speed. But if we go to, like, instantaneous responses and decision-making, this becomes a lot faster. Again, this is really evident when you think about something like that Globe example I showed. What you're really able to do there is take a task that would probably take you either an afternoon or evening or a number of evenings, and it's done in just a few seconds for you.
Speed & Personalization8:22
And then there's personalized experiences. You know, today we don't really have a lot of personalized experiences happening. We're starting to see elements of it. You know, I think OpenAI has started to launch a number of features that allow it to understand, you know, specifics of your world.
It could be your pet's names or kids' names or spouse's names. But really, I think, you know, where this goes to, and a lot of people push on this. I know, you know, two of my friends, you know, Bill Gurley and Brad Gerstner, they talk about this a lot on their pod where they really view personalization as the next major frontier.
And personalization and speed are going to go hand in hand if we're going to make that work kind of seamlessly for folks. I think next is kind of a universal natural language processing. And so if we think about our interface today to software, it's, you know, you know, we started with sort of point and click and keyboards.
Multimodal Interfaces9:14
We've gone to touch with our, you know, mobile devices. But really, you know, you start to see the power of this. And, you know, I think everyone's been super excited for the release of GPT-4.0, the voice agents. We don't think we've fully got there yet.
But I think we've showed the art of the possible there with what they were able to do with voice. And then that kind of mixed interaction, I would say, like, you know, we refer to it as sort of like XRX, where it's like any type of input, reasoning, and any type of output.
You know, the example I like to tell people there, if you're trying to order something, you may want to interact with an agent in voice, but you may want to see the responses in text. And so think about if you're trying to book your haircut and you want to say, "Well, tell me what times are available."
And then, you know, it tells you, "Well, there's 9:00 a.m. and 11:00 a.m. and 3:30 and 5:30." That's hard to remember if it's just coming back to you in voice. So you want to basically have these interactions that are multimodal.
That kind of touches on my second point there. And I think we're going to start to see a lot more of those interface changes as well. You know, advanced virtual assistance. This is like complex task scheduling. I think a lot of what we'll see in the back half of just this year is, you know, agents start to become much more complex and a lot of focus from LLM providers as well, I think, on making, you know, complex tasks something that are solved.
Agents & Collaboration10:26
It's interesting today because we measure the efficacy of an LLM through generally single shot. And I think we do that because, you know, going back to that where we, you know, the start of the conversation, which is the performance barrier.
But naturally, if you even take any existing LLM today and multi-shot it, its scores get a lot better. And there was a couple papers that came out recently that showed if you just had multiple agents working together on a problem, they can far of a less, you know, less parameter model, they can compete with higher parameter models just by doing sort of multi-shot reasoning or working together.
And so I think we'll see a lot more of that as the speed improves. And I think there's an incredible, incredible optionality there. You know, we saw the first, I think, first cut of collaborative AI agents with Apple AI, you know, where you see something maybe running on device, interacting with something off device.
It's, I think it's a very early implementation, and I think these things will get much more sophisticated and better. An area, you know, we've spent a lot of time within our career is like analytics and predictive analytics. I think today everything is, you know, pretty much action-oriented and drived off a human action.
So I think if we get to a place where the speed goes up, it can be a lot more predictive. You know, what does that really mean? It's just an agent that's always running in the background because the compute cycles are next to free.
Predictive & Context12:03
We don't see that today, but I think we get there as we get, you know, higher up the curve. You know, context-aware as well. And today, again, we are generally limited to how much context we can provide, and we're having to, even with models with bigger context windows, we still have to, you know, be conscious of, you know, how much compute cycles we're going to use.
But I think if that becomes next to free, it becomes quite powerful for us. You know, creative tools and customizable content. I'll focus on the second one here. This is an area where I think many of us would like to see things go.
Creative Content12:48
You know, the example I always like to, you know, one of my favorite shows was Seinfeld. And obviously, you know, it's not on anymore. But one of the things I like to do, you know, when I'm bored is go into, you know, LLM of choice and have it write a Seinfeld episode but made up of, like, modern-day things that are happening.
And if you ever try that, it's super fun because it does an incredible job of, you know, identifying which character in those scenarios that you give it would have, you know, sort of the funny or odd thing happen to them.
And so the idea of, you know, taking that beyond sort of writing and taking that to multimedia forms is going to be really, you know, really, really powerful going forward. You know, complex decision-making. You know, before our company was acquired by Groq, you know, we were building a company called Definitive Intelligence.
Complex Decisions13:18
So we spent a lot of air, a lot of time in this space, not only doing sort of, say, natural language to, you know, analysis of SQL,right? Text to SQL, as a lot of people would call it. But, you know, Rick, who's sitting here with us, like, you know, he was working on this really cool product for us called Pioneer, which was an automated data science agent where it's really meant to run almost endlessly on a problem.
And, you know, you sort of define a KPI. If you think about how a business runs, a business has a bunch of KPIs, and then a business has a bunch of data that's coming in, and then usually humans are taking that data and analyzing it to KPIs and creating PowerPoints and spreadsheets and telling either senior management or the world how well they're doing.
Well, there's no reason that just shouldn't happen automatically,right? And where there's an agent just constantly, you know, looking at the new data that's coming in, asking additional questions, diving into it. And I think we had a lot of interesting things emerge.
You know, we had let Pioneer loose on a data set of human workers and their performance reviews. And one of the things that we saw was it was able to correlate really interesting things that we couldn't think about in terms of, you know, depending on your age and depending on your performance review, it really affected your, I guess, your output, your productivity.
And so it was able to kind of discover that if you're of a certain age and you got a certain type of performance review, your productivity would fall off. And maybe Rick can correct me if I'm wrong later, but it was something along those lines, which was always an interesting example for us.
And then obviously a lot of, you know, really interesting things around dynamic optimization. You know, this is an area we're familiar from before. You know, when a bunch of us were at Ford after the acquisition of Autonomic, we really saw, you know, for the supply chain, if you think about how, you know, cars are produced and how they're shipped, you know, there's, you know, pretty sophisticated software that does this, but it's still not efficient,right?
And I think, you know, the art of the possible with sort of what we were talking about earlier could be very, very interesting for some of our old colleagues at Ford. I'll touch on a couple more things and then leave a couple minutes for questions if there's any.
But edge AI and decentralized AI, this is pretty cool. You know, there's a really cool project called, you know, hyperspace.ai. What they're doing is they actually have a lot of, you know, taking, you know, sort of like SETI at home or even Render, and where they're basically allowing people to take their unused GPU compute and make it available in the cloud.
Decentralized & Security15:54
Or I guess, yeah. And why that's interesting is there's certain use cases that necessarily don't require something to be real-time. And so I think we'll see a lot more of that. Now, this intersects really well with us getting more throughput and getting lower latency out of existing systems.
So I think we'll see a lot more of that as well, especially because the amount of power consumption that's required, if you distribute it, that you could be really interesting. And then a couple more here is enhanced security and privacy.
This is a big area. You know, I was talking to one of our colleagues last night, and he was subject to a really, really scary type of, I guess, maybe phishing call where, you know, someone had called in, sounded very formal, and had access to a lot of his information.
Now, you know, we've all seen there's, you know, these kind of people that run scam call centers and people that go and attack them. But these folks armed with AI are much more sophisticated because they can create stories and narratives that are much deeper than sort of the call center worker of past.
And now I think in order to protect against these systems, you'll almost need to have something on your side so that you can, you know, you can think about it. You know, with our colleague, he was just so confused because the narrative was so good.
The only way he could really figure out that this person was a scammer other than hanging up on them and saying, "Hey, well, send me some kind of formal message through the HSBC app, and then I'll know it's you."
And, you know, the person wasn't able to do that. And so I do think, you know, as voice cloning, as more of our information is online, we have to be really careful, and we'll need these protective systems that we can use, and we need them to run incredibly fast.
Education & Interop18:03
And so, and I think this is the last set of them here is, you know, education is something that's really important to us, you know, broadly at Groq. We think about this, and we think about, you know, making tokens available cheaper and more broadly and being able to personalize.
You know, Salla Khan has a very good TED talk from a couple years ago where he really highlights, you know, it's the two-sigma talk. And he said, "You can take any student at any level, the highest levels, or even someone performing lower.
And if you give them a personalized tutor, they can improve their test scores to standard deviations." And so imagine doing that, you know, obviously with AIs that are, you know, can be one very cheap to use, and that can be personalized to their learning experience.
You know, I was speaking to someone recently who was building an AI service for homeschooling. And what was powerful about that particular service is, let's say you have a young child and they're really into unicorns or ponies. And you want to teach them about, you know, math and, you know, math, subtraction, addition, multiplication.
It's a lot easier if you frame it in the context of those things. Hey, you know, you have three ponies times two unicorns, and what do you get from it? And so I never thought about that before, but for learning and customizing that for the interest of the person is quite powerful.
So we'll see more of that. And then just interoperability and compatibility,right? I think this is an area, if you've ever been in enterprise software, the majority of money spent in deploying and maintaining enterprise software is really related to, you know, inner connectivity and interoperability and compatibility.
And so, you know, having really fast and cheap, you know, AI technologies will help us really reduce a huge burden that exists on the enterprise today. So that's it. Hopefully you guys enjoyed that.





