Introductions0:00
Really excited to be moderating this panel between two of my favorite people working in AI. I'm Brittany. I'm a principal at CRV, which is an early-stage venture capital firm investing primarily in seed and Series A startups. Chris, why don't you give us a little bit about yourself?
My name's Chris. I'm currently the CTO at Prefect. We're a workflow orchestration company. We build a workflow orchestration dev tool and sell an orchestration, remote orchestration as a service. A little background: so I started kind of my journey into startup land and eventually AI and data.
I got a Ph.D. in math focused on non-convex optimization, which I'm sure a lot of people here are into. And then eventually, you know, data science, and then into the kind of dev tool space, which is where I'm at now.
Awesome. And Bryan, fill us in on your side.
I'm Bryan. I lead AI at Hex. Hex is a data science notebook platform, sort of like the best place to do data science workflows. I was going to say I started my journey by getting a math Ph.D., but he kind of already took that one.
That's kind of awkward. Yeah, I've been doing data science and machine learning for about a decade, and yeah, currently find myself doing AI, as they call it these days.
Awesome. So both of you are at relatively early-stage startups. And as we all know, early-stage startups have a number of competing priorities, everything from hiring to fundraising to building products. And one might say it would be a lot to kind of take a moment and just say, what is this AI thing?
AI Decision1:29
What the fuck do we do with this? And so I'm wondering, how did you decide that AI was something that you really needed to invest in when you already had, you know, an established business, growing well, lots of users, lots of customers, presumably placing a lot of demands on your time?
So Chris, I would love to hear from you on how you guys thought about that choice.
Yeah, so there are a couple of different dimensions to it for us. So we are, you know, a workflow orchestration company, and our main user persona are data engineers and data scientists. But there's nothing inherent about our tool that requires you to have that type of use case.
And so one thing, one dimension for us is,right, we assume that a big component of AI use cases are going to be data-driven,right? Like semantic search or like retrieval, summarization, these sorts of things. So just we wanted to make sure that, you know, we had a seat at the table to understand how people were productionizing these things.
And like, were there any new ETL considerations when you're, you know, moving data between maybe vector databases or something? So that was one thing. Another one that I think is interesting is when I look at AI going into production, I see basically a remote API that is expensive, brittle, and non-deterministic.
And that's just a data API to me. And so,right, if we can orchestrate these workflows that are building applications for data engineers, presumably a lot of that's going to translate over. And so, and I mean, last, like, you know, I'm sure the reason most people are here now is, you know, it was fun.
And so we just wanted to learn in the open. So we did end up just kind of creating a new repo called Marvin that I think Jason mentioned in his last talk, just to kind of keep up, you know, be incentivized to keep up.
And Bryan, you were literally brought on board to Hex to focus on this stuff. Would love to hear more about how that decision was made and how you've spent your time on it.
Yeah, I think a couple of things. One is that data science is this unique interface between sort of like, like business acumen, creativity, and like pretty, like difficult sometimes programming. And it turns out that, like, the opportunity to unlock more creativity and more business acumen as part of that workflow is a really unique opportunity.
I think a lot of data people, the favorite part of their job is not remembering matplotlib syntax. And so the opportunity to sort of like take away that tedium is a really exciting place to be. Also, realistically, any data platform that isn't integrating AI is almost certainly going to be dooming themselves to the now.
And sort of it'll be table stakes pretty soon. And so I think missing that opportunity would be pretty criminal.
Yeah, I totally agree with that. So you decided that you were going to go ahead and do this. You were going to go all in on AI. What criteria did you evaluate when you were determining how you were going to build out these features or products?
Building Criteria4:26
Did you optimize for how quickly you could get to market, how hard it would be to build, ability to work within your existing resources? What criteria did you consider when you were saying, okay, this is how we're actually going to take hold of this thing?
So for us, I guess there's two different angles. There's the kind of just pure open source Marvin project, not, you know, it is a product, but not one that we sell, just one that we maintain. And then we do have some AI features built into our actual core product.
And I think they have slightly different success criteria. So for Marvin, it's mainly just getting to see how people are experimenting with LLMs and just talking to users directly,right? It just kind of gives us that avenue and that audience.
And so that's just been really useful and insightful for us. So we just get on the phone. I mean, our head of AI gets, you know, talks to users at least, you know, a couple of times a day.
And then for our core product, so one way that I love to think about dev tools and think about what we build is failure modes. So like, I like to think of choosing tools for what happens when they fail.
Can I quickly recover from that failure and understand it? And so a lot of our features are geared towards that sort of kind of discoverability. And so for AI, it's kind of the same thing. It's like quick error summaries shown on the dashboard for quick triage.
And then measuring success there is like relatively straightforward,right? It's like how quickly are users kind of getting to the pages they want and how quickly are they debugging their workflows? So like very quantifiable.
Yeah, would love to hear from you too.
Yeah, my team's charter is relatively simple. It's make Hex, make using Hex feel magical. And so ultimately, we're constantly thinking about sort of what the user is trying to do in Hex during their data science workflow and making that as low friction as absolutely possible and giving them more enhancement opportunities.
So a simple example is, I don't know how many times y'all have had a very long SQL query that's made up of a bunch of CTEs, and it's a giant pain in the ass to work with. So we build an explode feature.
It takes a single SQL query, breaks it into a bunch of different cells, and they're chained together in our platform. This is like such a trivial thing to build, but it's something that I've wanted for eight years. Like I've done this so many times, so annoying.
And so thinking like that makes it really easy to make trade-offs in terms of what is important and what we should focus on. And so in terms of like how we think about, yeah, like where our positioning is, it's really just how do we make things feel magical and feel smooth and comfortable?
Hiring7:11
And how did you reallocate resources beyond, you know, they obviously hired you. That was a great step in theright direction. But what else did you do to actually get up and running in terms of operationalizing some of this stuff?
Yeah, I think we kept things pretty slim, and we continue to keep things pretty slim. We started with one hacker. He built out a very simple prototype that seemed like it should promise. And then we started building out the team.
We scaled the team to a couple of people, and we've always remained as slim as possible while building out new features. These days, I have a roadmap long enough for 20 engineers, and we continue to stay around five.
And that's not an accident. Basically, like ruthless prioritization is definitely an advantage.
And Chris, you guys wound up hiring a guy as well,right?
Yeah, we hired a great guy. His name's Adam. So he definitely owns most of our AI, but also,right, like anyone at the company that wants to participate. And so there was one engineer that got really into it and is, for all intents and purposes, has like effectively switched to Adam's team and is now doing AI full-time.
Yeah, so you guys are really dedicating a lot to solving this problem, including the hiring of two people on your side and one on Bryan's. So presumably, you're going to be looking for a return on that investment. So how do you think about what a successful implementation of an AI-based feature or product looks like?
Success Metrics8:14
For us, I would say that already we've hit that success criterion. So now the question is like further investment or just kind of keep going with the way that we're doing it. But so big thing was time to value in the core product that we can just easily see has definitely happened with just the few sprinkles of AI that we put in.
So we'll just kind of keep pushing on that. And then kind of like I said in the beginning, just getting involved in those conversations, those really early conversations about companies looking to put AI in production. And we've been having those on the regular now.
So I would say like already feels like it was well worth the investment.
What about you guys? You obviously just had a big launch the other day too. Curious how you thought about success for that.
Yeah, once again, it's sort of like how frequently do our users reach for this tool? Ultimately, magic is a tool that we've given to our users to try to make them more efficient and have a better experience using Hex.
And so if they're constantly interacting with magic, if they're using it in every cell and every project, then that's a good sign that we're succeeding. And so to make that possible, we really have to make sure that magic has something to help with all parts of our platform.
We have a pretty complicated platform that can do a lot. And so finding applications of AI in every single aspect of that platform has been one of our sort of like, you know, north stars and very intentionally so to make sure that we're, you know, making our platform feel smooth at all times.
Awesome. Well, let's move on to the next section, where you're going to talk about what, how you guys actually built some of these features and products. Since we're all here at the AI Engineer Summit, I assume we all have an interest in actually getting stuff done and putting it into prod.
Build vs Buy10:02
So when you were making some of these initial determinations, Bryan, how did you guys determine what to build versus buy?
Yeah, so from day one, I think the question, one of the first questions I asked when I joined is what they were doing for evaluation. And you might say like, okay, yeah, we've heard a lot about evaluation today, but I would like to remind everyone here that that was February.
And the reason that I was asking that question already in February is because I've been working in machine learning for a long time, where evaluation sort of like gives you the opportunity to do a good job. And if you've done a poor job of objective framing and done a poor job of evaluation, you don't have much hope.
And so I think the first thing that we really looked into is evals. And back then, there was not 25 companies starting evals. There are now more than 25, but ultimately we made the call to build. And I'm very confident that that was theright call for a few reasons.
One, evals should be as closest to production as possible, is literally like using prod when possible. And so to do that, you have to have very deep hooks into your platform. When you're moving at the speed that we try to move, that's hard for a SaaS company to do.
On the flip side, we chose to not build our own vector database. I've been, you know, doing semantic search with vectors for six, seven years now, and I've used open source tools like Face and Pinecone back when it was more primitive.
Unfortunately, a lot of those tools are very complicated. And so having set up vector databases before, I didn't want to go down that journey. So we ended up working with LansDB and sort of built a very custom implementation of vector retrieval that really fits our use case.
That was highly nuanced and highly complicated, but it's what we needed to make our RAG pipeline really effective. So we spent a lot of effort on that. So ultimately, just sort of where is the complexity worth the squeeze?
Totally. And Chris, what about you guys? How did you do that?
So I have a couple of different kind of things that we decided on here, and some of which are still in the works. Vector databases, million percent agree with that. Like we would never build our own. We haven't had as much need of one, I think, as Hex, but we've done a lot with both Chroma and with Lance, but neither in production yet.
So none of those use cases are in prod. And so the way that I've, the exposure that I've seen about people actually integrating AI into, you know, their workflows and things is there's a lot of experimentation that happens.
And then you kind of want to get out of that experimental framework, maybe by just looking at all of the prompts that you were using and then just using those directly yourself with no framework in the middle. And then once you're kind of in that mode, like I was saying before, like you're just at the end of the day, you're interacting with an API and there's lots of tooling for that.
And so I kind of see a lot of the decisions, at least that we had to confront on build versus buy is like it's another just kind of tool in our stack. Do we already have the sufficient dev tooling to support it, make sure it's observable, monitorable, and all this?
And we did. So we didn't do any buying. It was all build.
Yeah, and as you mentioned, I think, you know, you talked about many, many eval startups. I think we're all familiar with the broad landscape of vector databases as well. Are there any pieces of your infrastructure stack that you wish people were building or you wish people were kind of tackling in a different way than what you've seen out there so far?
Infra Gaps13:31
Either one of you?
Yeah, I mean, I think it would have been a hard sell to sell me on an eval platform. I think there was some opportunity to sell me on like an observability platform for LLMs. I've looked at quite a few, and I will admit to being an alumni of Weights and Biases, so I have some bias.
But that being said, I think there is still a golden opportunity for a really fantastic like experimentation plus observability platform. One thing that I'm watching quite carefully is Rivet by Ironclad. It's an open source library. And I think the way that they have approached the experimentation and iteration is really fantastic.
And I'm really excited about that. If I see something like that get laced really well into observability, that's something that I'd be excited about.
And Adam, your side?
I think
like small addition to what Bryan said, which is just more focus on kind of the machine-to-machine layer of the tooling. And so I think a lot, you know,right at the end of the day, the input is always kind of this natural language string, and that makes a lot of sense.
But the output, making it more of a guaranteed typed output, like with function calling and other things, I think is one step in the journey of making, of integrating AI actually into backend processes and machine-to-machine processes. And so any focus in that area is where, you know, my interest gets peaked for sure.
Challenges15:28
Yeah, totally. Okay, so you have all your people and you have all your tools, and then you're obviously completely good to go and fully in production. JK, we all know it doesn't work that way. What challenges, what challenges did you run into along the way?
Maybe ones that you didn't expect or that were larger obstacles than you would have thought.
So I don't, integrating AI into our core product, I would say from a tooling and developer perspective and, you know, productionizing perspective, none. Culturally though, I would say we definitely hit, you know, some challenges, which is that when we first were like, allright, let's start to incorporate some AI and do some ideation here,right?
A lot of engineers just started to throw everything at it. Like we should, it should do everything. It can monitor itself for like all of this stuff. And it was like, allright, allright, everyone needs to kind of like backtrack.
And so just that internal conversation of like, you know, getting buy-in on like very specific focus areas, which, you know, at the end of the day, where we are focused is that just removal of user friction, whether it's through design or just through like quicker surfacing of information that AI just like lets you do in a more guaranteed way.
But yeah, restraining the enthusiasm was the biggest challenge for sure. And it still exists to this day.
Everyone wants to be an AI engineer,right?
Yeah, exactly.
Yeah. What about you guys? Did you have similar or different issues?
That's interesting. It's like a similar flavor. It's a little different in sensation, which is to say that like, you know, I've never met an engineer that's good at estimating how long things take. And I would say that like that is somehow exacerbated with AI features, because then your first few experiments show such great promise so quickly, but then the long tail feels even longer than most engineering corner case triage.
Just such a long journey between we got this to work for a few cases and we think we can make it work to it's bulletproof is even more of a challenging journey. And I, yeah, this like over-enthusiasm, I think, yeah, slightly different instantiation, but similar flavor.
Whenever you're on that journey, how are you testing and tracking along the way, if at all, which is totally, yeah.
Yeah, I mean, to be a broken record, a lot of like robust evals, like trying really hard to codify things into evaluations, trying really hard to codify like if someone comes to us and says, wouldn't it be great if magic could do X?
We sort of pursue that conversation a little bit further and say like, okay, what would you expect magic to do with this prompt? What would you expect magic to do in this case? And kind of get them to kind of like vet that out and then sort of using this like barometer of could a new data scientist at your company with very little context do that?
And sort of that like, you know, cutting edge around what's feasible and what's possible.
Product UX18:31
Yeah, that makes a lot of sense. And, you know, one of the reasons I was excited to have the two of you up here together is because, you know, while Prefect has some elements of AI in the core product, as you mentioned, probably you're best known for the Marvin project that you guys have put out there, which is kind of a standalone project, which is a really interesting phenomenon that I'll say that I've kind of observed in this current wave of AI, which is, you know, companies that maybe weren't doing AI previously launching entirely separate brands, essentially alongside the core product.
So would love to understand more of what were your user experience considerations when you were building out, you know, Marvin as a separate product versus Prefect? What freedom did that allow you? What restrictions did you still have?
Yeah, that's a good question. So a few different angles there. I think one kind of philosophical angle is, you know, we try to do things that maximize our ability to learn without having to go full commitment. And so I think starting a new open source repo, like,right, we definitely have some ties to it now.
We have to maintain it, but past that, it's not all that high of a cost. But like if it, you know, it's all upside basically. If no one notices it, no big deal. We learned a little bit more about how to, you know, write APIs that, you know, interface with all of the different LLMs, for example, or something like that.
Or if it does take off, which, you know, it basically did for us, we got to meet all these new people who are working on interesting things like AI and data adjacent. Before the core product, this was maybe more, I guess, kind of interesting.
And Bryan, I'd be curious to hear about how much you had to like really focus some of your prompts to the use case that you cared about. So Prefect is a general purpose orchestrator. And so the reason I say that again is our use case scope is like technically infinite.
And so helping people write code to do completely arbitrary things is definitely not a value add we're going to have over the engineers at OpenAI or at GitHub or something else. So we knew that we couldn't invest in like that way of integrating AI.
And so then the next question was like, okay, so then what are just the marginal adds? And that's kind of where we landed, you know, where we are today. But there was, we did put energy initially like, can we put this directly in like the SDK or something like that?
And just very quickly realized that it was just too large of scope. And at that point, you might as well just have the user do it themselves. And like we're not adding anything to that workflow.
Yeah. Yeah. And on the flip side, you know, magic has been kind of a part of Hex basically, it seems like since inception from the outside. Obviously, we've all seen again a number of text to SQL players out there.
We can make arguments about whether or not those should exist as standalone companies. But I'm curious, you know, how you guys had to think about UX considerations when you were building out magic in the context of the existing Hex product.
Ultimately, I've been really fortunate to kind of like work with a great design team who sort of, they're just excellent. But the question about like how does magic feel? Magic is not its own product. I think that's one thing that's been important from early on.
Magic is not a product. Magic is an augmentation of our product. So it is a collection of features that makes the product easier and more comfortable to use. That is an easy sort of thing to keep in mind when deciding how to design, because it allows us to say, okay, like we don't want this to distract from the core product experience.
I can tell a story. We had one sprint where we designed something called Crystal Ball. And Crystal Ball was a really sick product. It did exactly what we wanted to do, and it felt wonderful. However, ultimately, it drew the user away from the core Hex experience.
And very quickly, our CEOrightly was like, I feel like this is kind of splitting magic out into its own little ecosystem. And that made it kind of clear that that might be the wrong direction to go. So even though Crystal Ball did feel really good and had a really incredible capability behind it, and frankly, the design on Crystal Ball was beautiful, the problem with that was it pulled us away from what we were really trying to do, which was make Hex better for all of our users.
Every Hex like consumer should be able to benefit from magic features. And that was starting to split that. And so we literally killed Crystal Ball, despite it being a really cool experience for that reason. So genuinely, we've really stuck to the like, it's one platform and magic augments it.
Yeah, that makes a lot of sense. And obviously, you know, Hex already had a relatively sizable user base at the time you guys launched this. So I'm curious, how did you think about the rollout, like just in terms of what users you gave it to and what timeline, what marketing did you do, all of those types of considerations?
Rollout23:19
Yeah, generally we start with a private beta and then we as quickly as possible expand that to a sort of like public beta. Our goal is to find people that are like engaged with the product and they are prepared for some of the limitations of AI tools.
Stochasticities come up many times, and ultimately we're expecting the user to work with a stochastic thing. Also, they're working with something very complex, which is data science workflows. So we're looking for people that are pretty technical in the early days.
Then we want to keep scaling and scaling to include the rest of the distribution in terms of technical capabilities so that we can make sure that it's really serving all of our users.
And on the flip side, again, you had a little bit maybe more flexibility with the rollout, just given that it was a new repo. I'm curious if that was different, similar to what Bryan's talked about.
Well, so yeah, well, the repo, no, it was, we hacked on it, you know, we had fun with it. We got it to a place where we felt proud of it, and then we clicked make public and then tweeted about it.
And that was like the end of that. So that was just pure fun. But for integrating AI into our core product, I mean, this isn't particularly deep, but it, you know, it's one of those things that I'm sure everyone here is thinking about and we'll continue to talk about, which is for us, a large part of our customer base are like large enterprises and financial services and also healthcare.
And so like very, very security conscious. And so we definitely had to make sure that this was like a very opt-in type of feature. But like, you know, we still want to have little like tooltips like, hey, if you click this, but also if you click this, we will send a couple of bits of data, you know, to a third-party provider.
So.
Yeah. And post rollout, just to go to kind of the last logical part of the conversation here, how have you guys thought about continuing to kind of measure the outputs? I mean, Bryan, you're the big evals guy up here, so I'm sure that'll be the answer, but I would love to hear more about how you think about that measurement and in terms of both the model itself, but also in terms of, you know, the model in the context of the product, which I think is also something that people, you know, need to think about.
Measuring Outputs25:17
Yeah. So I recently learned that there's a more friendly term than dog fooding, which is drinking your own champagne. And so I'll say, I drink a lot of champagne. I use magic every day, all through the day. One of the fun things about trying to analyze product performance is that you normally do that via data science.
And so I have this fun thing where I'm using magic to analyze magic, and I put a lot of effort into trying to understand where it's succeeding and where it's failing, both through traditional product analytics guided by using the product itself.
And so there's a very arroborous feeling, but ultimately good old-fashioned data science.
Love to hear it and appropriate with where you've come from.
Yeah.
What about you guys?
For us, it's, you know, I definitely don't have as much experience as Bryan on that side of it, but for a while, one thing we were doing when it was pure, just like prompt input, string output with no typing interface whatsoever, is then using that and then writing tests that again used an LLM to do comparisons and semantic comparisons and like write, there's obviously problems with that, but like it also kind of works.
But so then when we moved in kind of the typing world where like Marvin is for like type, guaranteed typed outputs, essentially, it definitely becomes a lot easier to test in that world, which is, you know, one reason that that's kind of the soapbox that I get on when I talk about LLM tooling, like bringing it into the backend is just like having these typed handshakes because, you know, you can write prompts where you know what the output should be and it should have a certain type.
And that's a very easy thing to test most of the time.
Yeah, totally. And one of the things I think has been, you know, most fascinating about this wave of software, and Bryan, you alluded to this a little bit earlier with your comments around, you know, being stochastic, essentially, is that it's not deterministic,right?
And also, I think that AI-based software doesn't have to be static either. It can be, you know, dynamic in a way that maybe traditional software isn't quite as much. And there's, you know, improvements that come along maybe on the UX side of things, but also the model.
We've heard a lot of people talk about techniques like fine-tuning, techniques like RLHF, RLAIF, all sorts of, you know, approaches to kind of continuing to improve the model itself in the context of the product over time. So I'm curious about how you think about measuring that improvement as you continue to hopefully, you know, collect data and refine your understanding of the end user.
Totally. There was a paper that came out in like June-ish or something that was like kind of splashy. It was from the, it was from Matthias from Spark. And it was like, oh, like the models are degrading over time, even when they say they're not.
And like what I thought was interesting was for like the people that are doing this stuff in prod, we already knew that. Like my evals failed the first day they switched to the new endpoint. I didn't even switch the endpoint over and suddenly my evals were failing.
So I think there is a certain amount of like when you're building these stuff, these things in a production environment, you're keeping a very close eye on the performance over time and you're building evals in this very robust way.
And I've said evals enough time for this conversation already, but I think the thing that I keep coming back to is what do you care about in terms of your performance? Boil your cases down to traditional methods of evaluation.
We don't need latent distance distributions and KL divergence between those distributions. We don't need that. Turns out like blue scores of similarity aren't very good for LLM outputs. This has been known for three, four years now. So take your task, understand what it means in a very clear, you know, human way, boil it down to binary yes or nos, and run your evals.
And to the people that say like, my task is too complicated, I can't tell if it'sright or wrong, I have to use something more latent. I would challenge you to try harder. The tasks that I'm evaluating are quite nuanced and quite complicated.
And it hasn't always been easy for me to come up with binary evaluations, but you keep hunting and you eventually find things. You talked about type checking and you talk about like type handshakes. And that's something that like a lot of people in ML have been preaching the gospel of composability for five years now.
You know, these are not new ideas. They're just maybe new to some of the people that are thinking about evals today.
Yeah. Well, so moral of the story is try harder, essentially. That's the takeaway.
Chris, did you have anything to add there?
I think the only thing I'd add is I don't have much take on actually how someone should do it or what they could consider, but I think, you know, you just described a highly non-deterministic, very dynamic experimentation workflow.
And like those are the sorts of things that just like our core product is meant for. And so like experimenting with those, like just knowing the structure of them is maybe the best way to say it is what fascinates me more than the actual like details of what metrics you might be using.
Yeah. Well, you know, I think the other reason I was really excited to do this panel is because we have kind of maybe two sides of the same coin as it relates to being an AI engineer here,right? One person coming from more of a traditional ML background, one person coming from more of a traditional engineering background, and both of you building these AI-based products.
Closing Takes30:47
So I wanted to give you a second if you have any last questions to ask of each other.
Yeah. So you work in this like data workflow space. And like I've thought a lot about composability and like data workflows. And I've long been a fan of sort of like workflow-centric ML. And so what I'd love to hear is sort of like when you think about building these agent pipelines, which are starting to get more into the like DAGs and the sort of like structured chains of response and request, what is the like one thing that like every AI engineer building agents should know from your sphere that'll make it easier for them to build agents?
So, oh, that's a really good question. I don't, I think the main thing is something that I kind of alluded to earlier, which is think about failure modes. I think that is the biggest thing. So like runaway processes, capturing potential oddities in outputs or inputs as early as possible with some observability layer.
And so the earlier you can get that wiring in, I think the better. And then caching is like the only time I will ever say this is definitely your friend in some of these situations. But it's also the root of all evil.
So you got to kind of, you know, balance that. But yeah, I think just thinking about the observability and debugability layer, especially with some of the kind of black boxy and like people who are pushing it and actually having like immediate eval of the returned code or something like having that monitoring layer, I think is just key.
Yeah. Chris, I know you've asked Bryan a bunch during this panel, but anything else you want to add?
Yeah, I mean, I'm just really curious, you know, I'm sure everybody asked you this, but the hallucination problem, like how, you know, obviously your users can just confront it directly. If it looks weird, they can see that it looks weird or it errors out, but just how do you think about it as the person building that interface for your users?
Yeah. Someone recently asked me for like references on hallucination and I was like, what are some good references on hallucination? And I Googled around and I found that generally the advice that people are giving is to fix hallucination, basically rag harder, just like make a better retrieval augmented pipeline.
And when I said that and I looked at myself, I was like, honestly, that's like kind of how we solved it. Like our reduction in hallucination for magic, which is not an easy problem, was that we had to think a little bit more carefully about retrieval augmented generation.
And in particular, the retrieval is not something that you'll find in any book, even the book that I just published. Like even in there, I don't talk about this particular retrieval mechanism, but it took us some additional thinking, but we got there.
Yeah. So again, moral of the story, try harder.
Yeah. Just think carefully.
Yeah. Allright. Last thing, just to wrap up, what is your hot take of the day for the closing out the AI Engineer Summit today?
I definitely stopped building chat interfaces. I think chat is a product, AI is a tool. And so finding ways to, once again, I know I've said this before, but like improve on the machine, the machine interfaces so that developers can actually benefit and use AI more directly as opposed to building chat everywhere.
Love that. Mine is a little bit mean-spirited, and so I apologize in advance. I think
a lot of the work that's in front of you as you're building out AI capabilities is going to be incredibly boring. And I think you should be prepared for that. The capability is really exciting. The possibilities are amazing.
And it's always been like this in ML. The journey feels very tedious. It's worth it in the end. It's so fun, but there's a lot of data engineering work in front of you. And I think people haven't yet appreciated how important that is.
Yeah. No, I think it's a very real and very fair take as all of us try to start hopefully moving into production with a bunch of this stuff. That's where the rubber meets the road,right? Well, that's all for us, I think.
Thank you so much, the two of you, for coming up here with me.





