Intro0:00
I'm Alexander Bricken. I'm on the Applied AI team at Anthropic, so I work very closely with customers to do technical implementation work, and I also bring that advice back to product research and model research. I'm going to pass it over to Joe.
Hey everyone, it's great to be here. My name is Joe Bayley. I work on the go-to-market team at Anthropic. I joined Anthropic over a year ago now, so I've seen our models evolve from 2.1 to today's capabilities. I think day to day what's really exciting is we're working with AI leaders who are solving real business problems that just seemed impossible a year ago.
So really excited about how quickly everything is moving. Okay, for today we will do a quick overview, you know, who we are, our mission, and then we'll focus a lot on implementing AI and best practices and common mistakes.
Alex and I actually didn't just take this from our own experience, but we talked to a number of our colleagues, so this is all based on hundreds and hundreds of customer interactions. So we hope there's some actionable insights to take out of this.
Awesome. So what is Anthropic? So we are an AI safety and research company building the world's best and safest large language models. We were founded a few years ago by some of the leading experts in AI, and since our inception we've not only released multiple iterations of our frontier models, we've done so while being at the bleeding edge of safety techniques, of research, and policy.
Anthropic Overview1:24
I'm going to pass it over to Alex to talk a little bit about our marquee model.
Awesome. And so some of you are probably familiar, but the most recent model we launched was Sonnet 3.5 new in late October of last year. You might be familiar with it because if you're a developer, Sonnet is actually one of the leading models in the code space.
So if you're familiar with evaluations like Sweebench, which is an agentic coding eval, Sonnet is still at the top of the leaderboard for that. I won't go too much into the details on the eval side, so let's keep moving.
So yeah, in addition to what Joe mentioned, we have a lot of different research directions that we're focused on, and these are really distributed but have overlap between, you know, model capabilities, product research, and AI safety. The one that differentiates us, I would say, is the interpretability.
Interpretability2:22
And this realistically is reverse engineering the models and trying to figure out actually how they're thinking and maybe why they're thinking, and then an additional capability in terms of steering them in theright direction depending on a use case.
So let's dive into that a little bit more. We're still very early in interpretability research, it's worth mentioning. As you can see, there's kind of like a longer timeline and we're really only at the first half of that, maybe even the first 25%.
But we're really approaching it in these stages that build upon each other. So these include things like understanding, so grasping AI decision making, detection, so actually being able to understand specific behaviors and put labels on those, steering, so influencing the AI input in some way, shape, or form, and I'll get to an example of that in a second, and then finally explainability.
And that's really where you unlock business value associated with interpretability methods. And so while we see interpretability in the long term providing, you know, a lot of significant improvements in AI safety, reliability, and usability, specifically our interpretability team uses methods to understand feature activations at the model level and then has published research on these in towards model semanticity and scaling model semanticity, which are two papers I highly recommend.
And then as the technology improves into kind of detection landscapes, for example, you can imagine having a much better grasp at the actual thinking and behavior of the model, or even discovering sleeper agents for safety reasons that might be buried deep within model capabilities.
So a good example of that is imagining you ask the model, what were the scores of the NBA matches today,right? And let's say it knows the answer and it says Steph Curry, you know, scored 30 points. This would lead to a feature activation of, for example, feature number 304, famous NBA players.
Realistically, that's a group of neurons activating in a recognizable pattern that we've identified across all mentions of famous basketball players when the model's answering a question, not just Steph Curry. And you also might have heard of Golden Gate Claude.
That was an example of us steering the model, basically amping up the activation in the Golden Gate direction, and thus whenever you'd ask a question like, what should I paint my bedroom, Claude would respond, oh, you should paint it red like the Golden Gate Bridge, and maybe it should have some like, you know, pillars in it or something.
I'm going to pass it over to Joe to talk a little bit about some of the customers we work with.
Yeah. So I'm going to frame this in two ways. One is sort of early on discussions, and the other would be just examples of customers that are doing really cool things. So in conversations, there's obviously a lot of noise and buzz and everything, and that's fantastic, but we often encourage our customers to sort of get back to the basics and how can they use AI to solve the core problem that your product is trying to solve.
Customer Stories5:04
We also get to work with a ton of sort of AI native or AI startups, and this is how they're thinking about their product. And I think you want to move beyond chatbots and summarization. These can be great options, but I'd be thinking more like, where do you want to place bigger bets?
And to give an example, if you just click one more time, fancy slide. Imagine you're an onboarding and upskilling platform. The problem that you solve for customers is you help them get ramped really quickly, and then you help them get to the next phase of their career by equipping them with skills.
So for instance, you might be public speaking, you want to get good at, or you might want to become a manager. And so it would be easy to say, okay, let's summarize course content, or
let's have a Q&A chatbot that answers questions along the way, and they could be helpful, but I would actually think about it differently. So what about if you could hyper-personalize course content based on each individual employee's context? Or if someone is like breezing through all the course content, could you adapt it dynamically to make it more challenging so they're actually getting more value out of it?
And then the last one that I particularly like would be, what if you could dynamically update course material based on people learning about the customer? So if someone was a visual learner, great, let's make visual content for them, and having the AI, having the ML, sorry, the large language model just do that automatically.
And you have to think, does that solve the problem more than summarization or a Q&A chatbot? So really good food for thought. And to sort of talk about some of the customers that we see achieving really industry leading results by combining their own domain expertise and our model.
So I won't read off each, but just a couple of callouts. One is AI impacting different industries. We have taxes, we have legal, we have project management. They're using AI to drastically enhance their customer experience. They make it more like easier to use, they make it more trustworthy, and so it's really improving the experience versus just being like a nice to have.
And then they're achieving a real high quality of output,right? You can't be giving, you can't be hallucinating when you're doing your taxes. It could just, you know, that could lead to all sorts of things. So we're thrilled that they're seeing these sort of like business critical workflows powered by AI driving really positive outcomes for them and also their customers.
Getting Started7:55
Awesome.
I can do this one. Yeah. So
getting started, I just quickly, there's two key points here. So if you go to the next slide, one of our products, we have our API, we have Claude for Work. Our API is for businesses that want to embed AI in their product and services, and then Claude for Work empowers your entire organization to take advantage of AI in their day to day work.
We also have, next one, we also have a partnership with AWS and GCP, and you can kind of get the best of both worlds here. You can access our frontier models on Bedrock or in Vertex. You can deploy these applications in your existing environment, and so you don't have to manage any new infrastructure.
So it really sort of breaks down like any barriers to entry. So you're getting the best of both worlds here. We talk a little bit about support throughout this talk. To us, it doesn't matter if you're accessing us through a third party or our first party, so I just wanted to call that out.
Applied AI8:56
Awesome. So now that we've talked a little bit about some of the customers, how do we actually set customers up for success when working with them at Anthropic? So just a preface on kind of what my team does.
As I mentioned, it's at this intersection of product research, customer facing interaction, and then also just actual research within the org. And we support the technical aspects of the use cases, so helping to design architectures, evals, tweak Claude prompts to get the best out of our models, et cetera.
And then we also bring whatever we see back into Anthropic, and we try to build some of the best products we can for our customers. So some examples of projects we worked on or things that we published include the building effective research paper that my colleague Barry published.
He's going to be speaking tomorrow. And then as well as that, we launched Model Context Protocol, which is an open source protocol for language models to interact with data sources. And Mahesh is going to be leading a workshop on that on Saturday, I believe.
So Anthropic as a whole, we try to effectively support our customers, but where we really start to embed, at least my team in particular, is we work closely with customers that are using Claude a lot, and they're facing really niche challenges in specific use case domains, and they need support from our team to try to apply some of the newest kind of latest and greatest research or get the most out of the models from a prompting standpoint, et cetera.
And so this approach is pretty additive. We often kick off a sprint once the customer is facing those tricky challenges, and that could be LLM ops architectures or evals. We help them define certain metrics that they deem to be important when they're evaluating the model against the use case.
And then finally, we help them deploy that kind of iterative loop, the result of that into an A/B test environment, and then hopefully into production. And so a part of that is the importance of evals, and I'll get onto that in a second, but I'm going to pass it over first to Joe to talk about some stuff that we did for Intercom.
Yeah. So sort of I think this is a good segue in what Alex was describing. So for those of you who don't know, Intercom is an AI customer service platform. They have an AI agent called Fin. By many measures, it's the best in the market, and it's a pretty competitive market.
Intercom Case10:53
So they had their product out for, I think, about a year or so, and when we spoke to them, they shared where they wanted to go, where they saw the future of like customer support and agents. And based on some of the capabilities of our model, we felt that we could have a pretty good impact on these metrics.
And so what we started with was the Applied AI lead met with their data science team, and we ran a quick two week sprint. We took their hardest prompt for Fin and we compared it against a prompt that we helped them sort of figure out with Claude.
And they saw really good results after the first two weeks. So much so we went on this sort of sprint of about two months where we were basically fine tuning and optimizing all of their prompts to get the best performance out of Claude.
At the end of this, they're able to look at all their benchmarks and see that Anthropic was outperforming the current LLM. It's also worth noting that they do a resolution based pricing model. So there's an incentive for everyone for the model to be really helpful and help customers solve problems and not be like a deflection machine where it's like, you know, we've probably all experienced them before.
And so at the end of this two months, they decided to move forward with Anthropic. They launched it. You can read about it. It's called Fin 2. And I think just some of the metrics are really like mind blowing, like can solve up to 86% of customer support volume, 51% out of the box.
Our support team thought about lots of different options, and they actually adopted Fin as well, and they saw very similar resolution rates, but also making it more human. So they, I think with our model, we can, there's much more of a human element to it, so they could do like adjustment of the tone, answer length, and then it was also really good at doing policy awareness, so like refund policy, for instance.
So unlocking some new capabilities. And we're thrilled to be partnering with them as they sort of, I think, march forward as a leader in this space.
Yeah. On a kind of separate note, one of the things I've seen recently is Claude on Twitter acting as some sort of therapist for a lot of people, and I always find that an entertaining example of like its character being expressed.
Yeah.
Evals & Metrics13:09
Cool. So let's get onto some best practices and mistakes that we see in the field on the go-to-market team. So firstly, testing and evaluation. I'm sure those two words have been mentioned a lot today and probably tomorrow too.
There's some typical common mistakes that we see customers struggling with. So the first one is they build a really robust workflow. They spent like a bunch of time building some architecture out, and then they're like, okay, now we need to evaluate it.
Like let's build some evals. That's not really how it should work in practice, because your evals are actually the thing that directs you towards a perfect outcome,right? You can't build a whole workflow without evals probably from the get go or very shortly after.
And so, you know, sometimes customers, as a result of struggling with data problems, might not be able to design their evals. You could use Claude to clean that up, do data reconciliation, or they're just, you know, trusting the vibes too much.
Maybe they run a couple queries, they're like, hey, it looks good,right? Are they really testing that on a representative sample though? Like do you have enough samples to say that the thing that you're looking at is statistically significant?
And like, or are you going to, you know, run a hundred things when it actually goes into prod, and then there's going to be like loads of outliers because you didn't actually predict correctly what the customer is going to ask of the model, for example.
So I challenge you to think about your use cases as this sort of latent space,right? Let's take this kind of chart here on the left-hand side of the slide,right? As you explore the latent space with different functions that you can apply to the model, let's say prompt engineering, prompt caching, stuff like that, you're basically moving your kind of position in that latent space around between attractive states, you know, and eventually you want to find an optimized point, but you don't really know where that is,right?
Like if you're changing an instruction, you don't know how the attention mechanism of the transformer is going to eventually result in some different outcome that might not be performant. And so the only way you can truly know that is empirically, and that's through evaluations.
And so I think that's why evaluations are so important, and a lot of people just don't understand that soon enough. In many ways, I actually tell customers, evals are your intellectual property. Like if you want to be competitive in a space, you need to be able to outcompete people by navigating that latent space and finding the attractive state faster than anyone else.
And so part of, you know, how you do that is, well, firstly, setting up some sort of telemetry to back test. Ideally, that architecture is set up in advance, but, you know, you should invest in it. Designing representative test cases.
So let's say you're working on that customer support agent eval. You know, you might have a kid come on your website, let's say you're building it for, and might ask some crazy question like, how do I, you know, kill a zombie in Minecraft?
Like it's totally unrelated to your product. That's still probable, and so you should probably include silly examples like that in your eval set to make sure that your model's actually approaching the response in an appropriate way or rerouting the question, et cetera.
Cool. Moving on to the next one, identifying metrics. So a lot of the time, you know, there's this intelligence, cost, latency triangle of trade-offs that people are trying to move in between. And most organizations can optimize for one or two of those things, but it's very difficult to meet three, at leastright now.
But realistically, that balance should be defined in advance, and you should know that for your specific use case, you're going to make a trade-off between those things. So let's say a customer support use case again. You care about your customer getting response within 10 seconds,right?
More than 10 seconds, I think there's been research done on this. The customer's likely going to just log off the page, and then they won't get the response, and then they'll probably complain about your product to their friends,right?
Whereas if you're looking at a financial research analyst agent, you probably don't care that it works for 10 minutes to come up with the actual response to your question, because the decision being made after that is very important.
It's an allocation of capital, for example. And so the stakes and time sensitivity of the decision should really drive your optimization choices, and, you know, maybe more instruction sets lead to longer latency, but higher performance, et cetera. The other thing is UX could be important,right?
So again, on that customer support agent, because we spoke about Intercom, you could have different ways of circumventing that 10 second to 15 seconds. Specifically, you could add like a little thinking box that bounces around. You could send the customer to another web page in the meantime, have them read something,right?
Like there's loads of ways to distract and kind of push on those boundaries, but you still need to know what that important indicator is, and you need to optimize accordingly.
Finally, fine tuning. So a lot of people, you know, I go into these calls and they're like, oh, we want to do fine tuning. I'm like, oh, here we go again. Fine tuning is not a silver bullet. So it comes at a cost, and most people aren't aware of that cost.
Fine-Tuning17:50
The cost is generally you're doing brain surgery on the model, and thus there can be kind of limitations to its reasoning in other fields outside of the thing you're fine tuning towards. So my encouragement is try other approaches first,right?
Most people, they don't even have their eval set when they're trying to do fine tuning,right? They need to have a clear success criteria in advance, and it's like only if we can't get that in our specific intelligence domain, do we then do fine tuning.
Don't try to boil the ocean in advance. The difference in fine tuned capabilities and the wide variance of which, you know, failure versus success looks like in fine tuning land means that you should be able to justify the cost of fine tuning and the effort of doing it,right?
Getting a team, fine tuning, working with us, for example, you should be able to justify that difference. And so in terms of best practices, you, you know, don't want to let fine tuning slow you down,right? You don't want to say, oh, I'm only going to convert this language model use case if we can actually finally fine tune our model.
It's like, no, no, no, pursue it and then realize that you need to do fine tuning, and then you can just sub in the fine tuned model and then explore other methods first. And there are loads of different methods that, you know, Anthropic as well as other companies are working on these days, and I just wanted to like flash up a few of them as we wrap things up here.
And so I'm not going to go through all of these, but alongside just base prompt engineering, which granted is very important, there are loads of different features or architectures that will change the success of your use case drastically.
So for example, you might not need to sacrifice on intelligence of your model in order to speed it up by like removing instructions if you can just leverage prompt caching and have a 90% reduction in cost and a 50% increase in speed,right?
Or contextual retrieval will drastically improve the performance of your retrieval mechanisms, and thus you feed the information to the model more effectively, and thus it has less of a time processing all the instruction set that you've given it.
So there are quite a few things that you can apply here, and some of them are even out of the box like citations. And then there are also architectural decisions like agentic architectures. You know, Barry, my colleague, he's speaking tomorrow, will have a lot to say on that, but that pretty much does it.
Thank you so much for your time. We'll be in the theater level lounge after this chat for follow-up questions. Anything else from you, Jeff?
Wrap-Up20:22
No, thank you so much.
Cool.





