About Cohere0:00
Um, hey folks, this is Vivek. I'm super excited to chat with all of y'all. We'll give a quick talk about what Cohere is all about, and we'll make sure we have enough time to chat about your production challenges with these models, or anything else you want to chat about.
Cool. Quick intro: we are a leading data security-focused enterprise AI company. Our focus is building trustworthy enterprise AI models with our partners for real-world business use cases. And we work with a lot of, like, strategics across various clouds, and our— have a bit of a Switzerland play when it comes to where we can ship and deploy, and I'll talk to that a little later.
So we have a crack team of ML researchers and seasoned enterprise operators. Aiden, who's our co-founder and CEO, was one of the authors on the Seminole Transformer paper. We have Phil, who's our chief scientist; he was an NLP lead at DeepMind and a professor at Oxford.
Nils, who leads a lot of our retrieval efforts, was built expert. And then Patrick was the co-author on the RAG paper, which is a big focus for us at Cohere. So, fantastic team that's helping us build lots of amazing things.
So here's a quick overview of our product line. We have two main arcs: the generative side and then the advanced retrieval models. On the generative side, we have two flagship models, which are Command R and R+. R is our world-class, super low-cost model that's great for most enterprise use cases.
Product Line1:37
And then we have R+, which is your more powerful, larger model for more complex, like, reasoning, tool-use, RAG use cases. And on the retrieval side, most people have used some form of an embedding model, and I'll get into that a little later in the talk.
But the ReRanker is something special that I haven't seen quite often in the market, but we think it adds a lot of value, especially to your RAG pipelines. So we'll get into that too. So when we built at Cohere, we have five guiding ethoses as to how we want to go about things.
Guiding Principles2:38
The first is, obviously, everybody's testing their models and all sorts of academic benchmarks. But for us, what is really important is the performance on enterprise use cases. So we've worked with a lot of our partners to ensure that we have an eval suite that is highly customized to enterprise use cases across, like, let's say, health, HR, finance.
And that's a bit of our goalpost as we ship each of these, like, model versions, and we constantly benchmark on how we're performing at each of these industries and use cases that our customers care about. And then when— the next thing is all about efficiency and scalability,right?
We're not particularly chasing the race for having the largest model out there, but what we really care about is the practical use of these models,right? How do these models get used, and how cheap is it, and how easy is it for you to run it as a customer?
The next big thing is, obviously, customization. You know, as much as we'd like for all of these models to work out of the box, there's always a certain niche that customers want to customize this for. And we offer a variety of things, some of which are pretty intrusive.
We've helped our customers with taking our base model and retraining that with their enterprise-specific data for domain adaptation. We can do a full retraining of the model for you with your data. And then, obviously, the pretty typical last few layers retraining, which is self-serve on our platform.
Data provenance and privacy is another big focus for us. We've worked quite a bit to ensure that all of the data that we've collected for building our models meets up to the enterprise standards, and we offer indemnification for any IP claims that you might run into as a customer.
And then, obviously, we don't ever use any of your data to train our models, so that's a guarantee from us. And deployment flexibility, as I mentioned, we're available on pretty much every major cloud provider. And then we also allow you to deploy on-prem or in your own VPC, wherever your compute and your data is.
Benchmarks4:56
That's where we'll meet you at. Cool. So just a quick look at, you know, your typical performance metrics. As I mentioned, something that enterprises, like, repeatedly tell us is they care about, like, multilingual, they care about, like, RAG and tool use for upcoming, like, agentic use cases.
So a lot of our focus has been in these areas. These are some benchmarks from Hot Pot QA, BambooGoogle, Berkeley function calling that our models are quite good at. And another example of, like, how we try to innovate is we try to make sure that as we are building these stacks, we're incorporating all of the features that people care about out of the box, and they don't have to do extra work.
Citations on RAG is a great example of this. For most people, you have to do a lot of work to actually build this functionality with, like, other APIs, but this really comes out of the box with, like, Cohere's models and APIs.
You don't have to do anything additional as a developer to build, get citations, and which is very important for any RAG-based, like, application.
On the multilingual front, we have one of the best performance when it comes to the fluorescent multilingual evaluation. And we also have a bit of a secret sauce with our tokenizer, which helps keep costs really low,right? And that's, again, TCO is, again, a very big thing for enterprises, and it allows our customers to take that same model, deploy across the globe with their customers, which is very important.
ReRanker Demo6:42
Switching gears towards the embeddings models, again, given Nils and his team have been innovators in this space for a while now, and we've done quite a bit of work to make sure our performance is great on, like, noisy data and at a super low cost in this particular space.
So we're actually pretty excited about what our embeddings models can do, and this is almost always, like, one of the top things that our customers are excited about. But embeddings is a pretty complex space and not without its challenge.
So we tried to build this, like, fun demo where we took all of the archive papers and we asked it a question: When was the attention paper built by, or paper published by Aiden Gomez, who's our founder? And we tried this across, like, a bunch of, like, embeddings models.
So some common patterns that we see is archive is a great example of, like, where you have different kinds of, like, data,right? You have, like, the title, when was the paper published, the dates, you have the various authors, you have, like, the actual paper itself.
And in many ways, this represents the kind of data you might see in enterprises,right? So when you actually build the embeddings for this, you get a fairly complex, like, vector space, and your search queries might not actually, like, map neatly to this,right?
And this is sort of, like, where our ReRanker comes in. So what our ReRanker does is, once you have all of these retrieved documents, it's a cross-encoder that helps you re-rank the retrieved set and make sure that that's the one that you send into your context with the generative model.
And here's the ReRanker in action. So what this demo is showing you is we search for the Transformer paper by Aiden. So you have three different types of, like, search patterns over here: lexical search, then an embeddings-based search, and the Cohere re-rank-based search,right?
And the various forms of these, like, retrievals obviously give you the responses, but they're stacked ranked at different places in the retrieval set, which means the overall accuracy of your RAG system might be low. And re-rank is what's helping you to make sure that isn't the case.
This builds on top of, like, other things that people care about, like chunking strategies, but making those more optimal. Another impact of this ReRanker is, again, total cost of operation, because most of the expense for your models is coming in from the input tokens,right?
And if you were able to, like, narrow in to theright context and do that quickly, you could pass in a very minimal amount of context to your large language model, which drives down your overall cost of, like, operation.
And that is, again, very important in the enterprise setting.
Deployment9:43
Yeah, and when it comes to deployment options, like I said, we have a SaaS API that we can help you manage run your workloads, but then we're also on all of the major cloud AI services: SageMaker, Bedrock, OCI, and private deployment across all of these cloud providers, and also on-premise deployments, if that's something you care about.
Security and privacy, obviously, is pretty top of mind for us, so we make sure we're compliant with the standards that our customers are often asking us for. And then the last bit is just enabling, like, developers. So we obviously have a pretty tight integration with things like LangChain, LlamaIndex, but we also have an open-source, like, toolkit that comes out of the box with various forms of, like, connectors, and that lets you ingest data pretty easily into your systems and lets you have full control over the things you're building and don't have to really look for a ton of different, like, options as you're developing your enterprise applications.
So that's it from me, and I'd love to take any questions or chat. And we also have Sandra here from Cohere, so she'll be happy to help.
Partnerships11:07
Thank you very much.
Yeah.
Maybe two questions. So on the first or second slide, you show your investors and selected partners,right?
Yeah.
I saw Accenture, I saw McKinsey.
Yeah.
Can you explain a little bit, like, how that partnership worked,right?
Yeah.
And then the second question, like, you can skip if someone else, like, has another one, is can you, like, maybe without disclosing customers at all, give us, like, a few samples of where your clients picked up Cohere versus other solutions and kind of why?
So explain where you win,right, on the enterprise world.
Yeah, absolutely. Happy to chat about that. So I think the typical challenge with all of these enterprise, like, generative AI models is the last-mile challenge,right? Like, there's so many different, like, arcs of, like, customization that's needed with the enterprises.
And there's a lot of, like, traditional, like, players who've been around for a while and have great relationships with the enterprises, have a deep understanding of, like, the various, like, business domains. That's where the McKinsey and Accenture and all of these companies come in.
So they've been really helpful for us to co-develop, like, the product, like, make sure that we're able to effectively bridge that, like, last-mile gap with them. And yeah, that hopefully that helps. And then onto your second question, I would say, in terms of, like, winnability, I think, like, the main aspects that have been resonating a lot with our customers is this control over the data and control over the compute,right?
Given we're available pretty much everywhere, a lot of, like, the customers care about that private cloud deployment, the ability to fine-tune in that private cloud environment, which is pretty big for a lot of people, and making sure that their enterprise data does not leave their own ecosystem.
And specific examples for that have been companies in, let's say, HR or healthcare or even folks who are trying to take their in-house, like, code and build a custom model that's working with their code base. Those are the styles of applications that we've seen a lot of, like, impact and success with.
Thank you.
Yeah.
Text Classification13:27
So I'm curious, for tax classification, what is the kind of latest best practice? I see on your website you have a classify endpoint,right? Build a classifier. Is that still the recommendation?
Yeah, that's a great question. The way I like to think about these things is there's always the arc of, like, and I think, like, Jerry in his earlier talk to this,right, which was what phase of, like, development you're in if you're trying to, like, prototype or you're trying to productionize, and what's your scale of, like, a production setting, which is important to consider for these things.
For if you're just trying to, like, get off the ground, like, quickly, I would say just using the generative model off the shelf is obviously always great. It gets you off the ground really quickly. When you're trying to, like, productionize something, that's when I would start thinking, hey, do I need a bespoke model or, like, the general model is good enough, and what are sort of, like, the cost of operation, like, differences and, like, the cost of, like, maintenance differences, and also the scale,right?
Like, I mean, if you're going to try to do something that's, you know, tens of thousands, like, TPS, then having a purpose-built, like, model for that is the route I'd go versus, you know, a more heavy general model, which might serve other needs.
So yeah, both of those are good options depending on what you're trying to accomplish.
Just a quick follow-up.
Yeah.
If we actually train a classifier with you, is that also a Transformer model, just more specialized?
Yeah, exactly.
Okay.
Yeah.
Got it. Thank you.
Nope, I think we're done. Oh, no, we've got one more question. Here we go.
Just a quick question. When we were looking at Cohere, you know, probably early last or middle of last year, one of the challenges with the models that we found were the input context size limits were quite small.
Context Size15:11
Yeah.
How has that evolved as you guys have sort of created the next, you know, sets of models?
Yeah, that's a great question. So our latest generation models are fairly competitive, 128k context input windows, and we're constantly looking to figure out how to, like, up them. So context window, I would say, should be not a problem for, like, most applications at the moment.
Yeah.
Cool. We're actually slightly early, but yeah, I'd like to thank you for giving the talk. Thank you, Vivek.
Yeah, absolutely. Thank you so much. Yeah.
Thank you.
Yeah.





