Intro0:00
Allright, well, hi everyone. Thanks so much for joining us today. Um, so today I get the great opportunity of introducing SambaNova and some of our capabilities around reaching over 1,000 tokens per second using Llama-3. Today I'm going to spend a little bit of time getting you oriented around SambaNova, some of the capabilities that we provide as an AI platform, and also some of the underlying technologies that are providing the means of achieving some of the accomplishments like 1,000 tokens per second.
Before I start jumping into the content, I want to take the opportunity to introduce some of my colleagues really quick. So joining me today, I have Petro Milan, who is a Principal AI Engineer. He's going to also be leading our workshop component and be here as we're getting hands-on with the technology.
I also have Varun Krishna, who is a Senior, uh, Senior Principal AI Solutions Engineer, joining me as well. And I'm Michelle Matern. I'm our Director of Solutions Engineering. I serve our global customer base at SambaNova. So before jumping into our content, I just want to cover what are we going to be talking about today.
Setup1:24
So we're going to start off with a little bit of housekeeping, uh, talk a little bit about the prerequisites we are going to be getting hands-on today, and also introduce our Discord channel, which is where we're going to be communicating with one another and also sharing some content and information that you're going to need, such as files and, and other, uh, like API and, uh, keys and things like that.
Um, I'm going to talk about SambaNova, just get you oriented around who are we and how are we achieving 1,000 tokens per second. Uh, then I'm going to pass it off to Petro. He's going to, uh, talk you through our workshop today, how we're going to get hands-on, get you oriented around, uh, how to get started, uh, and we're actually going to go through a live build and we'll support you along the way.
Uh, we'll also spend some time around questions, uh, just in case that, um, you have any questions about SambaNova, our technology, or anything that we're doing in the hands-on component. So before we get started, I want to talk a little bit about prerequisites.
So first of all, we are going to be using our laptops, so, uh, hopefully you have them today. Uh, and we're also going to need internet access to get through our workshop. Uh, we're going to be working in a Python environment, um, so hopefully we have Python set up and ready.
Um, and also we're going to be needing to install some packages, uh, through pip. Um, we are going to be working in Discord, so if you don't mind, uh, I'm going to give everybody just a second to hopefully get on Discord and, um, and join our channel.
So I'll just give everyone a quick moment to, to get set up.
Once you're set up, just maybe give me a thumbs up so I know.
Question.
Yes.
The Wi-Fi password. Time to build, capital T, uh, capital T, capital B. The general one, not the one that says speaker.
Okay.
AI engineer.
Okay, so just to repeat, uh, it's AI engineer. The password is time to build, with a capital T and a capital B.
No, no, capital T, capital T, capital B.
Okay, all it, it's all capitalized, each word, but concatenated.
Okay.
Okay.
Good? Were you able to join?
Yes.
Okay, awesome. Cool. Just want to make sure there's no problems. Awesome. We'll give everyone a couple minutes. We just got Wi-Fi access.
Allright. Um, in case you haven't got a chance to set up, maybe just take a picture of this really quick. We'll also go back to it, um, before we kick off the hands-on component. Um, but want to spend a little bit of time just getting you oriented around us as SambaNova.
Platform4:31
So SambaNova, we are a full-stack AI platform, and we've existed since 2017. Uh, we were founded out of Stanford University. So two of our, uh, co-founders are actually Stanford professors, uh, Kunle and also Chris. Um, both of them and including Rodrigo each bring a unique perspective to our founding team, uh, including previous, uh, startups that were, uh, acquired and had, uh, uh, various exits, uh, also building other AI startups that are pretty well-known in the industry, such as Snorkel, Together AI, um, and also a really, uh, depth of experience around building out hardware and chips.
Um, we are, uh, building the full stack from the ground up, so that means we build our own chip. Um, I'll go into that a little bit later, later, but also all the way through the system level and the software layer.
Um, we're on our fourth generation, uh, chip, um, and we have built a entire stack that allows you to both fine-tune models, pre-train models, and also deploy those models with really high-performance inference. Um, we have achieved over a billion dollars in funding from various well-known names like BlackRock or Google Ventures, Intel, GIC.
Um, and so we are really well-established to solve the challenge around bu-building and deploying AI hardware. So what exactly are we targeting, and what exactly are we trying to solve for? Our customer base comes from a wide variety of enterprises and also to government organizations.
So we're really aiming to deliver capabilities that can help service the enterprise-grade AI capabilities that companies and governments require to deliver, um, unique and differentiated capabilities, and also things like sovereign AI. Um, our ca our underlying platform is delivering the means to actually achieve the scale of a trillion parameters, uh, plus.
Um, and we're doing that through delivering full-stack capabilities. When I say full-stack, um, many folks have probably utilized many of these technologies on, uh, this slide. We're not necessarily trying to compete with every single layer, uh, involved here.
But what we are trying to do is ease the process of getting started and ease the journey along the way. And so instead of having to make a decision at every single one of these layers, we are actually integrating things into a very seamless experience from deciding on what chip is going to work with, uh, what compute, what compute and chip is going to work with what operation systems, what operation system works with what models.
You don't have to actually make each of these decisions and know that they have to integrate with one another. We actually create this very seamless experience along the way, uh, where everything kind of orchestrates and works very nicely.
And what we're doing at the end of this is actually delivering the capability to not only fine-tune but deliver really, really fast inference capabilities. So we're going to demo a little bit of this later, but, uh, recently we released a, uh, a demo, um, that you can actually go try live and we'll do so later, called Samba1 Turbo.
This is, uh, exceeding world records around, uh, speed of inference, especially when it comes to Llama-3. Um, and you can see that through some of the metrics that were re-recently published, um, through, uh, artificial analysis. Artificial analysis did a, um, a benchmarking exercise to understand the different capabilities, uh, the speed at which they are able to deliver inference throughput, um, for, uh, 1,000 tokens per second across various hardware providers.
And what you can see is that we are far out-exceeding as a platform, um, the through the throughput capabilities compared to some of the other, other providers out there.
And so I want to kind of talk a little bit about how are we enabling such speed. Um, so when it comes to the underlying technology, many of us have experienced some of these trends in the industry. Many of us got started with what you see on theright-hand side of the slide, which is the large monolithic model.
Model Trends8:29
This is the likes of like an OpenAI, for example, or a Gemini or a Claude. And many of us started our LLM journey or our generative AI journey using some of these technologies. But along the way, many, uh, many other capabilities in the open-source community started to pop up, specifically these smaller models.
And these smaller models allowed us to do things like fine-tuning and actually adapting some of these models to our enterprise data and our enterprise requirements. And so that really started to take off in the industry. And each of these started to see different pros and cons associated with them.
When it came to large monolithic models, when we actually started to put these into practice when it came to enterprise applications, one of the reasons many of us leaned into this is because of the broad capabilities that the likes of OpenAI brought,right?
And also the ease of integration in terms of OpenAI into the actual platform itself. It's super easy to manage, and it also is trained on the internet's data, and so it can handle a lot of different things. But when it came to actual enterprise applications, enterprises have unique capabilities required to deliver on some of the use cases and challenges they're trying to solve for.
For example, enterprises have unique data that they, they, you know, all oftentimes segregate from the internet or most of the time segregate from the internet, um, that's proprietary to them. And oftentimes they have spent the last 10 years trying to actually aggregate that data into the likes of data lakes and other, um, uh, kind of centralized capabilities.
And so now how do you actually transform that into AI capabilities that you can leverage? That was very difficult when it came to large monolithic models. The other challenges that we saw is that many started to become very concerned about security when it came to OpenAI.
They wanted to preserve, uh, data privacy. They wanted to also own and use the model as a differentiation for themselves. Um, and then also, as you start to see more and more adoption, that cost just started to skyrocket.
Um, OpenAI and, and a lot of these closed-source models charge on a per-token rate. And so as you start to utilize more and more LLMs, uh, and utilize LLMs, uh, more heavily, the cost just starts to go up and up and up and up, and it's really, really hard to control.
On the other hand, when it came to adopting the smaller expert models or the smaller, uh, open-source models like the likes of Llama-3, 8B, that we'll talk about later, um, we were able to address some of the enterprise accuracy concerns by actually pre-training, fine-tuning these models to adapt them to the enterprise requirements.
CoE11:04
Um, while we were doing this at a smaller scale to address like different capabilities and tasks that we needed to solve for in the enterprise, they weren't trying to solve like a broad set of tasks and also adopt a broad set of general knowledge.
Um, thus, manageability became a little bit challenging because now we have all of these kind of like micro-models that we have to orchestrate and have them work together. Um, but we were able to solve for some things like security, model ownership, uh, data privacy and data ownership.
Um, but again, because we had so many of these and we had to fine-tune each one of these, the cost also became very challenging. So what were what are we trying to solve for, uh, through our capability of Samba1?
We're trying to bring the best of both of these, uh, paradigms together, um, to deliver the, the capabilities of each, um, in a very simplistic way. And the way that we actually deliver this is through four core capabilities.
First of all, we take all of those expert models behind the scenes. So let's just say we've fine-tuned a model for our legal purposes. We've fine-tuned a model for our HR purposes. We've fine-tuned a model for coding capabilities.
Each of those have different groups and tasks and use cases that are going to consume those. But we want to really ease the experience of having to integrate those into the application. So we put all of those behind a secure single endpoint.
And so you only have to interface with one endpoint to gain access to all these various models. Now we need to determine how are we going to actually use those or consume those various models behind the scenes. So now we need, uh, capabilities around orchestration.
So one of the other capabilities we're delivering as a part of this is around, um, routing. So we're delivering the ability to determine based off of an incoming prompt, what is the best suited expert behind the scenes to solve that prompt?
And we're doing so through a router. We're also bringing the means of dynamically fine-tuning. Every expert is going to have a different cadence in which fine-tuning is going to make sense. Um, so maybe your finance model gets adjusted at annually when your policies are updated.
But maybe your coding model, because you're pushing code so regularly, needs to get updated on a quarterly basis. And so you want to actually be able to schedule your fine-tuning jobs and adjust and swap these models at the rate at which it makes sense to actually retrain these models.
And lastly, you have a bunch of models under the scenes. Not every application or group is going to or should be able to access each of those models. So now you need to figure out a way to, uh, manage the access controls for these.
So what we're also deliver-delivering as a part of this capability is model-level RBAC. So you can actually determine this application or this person or this group of people should be able to access this set of models. And it allows you a ton of efficiency from a computation and management operation standpoint with the security and fine-grained control that you need to actually manage ac-access to these different models and data under-underlying these models.
So within Samba1, we have two ways of delivering this. We have something called a flexible CoE that allows you to kind of determine exactly what models lie under the hood. And we also have a pre-composed version of Samba1 composition of experts.
So our pre-trained, um, or our pre-configured, I should say, our pre-composed version of this model has 92 underlying experts. And when I say experts, I'm really referring to a, a, a specific model, um, that can bring different capabilities, uh, associated with it.
And within those 92 experts, we have a broad range of languages that are covered within those models, a broad range of domains, and a diverse set of tasks that are very, very relevant to the enterprise. All of these models are supported by seven different foundation model architectures, including Llama-2, Llama-3, Mistral Falcon, BLOOM, and even some multimodal capabilities such as like LLaVA, CLIP, dplot, um, that is kind of starting to support some of the multimodal trends that are coming.
And of all these 92 models, we are actually delivering and partnering with organizations to, um, create and contribute back to the open-source community. So out of the 92 experts, 12 of them are actually ones that we've helped, uh, either develop ourselves or co-develop with organizations out there, including some of the language capabilities that we've delivered, such as models that can support things like Thai or Japanese, Hungarian.
Um, and through that, we've developed a lot of experience on how to actually adapt models to different languages. Um, we've also, uh, created a model for text-to-SQL capabilities and delivered, uh, really, really good results through that model for text-to-SQL.
And lastly, we have contributed back to BLOOM chat. Uh, it's the second largest open-source model, um, and, uh, it's also bringing a lot of the multilingual capabilities.
So organizations essentially, as they're constructing these different compositions of experts, can add as many expert models as they need. As I mentioned before, while we have that pre-configured composition, this is really intended for you to be able to construct exactly what you need in terms of models under the hood.
So you can add as many as you need. So our end goal with this is to really be able to bring the capabilities that enterprises need to, um, handle the diverse set of use cases and capabilities they need to solve their problems.
And so one of the things that we've created along the way to measure ourselves against this is an enterprise-grade AI benchmarking set. And this is really tailored to understand our capabilities against the best-in-industry models across various, uh, enterprise-specific tasks and domains that are needed.
Things like information extraction, that's where many, many enterprises are starting, um, but also broad set of capabilities like text-to-SQL, coding, function calling. Um, and we're measuring ourselves against GPT-3.5 Turbo and GPT-4. Um, and what we can see is along the way, um, we're meeting or exceeding, um, the capabilities that, that OpenAI is bringing.
Um, but alongside the capabilities, you also need to orchestrate these models. I talked a little bit about this before, but one of the deliverables as far as Samba1, uh, from a product standpoint is we're bringing routing capabilities so that you can actually, uh, take a prompt, determine what is the best suited expert, and then route to that.
Um, there's also scenarios where you may need to do something outside of routing. You may actually just want to directly call a specific model. Or in the case that is popping up really, really regularly now with agentic AI, you may need to do model chaining.
So that's another capability that we're bringing as a part of the product suite for Samba1.
Hardware18:37
And so how are how are we actually uniquely set up to deliver the composition of experts capability, um, and also the really fast inference speed that we're going to see, uh, briefly? So I mentioned earlier, but we are a full-stack, uh, AI platform, but we also build our own chip.
Our chip, we call that an RDU or a reconfigurable data flow unit. So instead of a GPU, we'll refer to our chip as an RDU. And our current version of that chip is called SN40L. And what is unique about SN40L is actually the memory structure of this chip.
So our chip supports a, uh, three-tiered, uh, memory architecture. So we have our on-chip memory, and then we have our high-bandwidth memory, and then we have a huge, uh, memory capacity in DDR. And this really allows us to store a ton of models, um, and have those models be swapped in and out of various tiers within the memory, um, to achieve really strong performance and also really, really efficient compute utilization.
So what does this look like? Um, as I mentioned before, we can store up to 5 trillion parameters on DDR. So that's like if we were to store like three to four OpenAIs on a single chip. And then as we need to actually execute those models, they can move up the memory stack.
Um, and this allows us to, one, take into consideration what models need to be used at what time, and then two, do so in a really performant way because the, the network connectivity between these three layers is really tight.
So just to kind of go into some of the specs, um, our on-on-chip SRAM, uh, has 4 gigs. Uh, our high-bandwidth memory has 512 gigs, and our DDR has up to 6 terabytes. So, um, lots of memory to work with.
And when we think about how does this compare to what you experience in the GPU world, when you think about the number, if you want to host a really large amount of models, when you have to do so on a GPU, you're basically needing to work with the memory that, that is that the GPU has on chip.
But with us, because we have the DDR component, we'll be able to work with a large amount of models in a single system and keep it coupled with the other memory, uh, tiers. And so with GPUs, you're often, if you wanted to host 500 different models, you're going to have to actually, uh, align those models to various systems, and you're going to have to basically call various systems to actually access each of those models.
With us, it's going to be one underlying system, um, because we're able to store, again, up to 5 trillion parameters.
Live Demo21:32
So, um, I'm going to hand it over to Pedro. Uh, what we're going to do next is actually, uh, get the chance to get hands on.
Um, yeah, thanks, uh, Rochelle for the presentation. Um, so what we will be doing next is a, uh, demo of our Llama 3 and Samba1, um, Turbo. And after this, um, we'll move to the, uh, hands-on portion of the workshop.
Um, so if you want to try out our Llama 3 endpoint, um, so you can go to our website, uh, sambanova.ai,
and then, uh, click on, uh, Samba1 Turbo. So this is where you can access, um, our, um, chat infra. And here you have, um, options to select, um, various models. You know, Llama 3, um, 8B, 70B, our CoE, Mistral, and even some of our, um, in-house, um, models which we train ourselves, like, you know, SambaNova and others.
Um, so yeah, we'll do a demo of Llama 3, um, 8B, and I'll ask it, uh, the following question. Um, so, you know, create a three-day a week workout schedule for intermediate fitness level. Um, let me actually, uh, redo it again.
You can see, you know, it gets, um, instant, uh, response. And then for the, uh, performance metrics, so you can see the, um, insights here. Um, so a few things, uh, to note. Um, so of course, we have a pretty high, uh, throughput, which is, uh, 1,000 tokens, um, per second.
But that's not only, um, the end of the story. You know, we also have a pretty small, uh, time to first token, which is basically the, um, input inference time of 0.09 seconds. And we also have a pretty small, um, end-to-end, uh, total inference time of 0.65 seconds,right?
So with our full-stack, um, platform, we can achieve high throughput and very small, um, inference time. And if you just want to do a comparison, let's say with, um, ChatGPT, uh, for instance, um, if you ask, uh, the same question, just copy it.
Um,
yeah, you can clearly see the, uh, difference in speed.
And the, uh, other, uh, cool thing which, uh, we built, so we have this, uh, real-time, uh, option where as you write your prompt, you can instantly see the model's response. And as you change the prompt, you can also see the change in the response.
Yeah, so let's say, I don't know, hi, I
want to write an email about blah, blah, blah,right? So but the point here is that, you know, with this, um, real-time option, you know, you can do, um, real-time chatting. And this can be helpful, for instance, if you're drafting emails or even if you want to do some, like, real-time, um, prompt engineering, let's say.
So yeah, that's it for the, uh, Samba1, uh, Turbo. I can give maybe a few minutes for folks to try it out before we move on to the, um, hands-on. Um, so again, you go to, uh, our website, sambanova.ai, and then you click on Samba1 Turbo.
Um, so you can do it from your laptop or from your cell phone.
Yes.
Another question. Can we tweak the generation parameters like?
Yes, yes, yes, yeah. Uh,
let me see how to do it on this, uh, UI.
Okay, I think, uh, from this UI, it seems it is fixed, but, um, in the hands-on, once, uh, you will be calling our endpoint, um, from the API key, we will be changing some of the, um, configs as well.
Basic Exercise25:39
So
yeah, any other, uh, question?
Yeah, were you able to, uh, try it out, uh? Okay, great. Okay, so I think, um, yeah, we can move on to the, um, hands-on, uh, portion. Um, so what I'll do first is just, um, introduce what, uh, I'll be talking about in this part, and then, uh, we'll dive in, um, to the, uh, hands-on.
So yeah, we prepared two, uh, exercises, um, for you today. So the first one is a basic, um, example to get you started. Um, so here, um, I'll be showing how you can load, um, environment variables, set up the, uh, SambaNova API key, um, initialize the LLM, and do a simple, um, uh, inference call in Python.
And the, uh, second, um, example is a more practical one. So we will be, uh, building and deploying a, uh, Q&A system, um, with RAG for, um, enterprise search, uh, using our platform. And we will also be using, um, other libraries and packages like, you know, LangChain, um, various data loaders, uh, E5-large-v2 embedding, um, ChromaDB vector store, and of course, the Llama 3, um, endpoint, which runs at the speed of 1,000 tokens per second.
So yeah, let's start with the basic, um, example. And, um, if you want to follow along, so you can, uh, yeah, go to Google, write AI starter kit, uh, SambaNova, and then, uh, click on this link, or you can just, um, write this, uh, URL.
So this is our, uh, starter kit, uh, repo. Um, we have a collection of open-source, um, examples on, um, GenAI apps. And, um, yeah, once you get there, yeah, you can go to, uh, workshops, AI engineer 2024, um, basic examples.
So I'll go over the README and then do a live demo, um, of the work of this, um, exercise, and then I'll give you, uh, some time to, uh, try it out. So, um, yeah, first, um, you'll need to clone, uh, this, uh, repo here.
And then, uh, you'll need to, uh, create a, um, .env file, um, in the repo root, uh, directory. So this is where, um, we will be, uh, specifying the, uh, Samba Studio, um, API key. Um, yeah, so let me show you how this is done.
Uh, it's going to be
yeah, I guess with the two mics, uh, I need to find a way.
Yeah.
Yeah, so I've already cloned, um, the repo yeah, the repo. And then, uh, at this level, this is where you will need to create the, uh, .env file. So vi well, in any case, yeah, you'll have to do touch .env.
And then
yeah, you can add, uh, the these information here. So for the first, um, hands-on, that's the, uh, um, only thing that you'll need. Um, and you can access or copy those from our, um, Discord, uh, channel. So
yeah, if you go to Discord,
events, um, yeah, so you can copy either, um, of these keys. So we have two dedicated, you know, endpoints, um, for this workshop. Allright. And, um, once we finish setting up the .env, uh, the third step is basically, you know, installing the, um, packages.
So for this one, you can either do it with, you know, Conda or a Python environment. Um, so first, yeah, we will go to the basic examples, um, repo. So just CD, and then the repo path. Then you can create, um, a Conda, uh, environment.
So I would recommend using, uh, Python, um, 3.10. And then you activate your, uh, Condo environment. Um, and here we name it basic_X. And then, yeah, you just, um, install the requirements with the, uh, pip install -r, um, requirements.
And then you can use this, um, line to just, um, link the kernel to your, um, notebook. Okay? So if I go to my terminal
yeah, so we'll go to, uh, workshop, AI engineer, basic examples.
Okay, so this is where you can create the Conda. So I've already done it, um, beforehand. So I'm just going to activate the, uh, environment. So
yeah, and this is the, uh, requirements file. So we only have, like, a few packages that is needed.
Yeah, and that's it, uh, for the installation. Um, so once this is done, yeah, you should be able to open, uh, the notebook. And again, you can do it, you know, from the terminal, which will route you to a browser, or, um, uh, through VS Code.
So if you want to do it with a terminal, you just write Jupyter notebook, and then
the name of the notebook. So it's going to be, uh, example, uh, with, uh, sambastudio.ipy, uh, 1b. I'm actually going to do the demo via, uh, VS Code, just because I can show you the, uh, time step. So I already have this, uh, set up.
Okay? So, you know, we are in the basic examples repo, and then example with Samba Studio. And again, if you're doing it, uh, with VS Code, just make sure that you have the, uh, kernel set up. Um, so this is a pretty basic, um, script.
So we will first, uh, let me just restart it. Yeah. So yeah, we will be loading,
uh, the libraries. Um, so we actually have a wrapper with LangChain. So this is, um, where, uh, we will be, um, initializing and calling our endpoints. The second step is to load the, um, environment variables. So these are actually the information which we added in the, uh, .env, uh, file.
And then, uh, we will initialize the, um, LLM. So yeah, we will be using the Samba Studio, um, wrapper. And, uh, we set the, um, Samba Studio API key, and then we specify the model config. So I had a question earlier about, um, the model configs.
Um, so this is where we can, uh, set this up and play with this. Um, again, I think most of you are familiar with these configs. So, you know, do a sample. If you set it to false, this is basically a, um, deterministic output.
If you set it to true, then it becomes, um, probabilistic. You can change the temperature and also the max tokens, um, to generate. And also, this is basically our CLI endpoint,right? So we have one endpoint through which you can actually call, um, different models.
And in this case, uh, we set the export to, uh, Meta Llama 3, um, 8B instruct. Okay? So I'll be running the, uh, this step. And yeah, we have now our model loaded. And now we're, uh, good to go.
So, um, I'll first show you how you can do an inference call using a simple, um, invoke method in LangChain. Um, so just write LLM.invoke and then add your prompt. And the prompt here is, um, what is the capital of, uh, France?
And what you'll, uh, notice is it gives theright answer, but also, um, give you other stuff which you didn't ask for. And this is a common thing with open-source, um, models because when you ask a prompt, you need to include the, uh, the special tags,right?
So in the case of, uh, Llama 3, you can get it from, uh, Meta model card. So you'll have to actually, in the prompt, add these, uh, special tags or tokens. Um, in particular, you have to let the LLM know that this is the beginning of text.
And this is where you will insert the user query,right? So and then let the LLM know where it, uh, needs to, um, answer. And once you have the special tags, um, inserted, um, now you should be able to get theright response.
So this is the first way to do the, um, inference call. The other way is to do it via a, um, LCL in LangChain. So basically, um, you can use the LCL to connect, um, a prompt template with LLM and an output parser.
And LangChain has very, um, templates that you can use. So, um, we'll asking the same prompt. It's just that, uh, the main difference here, we are adding the country as a placeholder. And then when you prompt the model, you can actually specify the value of this country,right?
Yeah, so that's it for the basic, um, example. And as I said, this is just to get you started. So, um, yeah, we can spend 10 minutes for you guys to try it out. Um, me, Varun, and Orshel will be here to help.
And, um, yeah, then we can move on to the second, um, exercise.
I can use any end if you're available.
Um, it might just be turning it up, I'm sorry.
Oh, that's it. That's it. It could be deleted. It could be, uh, issues.
Um.
Yeah, I'll move around if also people have questions.
Do you think it's a connected relation? It's not the endpoint, it's not.
Do you have theright endpoint set, basically? Uh.
Um, yeah, let me just bring my.
Can you open your .env? Well, actually, it's not working,right? I don't know.
Yeah, just, um, compare with.
Um, yeah, can you try the E4? The, the, the other one.
Oh, yeah. Actually, the one here is outdated, basically. So if you're using the one, the GitHub is outdated. So yeah, you, you yeah, you need to use the one from Discord. Yeah, because yeah, this is, like, public, so everyone can actually, uh yeah, sorry about that.
Yeah, so you can use, um, anyone. You have it working,right?
Yeah.
Yeah, yeah, yeah.
Yeah, yeah, got it. Um, yeah, got that.
And yeah, basically, for, you know, this workshop, the, uh, CLI endpoint, we've only, uh, activated the Llama 3. But eventually, um, like, if you were to activate the whole CLI against just the same endpoint, then you can just switch between models, basically, you know?
So, um.
What are the other benchmarks for the Llama 3 AP in terms of tokens per second? Like.
You mean, uh, performance?
Yeah, yeah.
Yeah. So I mean, there's the, uh, GPU numbers,right, which I think, uh, Orshel presented in the slide deck. Um, so we have, like, I think, 8 214x speed up with GPUs. And there are also, like, you know, other startups who have, um, also, like, high inference speeds, like, like, basically Grok.
But, you know, one difference we have with the Grok is they run with 576 chips. We only do it with 16 chips, basically. And you actually conserve the full precision. They do it with reduced, uh, precision, you know?
So, um.
This is, like, weird to get a response that quickly.
Yeah, yeah, yeah, yeah.
Yeah, that feels like a fake.
Yeah.
Were you able to, uh, get thatright, or were you not trying good? That's fine. But yeah, let me know if you have questions.
Okay, nice.
Yeah, maybe you can have our stuff there as well. Yeah, yeah. So now you can yeah. But this is simple if you want to try it out, you know? So yeah. Okay, so I'm ready for the migrate. Yeah.
Uh, which one are we using?
Let's try the two that are in here, the P2 and E4.
Yeah. Then maybe your internet connection, then.
I mean, I could pull up your webpage very quickly.
Yeah.
Yeah, no, it's not.
Uh, can you restart the, uh, notebook and then.
Yeah, you did that. Okay, let's just try the next one. I wonder if it's, um, if it's authenticating.
Yeah, like, this should be faster. But anyway, let's, let's.
No, yeah, I mean, this one should take, yeah, 3 seconds. Yeah, that's it. Yeah. But because this is without the special tags,right? So because it's a long, um yeah, but this one, yeah, should be just, yeah, quick, you know?
So yeah, yeah. Yeah, yeah, yeah. No, because
no, it's basically waiting for the whole thing to complete, and then we're outputting the answer, you know? So, um, yeah.
Yeah, so the same question.
Do you have questions? Were you able to try it out, or just, uh, okay. No, let me know if you have questions. Sure. Thanks.
Yeah, so the first one, yeah, it's the prompt template. So the same question has been wrapped into what is exactly the prompt? It's been wrapped into, um, so prompt template is basically, um, template. And using those tags. So every, uh, language model that we use contains this specific template.
So we need to provide an instruction to the model in that template. So that's how the model regresses the template. If you don't provide any template, the model just kind of gives you, um, another template. And this is, this is kind of like a, it's kind of specific to, um, this model modeling.
So if you provide this, the other, um, you know, it's a very particular way that you get. So that's what we're doing.
Okay. In the next, um, I want you to run with this one. In the next example, you can run with this. Uh, in the next.
We should move to the second one, or maybe just because I don't want to.
And that would be basically the.
Yeah.
Yeah, just, just follow the theory. And the key thing to, like, change the prompt and, um, you know, play around with it.
So what is the step of changing the prompt is the.
Yeah, so I, I mean, you can get it from the.
Llama card.
Yeah, so if you go to Google and then, uh, write Meta 3 Llama card, you, you get it there.
Okay, so it's not the entire.
Yeah. Yeah, like.
Yeah, different models will have different methods. Okay. Yeah, it's good. If you're using Llama 2 here, we have to use Llama 3.
I think we can move to the second one since everyone tried it out. Yeah.
Yeah, like, I yeah, yesterday.
Okay, cool. I think, uh, many of you were able to try out this simple, um, exercise. So yeah, we'll move now to the, uh, second, uh, one. Everyone, do you want to, um, yeah.
RAG Overview43:41
Yeah, so, uh, as I said earlier, this is going to be a, uh, Q&A, uh, system, um, with RAG. And, uh, if you want to follow along, um, again, from the same repo, uh, go to Workshop, AI Engineer 2024, and then EKR RAG.
And like the previous exercise, I'll go through the README, do a live demo of the installation and the run setup, and then, uh, give you some time to, um, try it out. And the, uh, app here, we have two versions of it.
One with a Jupyter notebook and the other one, um, with a Streamlit, which is a UI-based. And before I, uh, jump into the hands-on, just wanted to give a brief, um, overview of what RAG is. Although I'm sure many of you already know, uh, this concept, but just for, um, completeness.
So yeah, so RAG is a technique that, um, you can use to supplement, um, LLM with, um, additional information from, um, various sources to improve the model's response. And RAG is very helpful, um, if you want to use an off-the-shelf LLM to ask it questions, um, beyond its, uh, training data, or if you want to, um, have the LLM access to up-to-date information without, um, retraining it.
Also, you know, RAG can help reduce, um, hallucinations in some, uh, contexts. And a typical, uh, RAG workflow, um, consists of, yeah, the, uh, following steps. So we first have, um, document loading and parsing. So this is where, um, we can use a data loader to actually load the data into a, um, digital text that we can edit and format.
Um, and, you know, various data loaders are, um, available depending on the, um, extension of the file you're using. So in a PDF, text, um, PowerPoint. After this, we have a splitting step. So this is where, um, we will be splitting the document into, um, smaller chunks.
And, you know, the chunk size and the overlap, all of these are, um, hyperparams. And the, uh, next step is, um, vectorization. So this is where, um, we will be using a, um, embedding model like E5-large-v2 to map each, uh, chunk to a, a numerical vector.
And, um, we can store, you know, the vectors along with the content and the metadata in a vector store, like FAISS and ChromaDB. And today, um, we will be using, um, ChromaDB, which is, um, open source. And again, the whole, uh, goal of this embedding is that it allows us to do, um, like, semantic similarity and, uh, semantic search.
And in the retrieval step, um, this is where we ask, uh, the question, which is going to also be embedded into the vector, um, space. And then we have a retriever, which is going to, um, retrieve the closest, uh, chunk vectors to the query vector according to some, uh, similarity metric.
And we can also add a re-ranker, which can re-rank the retrieved, uh, chunks, um, per, uh, relevance and also remove some of the, um, unnecessary chunks. And the last step is basically Q&A, um, generation. So this is where we provide the LLM, um, with the query and the final retrieved, uh, chunk to get the, uh, grounded response.
And, um, yeah, maybe also would like to precise that, um, in this exercise, so the we will be using third-party tools for document loading, splitting, and, uh, storage. For the embedding model, um, you can either run it on CPU or on our hardware, and we'll be doing both to show the differences.
And for the LLM part, this is going to be done on our, um, hardware.
So yeah, that's it for the, uh, overview. I think, yeah, we can move on to the, uh, README. Um, yeah, so we first, uh, cloned, uh, the repo. I think if you've done the other, um, exercise, then you don't have to do, um, this step.
RAG Setup48:21
Um, same thing. Um, after this, we will set up the, um, environment, um, variable. Um, so for this one, we will be using, uh, Samba Studio, um, API for the LLM. And also, uh, we will be using a embedding, um, API as well.
So both are available on Discord.
Yeah, let me show you, uh, in the terminal.
Yeah. And actually, yeah, I would also recommend to deactivate your previous, um, environment.
Okay, so we go to the EKR RAG, uh, repo. Um, yeah, actually, for the .NET file, this one you'll have to put it, um, at this level, the AI certificate. Okay? So for this one, we will need the embed endpoint and API key and also the, uh, Samba Studio, um, endpoint and, and, and API key.
Allright, then we go back to the, uh,
EKR RAG folder, and then we are ready to go with the installation. So here, we will be needing more packages. Um, so first, um, Tessaract. Uh, this is our OCR, um, data extractor. So let's say you're using a Mac, um, just run brew, install Tessaract.
And this one, you can do it outside your local, um, environment. It should take you, you know, a few minutes to install. And you also need, um, Poplar if you don't have it already. Um, you can just do brew, install, um, Poplar.
And then, yeah, we will need to set our, um, virtual environment. So since this is a more complicated, um, exercise, I would just recommend to use, um, the, you know, default option. So the, uh, Python environment with a Python, um, 3.10.
So if you don't have, uh, Python 3.10, um, let's say on a Mac, you can install it using this, uh, command here. And then you can add the path to your shell, like bashrc or, uh, zshrc, um, using this command here.
And then you can just source your, uh, shell file. Okay? And then, yeah, we will, uh, go to the, uh, repo if you haven't done already, um, create your Python environment. Um, so again, if you added this step here, then your laptop should recognize the Python 3.10, then minus m, venv, and then the name of the environment.
You activate that environment, and then you run the install script. So this should take, I would say, you know, five minutes if you have, um, good internet. And once this is done, you'll also need to install ipy kernel and also, um, link your kernel to your, um, notebook.
So when we tested this on different laptops, um, you know, some folks were having also sometimes NLTK and SSL certificate. So you might also need to, uh, run this script here.
Yeah, it's in my bag, basically.
Yeah, so let's activate the, uh, uh, conda the Python environment.
Yeah, and this is the, uh, requirements file. And as, and as you can see,right, we have, uh, more packages here. And then, um, yeah, this is the file which you may need to also, um, run as well.
Yeah, and once this is set up, that's all you need to, uh, run. It's fine. I can do it. Okay. Yeah, uh, the, the notebook. And as I said earlier,right, so we, uh, have the app in a notebook and in a Streamlit.
So for the, uh, notebook, um, again,right, you can open it from the terminal or, uh, from VS Code. Yeah, let me do it, um, from VS Code.
Yeah, so you'll go to the EKR RAG, uh, repo, notebooks, and then, um, rag lcl.ipy1b. And this is going to be our main, uh, script, which is using actually, you know, files and models from other, um, files. So in particular, uh, we will be using the, um, the document retrieval.py and also, uh, some files, uh, from the vector DB, which I'll explain in, in more details.
Allright, so, um, let's go maybe first over the structure of the notebook. Um, so we first, um, import, uh, the libraries and set the required path. So you don't have to do, um, anything at this point, but yeah, just know that the Git directory, this is the absolute path for your EKR RAG.
And then the repo directory, this is the absolute path for the, um, AI certificate. So let's run this, uh, repo.
And then for the, uh, document, uh, loading and splitting, um, yeah, so we added, um, you know, various, um, data loaders, you know, like pypdf and unstructured. And which, uh, data loader you want, uh, you can set this up, um, in the config file, which is, um, here.
So yeah, I'm just going to do a test with pypdf2 for now, but we can switch to other, um, data loaders, um, afterwards.
And, um, yeah, for the experiment, I will be using the SN40L, um, paper. So this is an archive paper which we, um, recently, uh, submitted. So this contains, like, information about the stack, the hardware, and our, um, uh, COE.
Um, I can show you, uh, the paper as well. Um, we also have it in, uh, GitHub, but you can also, um, upload your own, um, PDF. And yeah, you'll have it to put it under data temp, and then
yeah, this is the, uh, paper that I'll be using, uh, uh, for the demo.
Allright, and, uh, let's go back to our VS Code. So this is where, um, you know, we will be using pypdf to actually load the content, um, into a list. And then, um, we are using the, uh, recursive, uh, splitter from, um, LangChain.
So what I'll do first is going to the notebook, and then we can go into the functions in more details if you are, um, interested. Yeah, so let's run this step.
Yeah, and, um, in the config, um, so I set the, uh, chunk size to 1,200 and then the, uh, chunk overlap to 240, but again, you can change those configs. So for this 15-page, um, PDF, we end up getting, um, 89 chunks.
RAG Demo56:01
The next step is the, uh, vectorization and storage. So this is where, um, we will be using the embedding model to map each, uh, chunk to a embedding vector. And as I said earlier,right, so we can actually run the embedding model either on CPU or on RDU.
So RDU is basically our AI chip. So if you want to do it on RDU, then you'll have to go to the config.yaml file and then, uh, set the type to, uh, Samba Studio. And then, uh, batch size.
So this one, you can have it either 1 or, uh, 32. So 32 means that we are actually processing 32 chunks at the same time. And this is a standalone, um, endpoint. So yeah, the COE is set to, um, false.
And I'll show later, um, if you want to run the endpoint, um, for the embedding on your laptop, uh, how you can change the configs, um, for that. Allright, and then, yeah, let's run the, uh, vectorization. And after this, we're actually storing or indexing the embedding vectors, um, into the ChromaDB vector store.
I think, uh, this should take around, uh, 10 seconds. Let me see what's happening.
Yeah, yeah.
Yeah, always good to, uh, restart the notebooks. I'll just go over the steps again.
Okay.
Yeah, so it took us, you know, four seconds to embed the whole thing. And, uh, yeah, this is where, um, we will initialize our, um, QA chain. Um, so again, we have different, uh, wrappers and classes, which I can go over it in details, um, afterwards, but for now, let's just execute, uh, the cell.
Yeah, now you're ready to go. Um, ask a question. And then what happens is,right, through this QA chain, the question gets embedded to the vector, uh, space. We retrieve the, uh, top K chunks. So in this experiment, yeah, we have this set to three since I won't be using a reranker, but I can also show you how to use the reranker.
And then, yeah, the question and the context are provided as, uh, basically, uh, context for the LLM to get the answer. So yeah, what is a monolithic model? And you can see that,right, the response is, um, instantaneous. Yeah, so that's it for the, um, uh, experiment.
Uh, let's maybe try ask it a bit more complicated question. So if I open the PDF again, uh,
I think, yeah, there was a table in the PDF.
Yeah, like this is a table, you know, showing, um, operation intensity versus fusion level. So yeah, let's see if pypdf is able to, um, uh, retrieve, uh, some of the information from this table here. So I already have the, uh, questions prepared.
Let's try to access those. Yeah, so yeah, I got the response,right? So 410.4, basically. Um, and again, if you end up having, like, more complicated, um, tables in PDF, um, in this case, I would recommend to switch to the unstructured, um, data loader.
And for this, yeah, all you have to do is just go to the config file and then, um, change, uh, pypdf to, um, unstructured. So yeah, that's it for, uh, the second, um, exercise. Um,
I can go into more details about each, uh, function if you're interested, um, or have you try it out first and then you can maybe, uh, come back and then, um, go over the, the, uh, functions. Um, yeah, so do you want to maybe try it out first, I guess?
Okay, great.





