Intro0:00
Thank you so much for coming to the workshop. My name is Gabriela de Queiroz, and I'm Director of AI at Microsoft. I have Pamela here.
I'm Pamela, and I'm a Python Cloud Advocate. It's a well-known name when you say Python. But I also worked in JavaScript for them for quite a long time, and I generally like lots of languages. And then you get a help here.
Hi. I'm Harald. I'm a PM on VS Code and GitHub Copilot chat.
Awesome. So today we are going to be talking or showing you how to run an AI application in minutes. So we are going to have a lot of hands-on, so be ready to do some coding. Not coding, but going through some coding, using different tools, GitHub Codespaces, Azure, and other tools that we are going to be talking about.
But just to give an overview of the agenda, I'm going to be talking about Microsoft for startups a little bit, some of the partnerships, some of the pain points, and then we'll go through the AI templates and hands-on.
So Microsoft has a program for startups. So if you have an idea, if you have a startup, you can apply to this program. And what I always tell people is you don't have to have a startup per se, but if you have an idea, that's enough to apply for this program.
And you get a lot of benefits, and benefits that can be I'll just skip. It can be like credits, so you get up to $150,000 in Azure credits. You also have third-party benefits, like a lot of different tools that you can use, and then, of course, GitHub, Microsoft 365, LinkedIn Premium, and more.
You can use all the different models from OpenAI, but also Llama, models from Cohere, Mistral, and so on. And the piece that I like the most is about the sessions that you can get, one-on-one sessions with people like me or Pamela, that we volunteer our time to share our knowledge with founders.
We can talk about maybe, I don't know, you are hiring, and then I'm an expert in hiring, so you come and talk to me, and I say, hey, these are some of the best practices for you when you are building your team.
Or you can go to technical sessions and ask more technical pieces as well. And inside this platform, we have several things other than the benefits that I mentioned and the guidance that I just mentioned. It's what we call Build with AI.
And inside, we have some AI templates that the idea is
we can help you accelerate the AI application piece with some kind of skeleton in a way, so you have something up and running in a few minutes.
So again, you get cloud credits, you have access to dev tools, you have the AI templates, you have the one-on-one guidance. And no matter where you are in your journey, if you have an idea, if you are already building, or if you're scaling, this program is for you.
You have access to all the cutting-edge AI tools, so you can innovate and streamline your AI development. And on top of the founders of this program that we have, we also have different programs that it's like the next step.
Let's say you are now scaling, growing, and then you use all the credits. What is next? There is a next. We try to guide you through the whole process. So there is something called the Pegasus program, where we help you to cost-sell, go to market, and so on.
And then there are some strategic VC partners and accelerators that we partner with. So we have partnership with Y Combinator, NEO, The Alchemist, et cetera.
Pain points for startups, there are a bunch of them. One of them is you don't have time. You cannot wait to go to market. You have to go as fast as you can. You have a lot of resource constraints.
We have some issues with scalability. You don't have the support and guidance, and that's where we are trying to help you with. So now we are going to go to the fun part. It's like the AI template. So that's where Pamela is going to show you all the amazing things that you can do with all the different tools.
Workshop Setup5:12
Allright. So our goal today is potentially having you deploy maybe even three different templates. OK.
So we have three different ones. I'll just show in the browser which ones we're going to be deploying. So we have starting. We're going to start simple with this chat application here just to make sure everything's up and working.
And then we've got two different RAG applications. One of them is a RAG on a Postgres database, like RAG on a Postgres table that does SQL filter building. And then we have RAG on an unstructured document. So here I've got a RAG on my personal blog, or a RAG on internal company documents, whatever it is that you're going to whatever kind of documents you're going to RAG on.
So those are the three templates we're going to be looking at today. And we have it all set up so that you should be able to deploy those templates without spending any of your own money and doing it all through our credits, which is yay.
Allright. So the first thing you need to do is get this URL. So everybody open this URL on your computer. So it's aka.ms/ai-workshop. It should open up a Word document in the browser that looks like the screenshot you see here.
So you can either type in the URL or scan that QR code and get that open on your machine. So let's make sure everyone's got it open.
Welcome, welcome. So go ahead. Once you've got your computer ready, put this URL in your browser. Harald, maybe you can just memorize it and then help anyone who doesn't have it. Yeah? AI-workshop. OK. So then let me go to the actual doc here.
So the first thing you need is a GitHub account. Does anybody here not have a GitHub account? OK. So everyone here has a GitHub account. Great. If you don't have a GitHub account, you can sign up for one for freeright now, and that should be fine.
The next thing you need is an Azure pass. So this is something that we've got for this workshop, for this conference, and this is going to let you deploy stuff on Azure without spending any of your own money.
So we got passes for $50, and they're valid for seven days. So if you do want to keep hacking after the workshop, you can keep using your pass. And after seven days, it'll disappear just like Cinderella and the Pumpkin.
So in order to get that Azure pass, you do need to have some sort of Microsoft account. So you can use your you can use a personal Microsoft account if you have one. So how do you tell which one you're logged intoright now?
I guess if you just go to outlook.office.com, maybe you know what Microsoft account you're currently logged into. And then you can see. Some people in the last workshop were logged into their kid's Minecraft account. So just you need a Microsoft account, and you might want to double-check to see which one you're currently signed into if you are signed into a Microsoft account.
If you don't have a Microsoft account, no big deal. You can make one on the spot. I made one this morning. So if you do need to make one, you can just make up a new Outlook address and set it up that way.
So you can also make it as part of this progress. So we're going to go to this AZ check-in URL, and that's linked from this doc here. So if you don't have this doc, if you just came in, we can help you get this doc open so we can get this URL.
And we're going to spend 10 minutes making sure we get through this step since it can be a little tricky. So when you go to this check-in URL, we put this in the browser. It loads. This is what you're going to see.
And it says I can either create a GitHub account or log in with GitHub. So I'm going to log in with GitHub because I already have a GitHub account, and I'm logged into this browser with it already. So I'm just going to click on that.
And so what that's going to do is create a pass for my GitHub account. And so we get a pass. So each of us will get a different code based off our GitHub account. So this is basically my Azure pass promo code, so I can copy that.
And then there's this button here that says Get on board with Azure. This is the next step, is to click this. And then we get this screen, which says, OK, this is you can start. And when I click this here, it says what my currently logged-in account is.
So this is where you should check to make sure you're happy with what account you're logged in with and you don't want to switch. I don't recommend using a corporate account. If you do have a corporate account, just don't use it.
It's going to be problematic for various reasons, because corporate accounts may have restrictions that won't let you deploy things. So we do recommend using some sort of personal account or making up a new account. So that's why you see I'm using my Gmail instead of my Microsoft.
So I'll confirm my account, and then I can enter the promo code. And that was from this screen, so I still have this screen open. So I just go there. I paste it in. And then we go S, 6XYYK.
I think it's case-insensitive. Submit. And then it's going to actually fail for me because I've already set this up on this thing here. And if you see this, it's because you've already actually gone through this stage. So for you, it should work the first time.
And then it'll create the Azure account for you. And if it works, then what we can do is go to portal.azure.com, so portal.azure.com.
And we'll see how it loads in. It does a bunch of redirects. And then we can click on subscriptions. And what we should see is there should be at least one subscription that says Azure pass-sponsorship. So that's our key that we have done this correctly.
And as long as we use this subscription, when we're doing our deploys, we will not get charged any money. Well, Microsoft will, but you won't. That's the important part. OK. So we're going to spend 10 minutes to make sure that we can get everyone through this stage so that we're all on the same page going forward.
So if you already got it, that's awesome. You can
Proxy11:44
look at Harald's Facebook profile or something.
I know.
So once you have that set up, the next step is the proxy.
So I'll just show that so that you can start playing with that. So here's it's the next link in here. So the reason we have a proxy is because normally when you're using Azure OpenAI, you actually have to fill out a form and say how you're going to use Azure OpenAI.
And then somebody says, oh, OK, yeah, that's a good use of OpenAI, because Microsoft doesn't want people to use AI willy-nilly. So we check to make sure that something adheres to our responsible AI principles. We don't have enough time for you to go through that process while we're in a workshop.
So we've set up an Azure OpenAI proxy that you can use during the workshop with the repos. And we have special instructions for how you can use this proxy with the repos since you can't use the actual Azure OpenAI.
So you can follow the link from the doc and log in with your GitHub account. I'll log out so I can show that. Log in with GitHub.
OK. And it says I'm logged in. And then we have an API key and a proxy endpoint. And that's all we need to be able to use an Azure OpenAI instance. Now, normally, I don't like to use keys, and I tell everybody to avoid them.
But in this situation, we are going to be using keys. And yeah, and these keys will expire at a certain point, so we don't have to worry about them being exposed. Typically, with keys, we'd have to protect them very fiercely so that nobody was using them.
So you can go ahead and log into this and see your registration details. And then you can even play around with the playground. This is really similar to the Azure OpenAI playground or the openai.com playground, if any of you played around with this.
You can see here you can play with the system message. That's how you say, oh, you're an AI assistant that constantly makes pirate jokes.
Yarr. And then we update the system message.
Oh, private. I wonder what it'll do.
There we go. And then let's see. Oh, enter my API key. OK. So we need to enter the key. I actually have never used this before. So we're going to enter the key, not save it, select a model.
OK. So we select a model over here. So we've got 3.5 Turbo.
Oh, I didn't know we had four, too. You set up four as well? Cool. We can use four. Four is better. Allright. And four is slower but better. OK. And then please tell audience about OpenAI. OK.
Allright. And you can see different parameters that we send. And these are all getting sent to the OpenAI SDK. So we say the model. Right here, we've set up two models, GPT-3.5 Turbo, GPT-4. Those are often the ones you're picking between with OpenAI, although now you've got GPT-4.0.
That's a good choice if you're doing something with vision, something multimodal. I wouldn't use it otherwise just based off of some experience we've had with it. But it is a great one. GPT-4.0 is good for vision. So here you can see with the combination of the system message and the user message.
So this is what we call user message. This is what we call system message. Those combined together, we get back a response like this where it describes OpenAI with lots of Rs and maybes and stuff. We can change different parameters here, like how many tokens it should send back.
The temperature is roughly the creativity. Top P is also roughly about creativity. And there's some more advanced stuff there. And you can see how many tokens you used on the way out and how many tokens you got on the response.
So you can play around with this playground
to try stuff out and make sure that you're able to use the key. So this is just linked off of this workshop. So if you go to the workshop proxy, you log in, you'll get your key and your endpoint.
You can go to that playground, and you can play around with the playground to check that that's working. But we just want to make sure everybody now has an Azure pass and is logged into the proxy so that you have a key and an endpoint.
So we'll just check to see if anyone had any issues with that. This step is hopefully level. OK. Allright. So here's the if you're looking for the models, this is generally the page to check. So GPT-4.0, GPT-4. And going down, those are the GPT-4 models.
GPT-5. You're saying there's a GPT-3.5 that supports vision? No, no.
Oh, 4 Turbo with vision. This one? Yeah. So we were using that one, but it's a lot slower. Yeah. So that's why I've started using 4.0. This one.
This is the G8.
Oh. OK. Allright. Yep. So you just want to compare those. So we'll just be using the basic GPT-3.5 and GPT just GPT-3.5 today, actually, and then also the embedding models. OK. So is everybody set up with the proxy?
OK. Allright. So now we're going to actually get something working. So we have this repo here. So you can follow the link from the doc. And it has READMEs for the three different projects that we can deploy. And these READMEs are specific to using them with the Azure OpenAI proxy.
First Deploy17:40
So normally, you can just use the READMEs that are on the repos itself. But because we are using this Azure OpenAI proxy, we do have to use a slightly different setup. So we've made READMEs specific for this workshop.
So we can start off on this OpenAI Chat app quick start and make sure that that's all working. So the first step is to open in GitHub Codespaces. So you can do that by clicking this button here. Have any of you used Codespaces before?
OK. A couple people. So Codespaces will open a VS Code in your browser with a developer environment for that repo. So you can actually use Codespaces on any GitHub repo. So you go to any GitHub repo, you click on code, and you can make a CodeSpace for it.
So it's a way that you can start hacking on any repo very quickly. So you can open this button here to open in Codespaces. And I'll just go ahead and make a new one.
And I'll say create Codespaces.
So this is going to take a few minutes to load, because what it's doing is that it's creating the environment for this repository. It's opening VS Code in the browser. And it's also just setting up VS Code. So if you actually have if you use VS Code locally and you've got extensions that you use locally, it's actually potentially syncing those extensions and enabling them here.
I should probably just not do that, because then it would load faster for me. But yeah, you can see in the bottom here as it's setting up. And we'll just wait for it. So the slowest part of using Codespaces is just the loading.
If you want faster Codespaces, there's prebuilds available as well.
Yeah. And I do have them on the third repo, but I think I don't have it on this one. So I should have remembered to do prebuilds for all the repos.
Right. And the slowest part is probably installing all the dependencies and the builds. It's basically doing all the things you would do when you install it locally, just automate it with a progress bar. And at some point, it will just light up.
Yeah.
Let's see what the you can even watch. Can we watch the logs for this one? Building Codespaces. Whoop. There we go. So if you like this sort of thing, like if you like watching Docker containers build, because that's what it's actually doing.
Everything's a Docker container. So you can actually watch it as it builds everything here. And now it's downloading all the requirements. So these are all the Python requirements. So all the examples that we're going through today have a Python back end and then some sort of JavaScript front end.
This one has what we call a vanilla JavaScript front end, as in I just wrote some JavaScript in a script tag. But then the other ones are much fancier. So they've got a full TypeScript and a build system and React components using the Microsoft Fluent UI web framework.
So you can kind of see the range of front ends there. OK. So you can see it's still going through the process, but at least now we can see the File Explorer has loaded. So we can explore the files here.
And I'll go ahead and show the code. If you're interested in the code, it is in the source folder. We are using a QWERT application. And I think nobody has heard of QWERT. But has anyone here heard of Flask or used Flask?
Great. So QWERT is just the async version of Flask. So it's literally built on top of Flask. And one day, it might be brought back into Flask. And you just take your Flask code and you put asyncs in it.
And then you've got QWERT. That's really how it goes. So if you haven't done async before in Python, async is a way that if you use async with your functions, they become co-routines. And then they can be paused and waited on.
And it's important to use async when we're building applications with AI, because we have these really long blocking calls to an AI API. So we make a call to an LLM and we send off our request. And these LLMs, they can take like two seconds, five seconds, 10 seconds, depending on what we're doing.
And while that's happening, we ideally want to be able to handle other user requests coming in. So that's why we use async frameworks. So if we use an async framework, then while we're making I/O calls, we can handle other user requests that are coming in.
So all of the ones that we see today have an async back end, either QWERT or FastAPI. Anyone heard of FastAPI? It's very, very popular these days. Yeah. So FastAPI is the one most people know of as the async framework.
So I like both of them fairly equally. So I use a mix of both. But I just want to make sure people know about the value of async frameworks. OK. So that's all in the QWERT app folder if you want to look at the code there.
So it is now finished. OK. Anybody else get their Codespaces loaded? Get a couple. OK. Great. Finished configuring. So I can just.
I have to add anything in the terminal?
Yeah. We are going to be using the terminal. And if for some reason your terminal goes away, sometimes it's half to the Codespaces, just click that plusright here. Sometimes my terminal kind of blanks out. So I just click the plus, and that'll give me a new terminal.
Right? Whoop. New terminal. OK. So here we are in the terminal. But actually, the first thing we're going to do is that there's a .env.sample. We're going to make a .env file based off of that. So I'm going to make a new file.
And I can do that using this little New File button up here. So I'll just click that, say New File, and I'll type .env. You could also copy and paste. And then I'm just going to paste the .env in there.
You could even rename .env.sample to .env. I think that's another way. And then we need to fill in these values to match the values of the proxy. So we'll go to the proxy. And let's see, where's my proxy open here?
So here's my proxy. So I'm going to go ahead and fill in this one. That's the endpoint. So the endpoint should start with HTTP and end with /v1 and look like that in the middle. So that's the endpoint.
That's where we'll be sending our OpenAI requests. Then we need the key. So we'll copy that. And it'll look like that or slightly different for you. And then the deployment is going to be the name of the deployment is GPT-3.5 Turbo.
And that's also the name of the model in this case. So has anybody used OpenAI.com? A few people. OK. So on OpenAI.com, you just pick what model you're going to use, and that's all you need. With Azure OpenAI, you have to make deployments based off of the model.
So you actually have a bunch of deployments. And you could actually have multiple deployments of a GPT-3.5 Turbo model that have different names. So when you're working with Azure OpenAI, you have to know the deployment name, not just the model name.
So that's one of the complexities of using Azure OpenAI. But it does give you more flexibility, because you can say, oh, this deployment's going to have 20 tokens per minute. And this one's going to have 30 tokens per minute.
And then you can say which of your colleagues can use what, like if they're all trying to use up your deployment or whatever. So it's more flexibility, but you do have to specify it. OK. So now my .env is set up.
So this is just so that I can run a local server. And I'm putting local server in quotes, because I'm going to run a local server inside GitHub Codespaces. So it's actually running a local server, not on my actual machine, but inside the GitHub Codespaces development environment.
So to do that, I'll grab the command here that's going to run the QWERT app. And just give it. And with Codespaces, you do have to allow. So you'll see this little thing that pops up. So if you ever want to copy/paste, you have to allow for the terminal.
And then I have OK. And then I paste it. And then you can see that it says it's running on this URL. Now, you can't just paste this URL in the browser. I'll show what'll happen. So if I paste in the browser, I'm going to get an error, because this is not running on my local machine.
This is running inside GitHub Codespaces. So you have two ways to get to it. One way is that if you just click on it, like option click, at least on my Mac. So I mouse over, it'll tell me what to do.
Mouse over, option click. So Codespaces will actually detect that you're clicking on a local URL, and it'll turn it into a Codespaces port URL. And it's this funky URL up here. Improved disco for me. And that's actually local for that GitHub machine.
And that's one way of doing it. Another way that you might like more is you go to your Ports tab. And you're going to find it listed here. And we'll see the forwarded address. And we can click on that, or we can even click the globe icon.
And we get to the same URL. So there's many ways you can get to this locally running URL and get to the special Codespaces URL of it. And you can even change your port visibility if you want to share it with a colleague or in a class.
You can change it to public. And then you could actually ping this URL to someone else. Now, this is not a deployed URL. You're not going to use this for your deployed URL. But it's fun. It's good for development.
So now I've got this running locally. And now we can type stuff and be like, what's the weather in San Francisco?
See if it's going to lie.
Oh, good. That was a good answer. I think this has been trained. It refused to answer. It's always good when it refuses to answer something it shouldn't know. So we could go ahead and I could change this now and change the system message.
And let's see, where's our system message in here? Soright now, my system message is just, you are a helpful assistant. You are an assistant that cannot resist a good pasta joke. Straight pasta jokes. I don't know what it's going to say.
I love LLMs. OK. So
what's the weather today? Is it going to make a pasta joke?
It's dying.
It's dying. OK. Allright. It looks like it might have been quite saucy today. Don't forget your umbrella. You might end up feeling like a soggy noodle. It's okay. So
that works. So here we go. So now this is running locally. And so this is a good one. When we're developing, we can just test things locally here. The next thing we're going to do once we're happy with it, we're like, this is the best app.
It makes pasta jokes. We're going to deploy it. So then we move on to the deployment instructions. So the first step of that is AZD Auth login. So this is going to log in to our Azure account that we made earlier.
So I'll do AZD Auth login. And this is going to give us a device code that we're going to paste into this OAuth browser flow. So let me go and open it up maybe over here. I think that's my Azure account that I'm using for this.
And then I go and I take this and I paste it in.
And I'm going to pick my account. I'm going to use this one. Continue.
And
OK. And then we're logged in. OK. Allright. So that was the device code flow. So you just want to make sure that you log into the account that you got with the pass. Whatever account you use for the pass, that's what you want to log into.
The next step is to create or, Gabriela, should I pause? Should we get through the local step first? Or should we keep going with AZD deploy?
Just a little bit. Like there's a couple.
Yeah. We can pause and see if everyone's got the local one running, actually. I think that might be good to do. OK. So let's just pause and see if there's any questions with getting the local one running.
Yeah.
So yeah, someone asked, can we just run this locally? You can totally run these locally as well. We like to use GitHub Codespaces in workshops, because that reduces the number of potential developer environment issues. If you want to run it locally, you can either run it just with a Python virtual environment, and you just have to install all the requirements.
Or you can run it with VS Code using the dev containers extension. And that will do the Dockerize environment for you. If you want kind of the benefit of the Dockerize environment without being in the browser and having to potentially pay for Codespaces.
So we should know also that GitHub Codespaces, you have a limit of some number of hours a month, either 60 or 120. It's 60? OK. I must have paid more. So it's 60. So you're not going to go over that today.
But eventually, you could go over that if you use Codespaces a lot. So if you're local, I think I have mine open locally as well. And yeah, locally, I'm just using a Python virtual environment. So you're also welcome to try these things out locally if you like local environments.
Just be a good person and make a Python virtual environment to manage your Python dependencies. Right?
Yeah. OK. So I saw a lot of local things. So I think we can move on to the AZD. So yeah, I did the login. So you saw me do the login here. And that's using the device code flow.
So you should see something like this happen from inside Codespaces. And the next step is to make a new AZD environment. So AZD is this tool we're using for deployment. So we make a new environment name. You can just call it like chat app, whatever you want to call it.
And then what that does is it actually makes this .azure folder. And it makes this chat app folder inside. And that's where it's going to store all of our deployment environment variables. So we need to configure anything we want to customize about our deployment.
We're going to configure that now. And it's going to update this file here.
So the next thing we're going to do is set all these AZD environment variables. So the AZD environment variables are different from the ones we just saw in the .env. The .env is just for the local server. AZD environment variables are for deployment.
Sometimes we use the same, but a lot of times we want our local environment to be slightly different from our deployed environment. So we have two different ways of setting those variables. Allright. So I first set these commands.
So this is just going to tell it to not create an Azure OpenAI, because we're using the proxy. And then we're going to set the name of the deployment to GPT-3.5 Turbo. Then we need to set the key.
So I'm going to paste this, and then I'm going to delete, delete, delete.
Gosh. That's what happens when you have Wi-Fi issues, actually, is you see it with the typing. And then I got to find my key again.
There we go. So that sets the key. And then I'm going to set the endpoint. I'm going to delete. How we're going to do this. There we go.
And get the endpoint.
Allright. So now I've set all these things. Now, if I've done it correctly, if I look at my .azure folder for that environment I created, I should see a .env that looks like this. So this is a .env that's inside the .azure folder.
So this is what is going to be used for the deployment. And it's going to tell it, this is how it's going to set up the Azure OpenAI connection. OK. And now I'm just going to do I'm just going to type, type.
Thank you. OK. Allright. And then I'm going to do AZD up. Here we go. So what AZD up is doing is that it's actually deciding it's doing several stages. OK. So I have to select an Azure subscription. In this case, I only have one subscription.
So you just press Enter. If you had two subscriptions, you would want to pick the sponsorship one. Then I select an Azure location to use. Typically, you just choose one that's close to you. So Central US is pretty good.
Now, what AZD is doing, the first step is that it's actually packaging up the code that is going to deploy later. In this case, we're deploying to Azure Container Apps. So it's packaging up a Docker container file. So it's actually literally building a Docker containerright now.
So if you do like working with Docker, Azure Container Apps is a great fit. And a lot of people like Docker. So we deploy a lot of stuff there. But we also are going to be using Azure App Service for one of the later templates.
So we've got lots of ways to deploy on Azure. So you can see it building up that Docker. The step after this is where it's actually going to create Azure resources. So it's going to create the Container Apps, create a Container Registry, create a Container Apps environment, and create a log analytics workspace.
So these are all the components of a containerized app on Azure. And it's multiple components, and we have to stitch them together. The way we stitch them together is using infrastructure as code. Has anyone used Terraform here before?
OK. So we have our own version of Terraform. It's called Bicep. And it is infrastructure as code, which means we're declaring what resources we want to make. So we say, oh, we want to make log analytics. We want to make Container Apps.
We want to make the actual Container Apps image. And then we're going to assign some roles. So all of that is declared in this Bicep file. So that way, you have repeatable processes for provisioning. And this is really helpful when you're making complex applications on Azure, because you might have like 10 different things you're using.
You might have a Postgres and a Key Vault and a Redis cache and log analytics and App Service. And you want them all to tie together. So you can declare what that infrastructure looks like and then put that in a Bicep file and then deploy it.
You can also use Terraform. So if you're really into Terraform and very comfortable with it, you could totally use Terraform here as well. I don't know Terraform. I haven't used it personally. So all of my examples do use Bicep.
But if you want to send a PR with Terraform, I'll review it and just stamp it, because I don't know how to reason about it. So what you can see here is that it is actually creating the resourcesright now.
So you can watch it here. You can also watch it in the portal. It's not really super exciting to watch. So this is the point where I usually fold my laundry, because it can take some amount of time.
Or you can even get an error. Oh, I used OK. So I already made one in Central US for the earlier demo. So I should have picked a different region. So for this Azure pass, there is a constraint of one Container App per region, which is why we said in the README that you should pick a region that you haven't picked before.
And I didn't pay attention to my README. So that one won't deploy. So what I can do is I'm just going to make a I'll just make a new environment. I'll just copy everything over.
Now, you shouldn't run into this, because this would be your first environment. So chat app 2. And I'll just copy and paste. We'll change it to chat app 2. And then West US seems like a good region. OK.
And then I'll AZDM select chat app 2.
There we go. And then AZD up. Yeah. OK. So then it'll do the up again. But I have one of these already deployed. So I'll just open up the deployed ones so you can see deployed. Deployed is going to look pretty darn similar to what it looks like locally.
Where's the one deployed? OK. So this one's deployed. It looks pretty much the same as what it looks like running locally. The difference is that the URL is a Container Apps URL. And you'll see this URL displayed in the terminal once it finishes successfully deploying.
You'll see this displayed. Let me see if I have that in my history anywhere from earlier today.
No. You never know. OK. So let's see how it's going here. Yeah.
Was it something that was generated or something that you made yourself?
I lovingly handcrafted them. Yeah. So yeah, I write the Bicep files. Some of them, all the ones in core, are actually from a shared repo that we just copy and paste from. We're trying to move towards something called AVM, Azure Verified Modules, which are Bicep files that are maintained and have security best practices in them.
So we'll gradually be moving over. But basically, with Bicep files, you can use ones from a central registry. You can use ones from your own private registry if you're doing a lot of them. Or you can just use ones inside the folder.
So there's a lot of techniques you can use, depending on how much Bicep you're using.
OK. So now it's starting over and deploying again. Allright. So let's walk around. Or any questions on what I showed here? Allright. So I saw a lot of AZD deployments are going. I saw some issues with naming, which I run into all the time.
Azure has very obscure naming rules. The safest thing is to do short names with no symbols in them and nothing fancy. If you do have a naming rule, you can just always do AZDM new and make a new environment and start over.
And that should be OK. But generally, the issues you run into with deployment are usually related to naming, region constraints, account constraints. And that's probably, yeah, the ones you might run into. Allright. So we're giving it 45 minutes left.
RAG Explained41:08
I'm going to show you the other ones. And these are ones that you can also start trying to deploy now, following very similar READMEs. So the first one actually, these two are both about RAG. So first, I'll talk briefly about RAG.
So let me first motivate it. So let's see. Tell me what Pamela Fox likes to code on. I don't know. Let's try this. I'm trying to get it to lie.
This is the pasta one. OK. So this one clearly lied. Spaghetti Python. Great. Allright. But then if I go to
this oneright here, tell me what Pamela Fox likes to code on. This one will hopefully be more accurate, at least have less pasta jokes here. So basically, what we're trying to show is that if we just ask an LLM to answer a question, it's very possible that it's just going to make something up, if that's what it seems like there.
Oh, good. I mean, in this case, it says it doesn't know what I like to code in. I think I should have said code in. Like here, what Python frameworks does Pamela use? Let's try this one. So if it doesn't know the answer, it'll say, yeah, in this case, it does know the answer, because this is actually using the RAG technique in order to answer questions based off a knowledge source.
So those are our last two samples are about RAG. So the general approach of RAG is that we get a user question. We use that user question to search some sort of database or search engine. We get back matching search results for that user question.
And then we send those to the large language model and say, here's the user question. Here are the sources. Now answer the question according to the sources. And so now we can make customized applications that can actually synthesize and answer questions for any domain.
So we've got two RAG samples here. So one of them is RAG on Postgres. So this is for the use case if you've got an existing database and you want to be able to ask questions about that database and have the LLM answer accurately based on that.
So for the example database that I'm using, I have product. So this is a chat on products. So there are tables storing all the products for this website. So I can say, OK, what is the best shoe for hiking?
So then it's going to go and search the database rows and get back matching rows and then come back and say, OK, this is blah, blah, blah, blah, blah, blah, blah, blah, and it's going to include citations. So one of the key points of RAG is to have citations so that users can verify where the information comes from and see that it's actually legit information.
And we can also look at the process for this RAG flow here when we look on the thought process here. And as we'll actually see, is that this RAG flow is a multi-step process. So the first process is actually what we call the query rewriting phrase or the query cleanup phrase.
So that's where we take the user's question and we ask the LLM, like, hey, here's a user question. Turn this into a good search query. Because a user question may not be that well formulated. Please tell me about the best shoes for hiking.
No. OK. So there's a user query. And that's probably not the optimal search query for a search. So if we look now at the thought process, we can see that the LLM actually turned that whole long thing into best shoes for hiking.
So that's a better query. So that's our query rewrite phase. So that's an LLM call. Then we get back the resulting rows from the database. And then this is our call to the model that says, hey, you need to answer questions according to the sources.
Here's how you should cite your sources. Here's the user question. And then here is all the sources. So this is basically RAG. And then we're able to use it with different sorts of data sources. So that's RAG on Postgres.
So you can get that set up following really similar steps to the other one. And you can even run that one locally first as well, just on a local Postgres database. So here, you can run the app locally.
This one is a little more fancy, because you've got a React front end there. Then you can deploy to Azure. You're going to set similar variables and run it up. So if you're interested in that, you can start going through those steps.
And then you can customize it. The other kind of RAG that we have is RAG on documents. So if you're trying to ask questions about unstructured documents, like you've got a bunch of PDFs or Word docs, Excel files, anything like that, you can actually put those into a search index and then search that.
So the example we have for that is RAG with Azure AI Search. And it's a really, really full-featured example. We've had it for the last more than a year now. And we've had thousands of developers deploy with it and put it into production.
And so it's been used for a ton of use cases. And it's got a lot of features: speech, voice, vision, user access control, lots of cool things in it. So let me show that was the one I was actually showing earlier with my blog.
So I made a version of it that's just based off my blog posts. And it can cite my blog posts. I've also got this one here, which is for an internal company handbook, which is a very popular way of using it as well.
And so you can see for each of them, we can click on the citations. And yeah, so now this is a bit more complicated, because here we have a multi-page document. So we've got a 31-page PDF. We can't just send an entire 31-page PDF to the LLM, because for a lot of our LLMs, it's going to go beyond the context window.
A lot of our LLMs have a context window limit. So typically, that's around 8K, 8,000 tokens. It can go up to 32K, even 128K, we're seeing. But typically, they do have some sort of context window. And even if they don't have some sort of context window, LLMs can get lost if you give them too much information.
There's a research paper called "Lost in the Middle," where they did a study to see if they throw too much information at an LLM, at what point it stops paying attention. So we generally want to send the LLM the most relevant chunks.
So what we do is that we first have this data ingestion phase that will take a PDF or whatever document, takes the document, it extracts all the text from it. And we do that with Azure Document Intelligence, which is very good at extracting text from all sorts of documents.
So we extract the text from it. We chunk up the text into good-sized chunks, usually around 500 tokens each. Then we store each of those chunks in the search index along with their embeddings. And that's what we actually search on and send.
And then we send. So if we look at the search results here, we can actually see that the search results are just chunks from the PDF, where we say, here's the chunk. Here's the embedding. This is the page it came from.
And this is the file it came from. And we just send back those chunks. So this is the most complicated of our architectures, because we do have to have that data ingestion phase. And that means we have to have a script or a process that does that ingestion stage.
And here, we can do it locally or in the cloud.
So those are the two RAG samples. So we have another 40 minutes. And we have a good ratio here of how brief to y'all. So if either of those sound compelling to you, sound like a use case that you're interested in, then you can try to deploy them now and see how they work.
So once again, you just go to the app templates workshop repo. And you can either pick RAG on Postgres or RAG with AI search and then start going through the steps to try it out. These will take longer to deploy.
So it's good to start the deploy now, because they've got a lot more infrastructure to set up. And then for the AI search, it's got to do the whole ingestion step. And that ingestion step takes a certain amount of time as well.
So yeah, any questions before?
For the ingestion step, are you using any libraries for the chunking and all that stuff?
Yeah, that's a great question. Are we using libraries? So when this sample was first created, it was like last April. It was before there was really good established libraries. We kind of used LangChain, but not heavily. So all of ours, it's actually custom coded.
Now, if you're going to use a library, the big thing I would make sure you're doing is using a token-based chunker. A lot of the splitters out there are doing character-based splitting, which is probably fine if you're doing English-only documents.
But we do have lots of international customers. And as soon as you start doing non-English documents, then you really want to do stuff based off of tokens and not characters. Because imagine you take a Chinese document and you say, oh, my chunks are 1,000 characters long.
That's a lot of tokens. You can go over the context window really fast. So we have token-based chunking that we've implemented here. There is token-based chunking available in LangChain. So if you're going to use LangChain, the thing to do is find my colleague's blog post where he talked about it.
OK, yeah, working with CJK, especially if you're doing anything non-English. He basically analyzed all the splitters from LangChain to figure out which of them properly worked with token-based splitting and with CJK languages in particular. So we've implemented this ourselves.
My manager, Anthony, he worked on it. But LangChain and LlamaIndex both do a lot of this stuff. They take care behind the scenes. So what you need is you need the splitting. And you can get that basically from LangChain, because LlamaIndex uses LangChain.
So I would just say use LangChain probably with this one, so you can specify the chunk size. And then you just have to vectorize. So that's easy. You just use the OpenAI SDK. And we do the batch embeddings with that so that we can do a bunch at a time.
And then you just store it in AI search. So the hard part is really extracting the text. So there, we either use Azure Document Intelligence in the cloud, or we do have some local parsers too. If somebody doesn't want to use Document Intelligence, we use PyPDF.
We use our own CSV parser, because that's straightforward. HTML for my blog, I just use Beautiful Soup, which is the Python package that does HTML parsingright, because I thought I could do a better job at it. So this one, I just use Beautiful Soup to extract the text.
So you can do that as well. And we've got Beautiful Soup in there. So yeah, there is actually a surprising amount of things that we've written ourselves for the AI search repo. If we were going to do it today, we'd probably use the LangChain splitter at least.
Yeah, good question. Sorry, long answer. Other questions?
So the
ABM, was it or wasn't it?
ABM?
ABM. I noticed that it has five new services.
If I set it up once with a certain deploy and I decide to change later the ABM, the byset to something else, is it smart to reconcile the differences?
So generally with byset, what it does is that it tries to figure out and byset is really compiled down to ARM. And ARM is just JSON. So what you're actually doing is what's called an ARM-based deployment. So with ARM-based deployments, what they try to do is figure out what does your resource currently look like, what are you saying you want it to look like, and what changes does it need to make happen.
So yeah, we'll probably switch over to ABM in a lot of our samples. And we're probably just going to make sure we're trying to make it not have a change. But if you want it to change, then that's fine.
So you should totally be able to switch between ABM, not ABM, as you decide, as you see fit. And the important thing is just it'll figure out the difference and just make sure you are on board with any changes that come up.
So if you're doing there's this AZ deployment command that does what if. And that actually tells you what resources will change. I want to figure out how we can do that with AZD. I think AZD maybe has a dry run command.
So that might be what we try when we consider switching to ABM, because we want to switch to ABM so that we don't have to maintain our own modules. But we just want to make sure that we are aware of any configuration changes that could happen.
I'm taking note that ABM is just basically sitting on top of AZD, or they are completely two different things?
They're just different things. Yeah, AZD is a command line tool that
does the ARM-based deployment and also does code deployment, code upload. So I have this azure.yaml here. I didn't show this. So azure.yaml says, this is the code that you're going to deploy to this host. So AZD does multiple things.
It does provisioning, which is basically doing an ARM-based deployment, which is equivalent to if you're doing AZ
deployment, if you know the Azure CLI, it's this AZ deployment command. So it's doing that. And then it's also doing packaging and code deployment. So if you've ever done I don't know if you've ever done web app up, that's where you deploy code up to app service.
AZD will also do that for us. So AZD is trying to do the whole workflow of you need to provision your resources and you need to deploy your code. And we're trying to make this central way of doing it across all of our offerings.
Becauseright now with Azure, if you know Azure, we've got a billion different ways of doing things across all the different things. And AZD is trying to make a more common way of doing it. So if you look at my GitHub repo, I'm kind of a huge AZD fangirl.
So you can see all of these repos are all AZDified, almost all. That's what this AZD column is, because to me, it's the best way to deploy, because it's repeatable. So if you are looking for examples, I have quite a few here.
But yeah, so we should be able to do it on different hosts, container apps, functions, app service, Kubernetes, et cetera, and with all the different possible byset.
Other questions?
Production & Eval56:30
So what happens after we go to production? Observability and all that good stuff.
Oh, yeah, great question. So we do have a generally, there's lots of docs under Azure Search OpenAI Demo. So we do actually have a productionizing guide. You also asked specifically about observability. We do integrate with Application Insights with OpenTelemetry.
So that's what we use by default. If you want, you could use LangFuse. I actually like to use LangFuse. I don't know if you've seen it, but it's an observability platform. So you could use LangFuse. But by default, we're using Azure Application Insights with the OpenTelemetry packages to bring everything in there.
But we do have a whole productionizing guide that talks about how are you going to scale things, if you need to load balance your OpenAI capacity, if you want to do VNet deployment, if you want to do user auth, how to do load testing.
So I've run quite a few load tests for this one. And then how to do evaluation. So I like to do evaluation. Everyone should do it. It's basically like the new form of
testing for this world. Let me see. I think I closed my evaluation repo. I can open it. But basically, you want to be running evaluations to see if you are getting quality results from your LLM, because a lot of times you might run here's the thing.
I show those sample questions all the time. And they perform great. But you cannot trust your sample questions. And you can't even trust them. You might make a prompt tweak and be like, oh, this prompt tweak was so good.
I'm getting such good results. You cannot trust it. You have to run an evaluation across a huge number of samples to make sure that it's actually running. I run it across like 200 samples. It's probably the minimum of what you should do.
But you have to run evaluations in order to see I'm assuming you run evaluations for Copilot chat,right? How many samples do you run off?
It's always generated off public repos.
Oh.
So it's never growing.
That's cool. You want to come up and talk about evaluation? Because you're like Harald is running an actual because basically, you're making Copilot chat. Here we go, Copilot chat, which you can see I use it a lot. So you're like running a real RAG.
Yeah.
Yeah.
So it's different reqs. So there's not just one RAG. So if you do add workspace, if you ever try that in Copilot chat, we actually run a local sparse index. So that's basically your classic how Google works, just looking up words on the internet and documents.
And it's faster. And we can do it locally. So that will always work. We also do a
semantic index against an index that GitHub.com maintains and ranking those in using another LLM call. So RAG basically becomes a series of indexes. You do Postgres. You showed that you created some keywords up front based on the search.
So that's what we also do locally. But yeah, anytime we have changes, we have one test set that can run on each PR, and then a larger test set that we run daily that has a lot more repositories from across different languages.
So you run the evals in the PR?
Some, yeah. So we have a subset that's more unit test driven, where it's like, can it answer questions for this? Does it hit any issues we've seen in the past? So it's more unit test style, where it's like, does it behave as it did before?
So yeah, it's important to know. That's the first biggest thing we invested on early on, because we found it's so easy to get lost in prompt crafting and assume how RAG works and assume how it works in the wild.
And the local sparse index, is it SQL-like?
It's TFIDF.
Oh, TFIDF.
Yeah.
I don't.
Yeah.
That's cool.
So yeah.
OK, so let me show I don't know if you can show your evals. But here, I can show evals on the Azure AI search. So let's see.
I'll summary. OK. Allright, so we'll look at them for my blog. Let's see. So these are ones I've run before.
Oh, probably Pamela's blog. Pamela's blog? Yeah, Pamela's blog results.
OK, allright, so these are a bunch of evaluations that I've run fairly recently. So with these evaluations, I do GPT metrics. And then I also do basically regular expression metrics. Are your metrics usually GPT metrics or?
No.
No.
It has code tests.
Code tests. OK, OK. So with these GPT metrics, what they're actually doing is sending the original answer, sending the ground truth answer, which is generated synthetically, and then also sending the new answer to a LLM and saying, hey, rate this from 1 to 5.
And then we can see the results. And this is the actual prompt that gets sent. It's like, OK, rate this from 1 to 5. Here's some examples. So we do that for groundedness. We do that for relevance. And then I also check whether citations match across ground truth and not ground truth.
And that's actually my favorite metric. And that's just a regex. So my favorite is just this one, this citation match here. So I'm just making sure that the answer contains at least the citations that were in the ground truth.
So I run metrics like this. A lot of times, I'm looking at retrieval parameters, because for RAG, the retrieval makes a big difference. So here, I was comparing stuff like, what if I use text only? What if I do vector only?
What if I do hybrid? What if I do hybrid with Ranker? And that's super interesting. I was trying with different retrieval amounts, like if I retrieve 5 versus 10 versus 3, what do I get out? What if I change a prompt?
I just say, I've tried so many tweaks on our prompt. And I've never managed to actually get improvement in the overall stats. So we still haven't ever changed the prompt, because I haven't really proved that anything is sufficiently better.
Or I'm just a really bad prompt engineer. I don't know. None of my prompt engineering ever moves a needle. For me, the only thing that moves a needle is retrieval parameters, like how you're working with your search engine or changing the model entirely.
Changing to GPT-4 has a big difference than GPT-3.5. So that should be part of your putting in production for sure, is to make sure that you've got some sort of valuation set up.
Other questions?
Is anyone trying to get one of the RAGs up? Anyone getting RAGs up?
I tried to take a lot of things to production, but they were all throwaway toy projects with a fast API. And I had a hard time evaluating which vector store to use and how much to chunk and whether my results are good or not.
So they were all stuff that's where I felt like I didn't make the most progress. I was like, I know I should be testing this.
Yeah, so yeah, I mean, it's hard. That's part of why I do it as well. I've also run all those on our sample data too. But I think what I've discovered is it really helps to run the evaluations on stuff that you know, because then this is the summary.
You can kind of look at the summary and be like, OK, I guess things went better. But then what I usually look at is I actually look at the changes between two runs and be like, OK, well, what was the difference between the baseline and then
the maybe what was it, vector only? Vector? No, Ranker? OK. And then I'll just look at things that changed on citation match. OK, so this is what I usually do, is I look at the overall stuff. And then I look and I compare the answers across my ground truth and the new one with the parameters.
And so then I can better reason about it. But you really have to know your domain in order to be able to evaluate your evaluations.
But it also helps if other people have run it for you. So this is a really good blog post from the AI search team that I always reference, where they ran massive queries looking at hybrid search versus vector search versus text search.
And they found that hybrid retrieval with semantic ranking outperforms vector-only search. So I ran my own versions of that and recently blogged about it. But it's basically the stats that I was just showing, where what I found actually for my use case, vector on its own did horribly, like really, really, really badly.
Where is it? So vector only got a groundedness of 2.79, which is really, really low. Text only got 4.87. So part of that is because Azure AI search is really good at full-text search, like incredibly good at it.
It does all spell checks, stemming, everything you could imagine. Hybrid, which is where you take vector and text and then you merge them using this algorithm called Reciprocal Rank Fusion, which you can actually just see the algorithm is just this.
It's just you're just doing a little math here to combine rank scores. So just a basic hybrid like that, the groundedness is only 3.26. So you can see hybrid on its own is worse than text only. And that's because vector results can add so much noise.
You accidentally grab the wrong, distracting things. What I found actually is if I ever accidentally vectorize an empty string or something close to an empty string, it's similar to everything. I don't know what this is about the OpenAI embedding space.
But if you accidentally vectorize an empty string or even we have vision as a feature in the Azure OpenAI search demo. I was helping a customer this week. And they were finding that so many of the results were getting this blank blue page, because apparently, this blank blue page, the vector for it and this is a vector via a different model, the Azure Computer Vision model.
The vector for it was just matching everything. So you've got to be really careful with vector spaces. It's so easy to accidentally add noise to them and for there to be distractions. So hybrid on its own only got like 3.26.
Once I used hybrid with semantic Ranker, then I got the best results, but only by a couple percentage points. Now, hybrid with semantic Ranker, semantic Ranker, that's a feature of Azure AI search, which is actually another machine learning model.
It's called a cross-encoder model. But basically, they actually had humans rank results according to queries. They use it for Bing. So they said, hey, humans, here's 10 search results for a query. Rank these from 1 to 10 and tell us what's the best.
So they train a whole model based off a bunch of human data. And then they get back this model that they can then use for any arbitrary ranking of user query along with results. So basically, hybrid with Bing Ranker gets you the best.
But if I was going to have to if I was on a desert island and I could pick between vector and text, I would use text, at least for Azure AI search. It's going to depend on how good your full-text search,right?
If you're doing full-text search with SQLite, which I don't even know if supports full-text search, it's not going to do very well.
So what's actually powering that full-text search? Is it TFIDF?
So you're using TFIDF for your Copilot chat, you said,right?
For the.
For this one, for the at workspace,right? Yeah, for Azure AI search, they're using several things. But one of the things they use is Lucene, which is
a search library. And it's got stuff like spell checking and tokenization and stuff like that. So they're doing a lot. And they're also using BM25, which I think is basically TFIDF,right? OK, yeah, so BM25, that's what you want to look for
is oh, we got a search result here. Yeah, so if something is using BM25, I think that's basically the best for full-textright now. So that's what you want to look for, is you just want to look for a good full-text option.
Yeah. Yeah, it's overwhelming. That's why I love when people put out research. So we'd be like, OK, great, because this also has the optimal chunk size. That's why I was saying we do 500 tokens, because they did the work here and said, OK, the optimal is 512 tokens.
Great, that's what we're going to use. Now, obviously, for your particular use case, it can be different. But we can't all run like 2,000 different tests to see what the optimal thing is. So it's really nice when people document what worked well for them.
Cool.
Anything else?
How often do you update the vectors? If you have a lot of
data that gets updated a lot, how often do you choose to update the vectors?
Well, we only need to update the vectors if the data changes or if we're changing our embedding model. So if we change our embedding model, we have to update everything to use a new embedding model,right? Because now OpenAI has these new embedding models.
I need to do some tests with them to see if I can get better results for them. So in that case, I would rerun everything. So probably what I want to do is set up a separate AI search index for this one, which uses one of the new embedding models, embedding 3.
And I have to decide how many dimensions to use and then compare it to see how much better results are. I'm told that generally, the results are better. But have you tried any of them?
I'm switching it.
Oh, you're switching to the new one? What dimension are you going to use?
Small, probably.
So you're going to use small, 256? Wow.
512?
Yeah, you can do 512 too. Yeah, yeah, so you can that's the thing, is it's so many options now.
You're going to test.
Yeah, you're going to test.
Run A/B tests.
Yeah, oh, and you can run A/B tests. You have customers, yeah. So yeah, but you're going to have to re-index everything.
Yeah.
So that's when you would have to update stuff, is if the content changes or if the model changes.
And then test that. Yeah, I do want to try out the new ones. They should redo this one too.
There's too many decisions.
Cool. Any other questions? Harold, do you want to show stuff in Workspace?
Again?
We'll close out on this.
So see RAG in action.
Copilot Chat1:11:43
So if you ask a question in Copilot Chat, so that's the Copilot Chat panel version. There's also another one that's inline. So if you open up, this is a natural input. We call it inline chat in your code, basically letting you apply code directly or natural language directly to your code, which is always nice.
You don't have to think about the response. You have to think about just you know what you want. And you want an AI to do it for you. But in the side panel, most of the time, what you will run into is you're going to run things.
Let's pick a function. These are tests.
So compare the tests. Now I've code selected on theright. And on the left, I can ask things about the code I have selected. That's the surefire way to get good results, have code selected, and talk about it. And you already see that we do some magic in our responses.
So everything is code highlighted. So you can actually jump to the different aspects that are being used and even to dependencies. So it found that there's a dependency. So you can also jump to that. So now going back to here, let's see which tests actually are defined in a repository.
And because I want to talk basically about the whole workspace, and that's where I can just say which
tests are defined or how are benchmarks being run. So a general question that you would go otherwise to a colleague who hopefully knows this. And hopefully, they're in the same time zone and they know this. But now I can actually send this to Add Workspace.
And that's where we're kicking in this whole RAG agents scheme. So this repo is probably not indexed on the GitHub.com site. So if you're in Copilot for Enterprise, you will get a semantic index that GitHub keeps updating for you.
They also have a few open source repos indexed. But in this case, this is all happening now in VS Code itself. So this is mostly sparse indexing. And actually, we see that sparse indexing is usually on par, similar to what you see with the text-based retrieval, that this works really well.
So this did a TFIDF of the whole repo?
Yeah, yeah.
So just basically figuring out all the most repo forward.
Yeah, so first we do same as you have in Azure Search, where it does find more words for what you're potentially looking for that are fitting with the repository. So we also do stemming. But that's one. That's the first LLM call.
Then the TFIDF will find all those results. And then we do the re-ranking on top. And that actually gets us actually mostly better than doing a full vector search on the same topic.
How do you do the re-ranking?
Different ways. So that's where we experiment a lot. But it's another GPT 3.5 call, I think.
Oh, OK, so use the LLM as a re-ranker.
Yeah.
And you see what's being pulled in. So these are all the things it found and the chunks it found within. So what we do, what you'll see is we actually do semantic chunking. So for most languages, we look at function segments.
We look at specific blocks of code. And that's where we found the most impact as well. So people brought up chunking as a big area of improvements. And that's what we also have in our code, that the chunking is the biggest impact, I think, from what we've seen.
Semantic chunkers?
Yeah,
which helps that we have all the languages, the knowledge around the team, like Python,right? But yeah, so that's the basics. And you'll see that it works everywhere. That's the nice part. It works locally. And it works slightly faster if you already have an online index where we can retrieve the semantic index from
in action.
Questions?
What is the .prompty?
Oh, so .prompty is a new prompt format. You can show the .prompty file.
Yeah, all those .ones, those ones, yeah. So this was announced that was called build.
Build, yeah.
Yeah, so now it's a build. Scroll up to the top of it, yeah.
There you go.
So it's a way of it's like an artifact of prompts. Becauseright now, you might store your prompt as a multi-line string variable. We store them in all kinds of formats across the repo. So this is like a standard way.
So it's actually a Jinja template plus this YAML at the top. So the YAML describes the metadata of the prompt. And then the Jinja template, it's a template that you can pass things into. So this is used by PromptFlow.
But it's also used by Azure AI Studio. And the goal is and I think maybe LangChain might have support for it now or soon. But the goal is just to have a common way of representing prompts. So we'll probably try to use this in more of our stuff going forward.
There you go. Just ask Copilot.
Yeah, so I'm using this is using the PromptFlow evals package, which has a bunch more things as well, other kinds of evaluation. I wrote my own CLI and UI on top of this. But they have one too that you can use.
Yeah.
Prompty.
Do you run the evals in your CI pipeline somewhere? Yeah. If you look at Azure Devs, this one, it does actually run them. So I run themright now, I'm just running them as a smoke test for this repo.
But you can see what I've done is that I have a target URL. So that's generally what you'd want to do is you'd need to run the eval against your live or for you, you're doing PR builds. So there you want to run it against your PR build.
So the tricky thing is just making sure you have a way of contacting your thing with everything, all the production setup, all the Azure stuff in it. So yeah, I would ideally have it as a CI step for every one of our repos.
And I'm just figuring out theright way of setting up the target URL and all that stuff, especially because most people aren't making public-facing apps. Most people are either putting it behind user auth or putting it in a beta.
So we need evaluation flows that both can use your production resources, because that's how you know it's working, but then also work with however your app is deployed. So I think you can certainly figure out how to set it up for your situation.
I'm still figuring out how to set it up in the general case. But the thing to keep in mind is that evaluations are slow if you're doing GPT metrics,right? Or I mean, generally, they're slow because all of these calls are slow.
You saw how much time it took to get back a response,right? So generally, they're slow. They're much slower than traditional unit tests. So you do not want to casually run an evaluation. They're also expensive, especially if you're doing well, first, because the LLM calls happen behind the scenes.
And if you're using GPT metrics, because I'm doing all these GPT metrics like relevance and groundedness, that's another LLM call. So you want to have a higher barrier to running than with normal unit tests,right?
And caching.
Caching, oh, you cache? How do you cache? How do you know that something hasn't changed?
Based on a prompt and the test, yeah.
Yeah, I guess, yeah, if it's all within one repo. This one is like a repo that works with other repos. You don't know if the app is changing behind the scenes. But yeah, yeah, so caching, if you need to cache, that's good.
Cache what exactly?
So we look at each test. And we only rerun them when any of the prompts, when the input, basically change. So if you imagine like an OpenAI proxy that you could set up, if it's the same, similar to what they do, I think OpenAI has like the seed variable, which is basically caching.
But they don't tell you. And that's basically, if nothing changes in a prompt, it just sends back the old response.
Oh, so you implement caching in Copilot Chat, you mean?
Not in Copilot, in our testing infrastructure.
In some people also implement caching in the RAG application itself or maybe in the embedding step or something. I still don't know how often you're going to get the same question.
For tests, it helps. For tests, always send it, yeah.
Yeah, yeah.
Yeah, I haven't played too much with that.
It's good. So can this also, is this just for OpenAI? Can this work with Mistral and all the other ones I mentioned in the beginning?
Good.
Yeah, I mean, mine, just this one, I just hit up a URL and get back the answer. So the URL is just of your deployed app. Oh. Sorry, I meant like the starter templates. Oh, yeah, good question. So with the starter templates,right now, they're all configured with OpenAI.
And so you can swap out different OpenAI models, so you have GPT-4. But they don't work with the new non-OpenAI models because we can't necessarily use the OpenAI SDK with them. I think there is actually a way to use OpenAI SDK with them.
But we're supposed to pretend we can't. So there is this new SDK. And I haven't messed with it yet. I don't know if you have. But Azure AI inference, have you seen it? I think this is the new unified SDK.
And yeah, so this is what to use for everything that's not OpenAI. Oh, it says it can even do OpenAI. So we might have to port to this. The thing I don't love about this is that this is Azure specific.
Becauseright now, we use the OpenAI SDK, which is like not Azure specific exactly. And so it works with like Ollama and stuff. I don't know. But we might end up porting for this. So if we ported to this, then probably it would just work with everything.
So this is really new. This came out at build. So we just have to decide whether to port everything out to this so that we can use all the models.
Yeah, everything changes all the time.
Yeah, but we would also need to make the Bicep for it. The other thing I haven't done is I haven't because I try to set up Bicep for everything. So typically, Bicep creates your Azure OpenAI instance. If you're using Mistral or Llama, you would want Bicep to you would probably want Bicep to create that as well.
And so that would be different Bicep in addition. So we'll probably end up adding it. Because basically, what you do is you go to the issue tracker. You file a request. And then if enough people ask for it, we're like, OK, guess we're going to do it then.
We can.
Yeah, yeah. But that's how we figure out what it is that people are looking for. Because it is really nice to be able to swap out models. Becauseright now, all of the samples do work with Ollama. So if you have Ollama running locally, here's my little Ollama up there, you can run like Phi 3 and stuff like that.
You just go to your terminal. And you're like, Ollama, is this it? I think I don't know if I typed Phi 3 correct.
And so they do all run with Ollama things. But none of the Ollama models have really been sufficient for RAG, in my experience. I run them just to check. But they all fail to follow directions, in my experience.
Because I just think they don't have enough parameters. These are like 3B, 7B, et cetera. So they don't provide citations correctly. I don't know. Have you had more success with like Phi 3 mini?
We see every model requires prompt changes.
Yeah, and I'm bad at prompt engineering. Yeah, so out of the gate, I haven't had success using any of the small language models for RAG. I'm sure the big versions of them would work much better. So I do want to try out like the 70B.
70B, yeah.
I've done up to 7B. Because that's like how far I can go up locally. I can't go much more than that just for space reasons. So for them, what happens is that they'll answer the questions fine. The issue is that we need citations to be in good format.
Because these actually come back as bracketed square brackets. And they just don't reliably come back with square bracketed citations, which doesn't sound like a big deal. But we're trying to make clickable citations here. So that's the issue I've had, is that I think they're fine at synthesizing the information.
But they don't follow the syntax directions in terms of the citations. And they're kind of maybe more likely to make stuff up if I ask an off-topic question. That's been my experience there.
Good for reranking. I think finding theright spot like reranking or the GPT judging. I think that that's where finding this one thing, maybe not like the expert full answers that follows to format, but one of the smaller tasks.
But you can't so most of them don't support function calling off the bat. So would you do reranking with just a simple you'd have to figure out what syntax they come back with.
But they're all good at coding.
Yeah, that'sright. That's true. Everyone's good at coding. If you can turn something into a coding task, you're good.
Yeah, that's another form of RAG. I was telling someone last time, these are all doing kind of RAG on just a few documents at a time. If you're trying to analyze a whole database or a huge number of documents, then you really want to actually use a SQL query with aggregate functions or do a Pandas query,right?
So at PyCon, we did a demo where you upload a CSV. And then you say like, oh, I want to count the top restaurants in it. And then it just comes up with the Pandas code. And then it runs the Pandas code in a sandbox environment.
So that's another increasingly common form of RAG, where if you want to come up with insights and analysis and that sort of thing, then you want to consider a different architecture, where you're actually going to have the LLM generate Pandas code or SQL code.
It's very good at both of those. And then run those in a safe way.
And what I've done for the syntax is I provided them a text string definition. And then the text string definition is I don't try to do the syntax. I just have them give me the explicitness of what I'm going to report.
And then now, yeah, TypeChat.
Yeah, are you actually using TypeChat? Because that's basically what TypeChat does.
I typically read pretty well.
OK, OK. So yeah, but what you're describing, that's the same approach.
Yeah, I've seen it. It's a lot more better execution.
Yeah, so that would be an yeah, we're trying I know Daniel's actually experimenting with Daniel's the creator of TypeChat. He is experimenting with TypeChat with the local models.
Because we did also try TypeChat with Phi 3 locally to see if we could use it instead of function calling with OpenAI. And we were having a hard time with it. But I think Daniel maybe has to tweak the prompts.
And maybe they'll end up working better.
Do you know if Phi 3 has a response format JSON?
Maybe the bigger one, but not the smaller one.
Not the smaller one.
Yeah, because that was the winning part for me. Because I eventually had to just find trying to get new text, just give me the JSON. And I got stuck on a few.
So you just told a DD JSON.
Not for Phi 3, but like GPT-3.
Yeah, yeah, and that would work. Yeah, yeah, that makes sense. Yeah, yeah, with prompt engineering like that, then.
It'll work.
Closing1:28:16
Nice, cool. Allright, everyone, well, you have this passes for seven days. So feel free to keep deploying. If you have any feedback for the workshop, tell us. Or we have a survey there, which I assume is anonymous.
It is.
Yeah, it's anonymous. And so you can fill out that one as well. You can take a picture, fill it out later. And that's it for today.
Thank you.





