Welcome0:00
Thank you so much for coming. My name is Gabriela de Queiroz, uh, I work at Microsoft, as you can tell. I'm Director of AI, working on— with startups and doing events like this, outreach, talking to founders. And then I have Ash, my colleague.
Hi. I'm a Senior AI Advisor with Microsoft for Startups, so pretty much spending my day-to-day talking to startups, helping them build their AI tech stack, primarily on Azure, but, like, no bounds.
Uh, I'm Pamela. I'm a Python Cloud Advocate.
Just click on this thing too.
And, uh, I spend most of my time working on open source repositories, like the ones we'll be using today, and then also doing live streams and conferences and all of that sort of fun stuff, like here.
Awesome. So, uh, we'll have a lot of, like, hands-on. But before that, we're going to set the stage, share a few things, then we have the instructions so everybody will follow. So today we're going to talk about the AI Templates, so it's a way for you to run your AI application in minutes.
Uh, and the agenda is more or less like this: we're going to talk about Microsoft for Startups, something that we call Founders Hub, and then the partnerships, the, um, the— then the AI Templates, and then we go to the hands-on workshop.
How many of you have heard of Founders Hub?
Founders Hub1:36
One, two. Okay.
Okay. Uh, well, uh, we have something called Founders Hub where we offer you a bunch of things, uh, not only credits but a bunch of other things that I'm going to talk about. And the sign-up process is very easy, it takes you, like, less than five minutes.
Uh, it doesn't matter where you are in your stage, if you only have an idea, or if you already have a startup and you are incorporated. It doesn't matter if you have funding or not funding at all. Um, and there are, like, several cool things about this platform, the Founders Hub platform or product.
Uh, you have the benefits piece, but you also have one of my favorite pieces, which is, like, the guidance. So you have one-on-one, um, calls with experts to ask technical questions, to, like, anything like go-to-market, or, like, how do I go about my strategy, anything that you can think of that you can leverage.
Uh, and then there is something new, and Ash, please chime in, uh, if I'm missing something. But there is something new that we added to the platform, which is called Build with AI, and that's where we are going to be focusing.
Um, it is an open source piece, it's all on GitHub, but you can access through the Founders Hub, which again is the Microsoft for Startup platform that everybody can sign up and, and, and join. Uh, one of the cool things is, like, you get a lot of credits,right?
So, like, everybody likes free credits, especially when you are trying things out. Um, and we offer up to, up to $150,000 in Azure credits that you can use across Azure services, including, which not a lot of people know, Azure OpenAI.
So we have the same OpenAI APIs on Azure with the whole security compliance capabilities of Azure, plus a lot of, like, other benefits like GitHub Enterprise products, uh, Microsoft 365, LinkedIn Premium, and more. Uh, and then you can use, uh, Azure, for example, Azure AI Studio, you can use models from OpenAI, as I mentioned, but Llama and others.
Uh, so one of the main slides, if you are looking for credits, guidance, and so on, this is, like, the place where you take a picture and you get the URL and you apply in minutes.
Um,
do you want to talk about, a little bit about more about the cloud credits?
Uh, yeah, I think one of the most important things that I would.
Sorry.
Okay. One of the most important things that I personally like about the entire Microsoft for Startups program, it's not quantitative. It's not just that we throw, we're throwing out a bunch of credits at you and you're like, "Build it yourself," or like, you know, "Figure it out yourself."
There is a lot of cross-collaboration that's happening. So, for example, like, we spoke about GPT-4, GPT-3.5 Turbo, etc.,right? Our product team work very closely with the startup, so if there is any, um, private preview happening, we are working personally with startups to have them on board, to, for them to, like, try out these products, give us feedback that, "Hey, this is something, this is one of the feature capabilities I would like to have."
And that gets, like, passed on to the product team for them to work on. There is my team, which is AI advisors, which is working on, with startups one-on-one. So you're getting, like, expert advice. And as a startup, I think, like, getting advisors, getting time, timely advice, getting, like, timely, uh, resolutions is really, really important.
And those are the qualitative benefits that I personally feel is really important when you're working with Microsoft for Startups. It's not just AI advice. You can get a bunch of different, um, experts by, by you. Like, there is a, there is a tool within Founders Hub which helps you match with different experts.
You have a question about infrastructure, you have a question about Kubernetes, you have a question about how do I think about, like, my go-to-market, how do I price my subscription model. There could be a ton of different things that a startup is going to be doing, and you don't have enough resources to help you with that.
And that's where you get experts from Microsoft who are working in a variety of different products to help you navigate these challenges. So think about the qualitative benefits and the support that you're getting, apart from just the credits.
So that's, that's something I would like to add.
Yeah. And then, uh, we go with you throughout all the stages. If you only have an idea, or if you are building, if, if you are in the scale, uh, phase as well. So it's for any startup at any stage.
Um, as I mentioned, we are helping with all the cutting-edge AI tools and helping you streamline the AI development. And then we have something, like, very special as well, which is go, goes beyond the Founders Hub, which is the partnership that Ash is going to touch a little bit.
So one of the things I was talking about,right, Founders Hub is a platform that becomes, like, an intake for all the startups. This becomes your, like, if any of you are, like, familiar with, like, YC's platform or, like, port-portal,right?
Partnerships6:35
Similar to that, like, we have Founders Hub as a portal which has, like, end-to-end everything. It keeps track of how much credits you have consumed, what are the products that you've been using, what are the benefits that you've gotten.
It's not just the Azure credits. There's a ton of, like, third-party credits that you get and productivity tools that you get for free. And as a startup, you need all of these tools to, like, help you build that ecosystem.
So that's one of the most important things. Now, as a startup, you're getting started at Founders Hub. That's where you get, like, the $150,000 in credits. As you grow in your journey, if you are getting involved with one of the Microsoft strategic VC partners, which includes M12, Y Combinator, Neo, Alt Capital, Alchemist, to name a few, then you get a bunch of extra credits.
So outside of the $150,000, you get an extra credit, which becomes a total of $250,000 in credits. And then you get, like, um, startup development managers who are, like, you know, dedicated cloud, uh, solutions architects to work with you.
Pegasus program is another, uh, elite-level, uh, partnership program that we have with startups, which helps them go-to-market, which helps them get you onboarded on Azure Marketplace if you need help, like, you know, to amplify what you're doing. So you're coming and telling us that, "Hey, Microsoft, help us reach your, like, you know, millions of, uh, viewers, or, like, millions of people who follow you on your LinkedIn page or on your blog, and help us amplify our startups."
So we do blog pieces with you, we do video interviews with you, we do YouTube interviews with you, which helps you amplify that product. Outside of that, we're also getting startups involved in conferences. So, uh, Microsoft takes up a lot of, like, speaking slots at conferences, and we get our startups that we're working with come and talk about their products on stage.
Yeah, we have, we have one. For example, Nichola is coming to talk about their product, and they got, um, integrated to Azure so you can run the time gen, uh, uh, that they have. So they are coming to talk.
Yeah. So exactly. Like, so it's a lot of, like, amplification that you're getting from Microsoft's perspective as well.
AI Templates8:53
Yeah. Um, I think we briefly covered about some of the pain points. Uh, I'll give you, like, a quick TLDR,right? As a startup, if I'm building a startup, I really do not have a lot of time, uh, to, like, spend two weeks on getting a support ticket cleared off,right?
Like, I need quick solutions. I need really, really fast experimentation pace. I want to try out different things, different tools, different models very quickly, and then decide for myself what's the best fit for me. And that's one of the places where, uh, one of the pain points that we've heard from startups is that, "Hey, it's really hard to go through, like, tons of documentations and figure out how to get started, like, the first quick start program, uh, to, like, run a particular application.
It takes us weeks altogether to run our first end-to-end application." That's the solution that we wanted to target with AI templates. So, and one of the ma-major disclaimers, it's not 101-level examples. These are really, really complex examples. So one of the ones that we're going to be showing is RAGs with AI search.
So they are re these are really complex examples that you can run within minutes. And that's, that helps startups get started very quickly and run these applications, customize it as they want. Because it's open source, any of the issues that you're facing, just, uh, put it on GitHub and then it'll get resolved.
We have, uh, we have an entire, uh, cloud advocate team working on these templates. If you have any requests that, "Hey, help us with this particular other template," we're, we're here to help you on that.
Allright. So now it's the fun part. So now we are going to be doing, like, hands-on. Uh, we have the setup instructions in, in this URL. You can go over there or, or scan the QR code. And we are going to go through all the setup instructions.
Workshop Setup10:28
Actually, Pamela is going to go through it.
Okay. Well, let's make sure everybody has this URL because you are going to actually want to have that doc open. So it's aka.ms/aie-workshop, or you can scan the QR code. And it should open up a document that looks like the little screenshot there that says, uh, you know, "Quick start with AI templates with setup instructions."
And if you have any trouble, any issues, we are three, so you can raise your hand and then we'll go and help you.
Does anyone not have the doc URL yet?
Still working. Okay.
We need to just have a whiteboard. Remember whiteboards? Whiteboards were great. Yeah, this is the whiteboard. Do you think they'll mind if we just did you bring spray paint?
I'll get some shine.
Okay, great. Thanks. Allright. We good? Allright. Okay. Just holler at us. Because we're, we're I'm going to step through it first. So you still have time, but we want to make sure you all have that doc. Okay. Allright.
I just remembered how PowerPoint works. Okay. Allright. So, um, okay. So if we look at that, those instructions, um, the first thing is that we, you do need a GitHub account. Uh, so if for whatever reason you don't have a GitHub account yet, this is a good time to get it.
You should be able to get it for free. And if you do have a GitHub account, just make sure you are logged into your GitHub account. And what we've provided today for this workshop is two different things that are going to help you, uh, deploy these templates.
So one is an Azure pass. So this is really cool because normally things cost money. Um, but we're giving you an Azure pass, which will create an Azure subscription for you, which will give you up to $50 worth of credits, and it will expire in seven days.
So I guess you can keep working with it for the next seven days, and it'll certainly give you enough to get started during this workshop. So that's really cool. So you none of, none of you should have to worry about any of this costing you money.
Um, we don't want, we don't want that to happen. And then the other thing that we've got you is, uh, normally when you're using Azure OpenAI, you have to sign up to get permission to use it. You actually have to go fill in a form and get approved for it and, uh, get your account opened up for it.
So since it takes time to fill up that form, and the reason we do this is for responsible AI reasons. We want to make sure people are using Azure OpenAI for good reasons, which is good. I like, I like that about Microsoft.
Um, but it does take time. So we have set up this Azure OpenAI proxy that all of you will be able to use during this workshop. So you're going to use an Azure pass so that you can freely deploy these things, and then the Azure OpenAI proxy so that you don't have to worry about getting permission to use it.
Uh, so that's what a lot of this setup is. So the first step is to get that pass set up. So I'll demonstrate, uh, from here. Let me go to my Firefox.
Okay. So when you go to this Azure check-in link, uh, it's going to ask you to log in with GitHub. So then I can say, "Log in with GitHub." Uh, oh, let me do it in my I've got three browsers openright now.
Here we go. Let's do it in Chrome. Yeah.
You know, the, the initial link that, that takes you to the OneDrive, it says request is blocked.
Request is blocked? Uh-oh.
That's old. Everybody else?
Is, is was anybody else able to open the doc? Yeah.
I was trying to eventually open it.
Okay.
Yeah, the Wi-Fi was.
Oh, sometimes it you mean maybe the Wi-Fi's bad. It misinterprets it.
Yeah, mine wasn't even connected.
It could also just be that the URL was slightly wrong. Uh, I think that's what happens for AkaLinks.
It could also happen.
Okay. Allright. Cool. So then that has a link to this check-in website. So on the check-in website, we see that it has two options. Either create a GitHub account if you don't have if you don't have one, or just log in with GitHub.
So I'm going to click log in with GitHub. And there we go. I was already logged in in this browser. And then you'll see this Azure pass. So this is just a, a promo code that you're going to use for the next stage, and you can copy to clipboard.
And there's this button here that says get on board with Azure. So when you click on that, then you get brought to this screen here.
And, uh, it says what Microsoft account I'm currently signed in as. So it, you know, so at this point, you have to figure out what Microsoft account you want to use. Uh, so you could make up a new one.
You can just, like, make a new Outlook address. I made one this morning. It's no big deal. Uh, you could use a Gmail account. Um, I don't recommend using your work account if you do have a work account, just because work accounts tend to have a lot of restrictions on them, and I just think things will not work out.
So I recommend some sort of personal account. Either create a, a brand new Microsoft Outlook account today or just use your Gmail, um, whichever those you, you want to do. Um, so you do need to be logged into some Microsoft account.
So it says, "Okay, I'm going to be logging in with my Gmail." I confirm. I enter the promo code here. And then I have to do this horrible I don't like, what is oh, okay. DVW. Okay.
Hmm. Okay. And then in this case, I get an error, and that's because I did already redeem my pass, uh, for this, uh, for this account. Um, but you should get a success if you haven't redeemed your pass for the account yet.
And then that should give you set you up with this new subscription in your Azure portal. So you can see here now on my pamela.fox@gmail.com, I've got actually two subscriptions. Because I'm actually a paying user of Azure. So this is my, uh, paid Azure subscription.
You can see I'm spending $40 this month. And then here's my, uh, sponsorship, which I, you know, won't be paying for. Um, so that's and if you're doing a brand new account, you're only going to see this sponsorship here.
So that's the expected flow, uh, to being able to get set up with this Azure pass. Uh, you don't have to do itright now. You could wait until we, you know, break into hands-on time. Uh, because then it probably is easier for us to walk around in case there's any issues.
And just to be careful, make sure that in the next step, wherever you're using the subscription ID, you use it for the Azure sponsorship one and not the other one, or you're going to get billed on your, uh, on your credit card.
Yeah. Yeah. So if you're worried about that, just make a whole new account. That's what I have in my let's see. That's Firefox. So in, um, in this account, uh, in, in my other browser, this is an account I made, uh, this morning, uh, with just a new Outlook account, pamelafoxaifair@outlook.com.
It's pretty good. Allright. And that one only has in this case, I only have the sponsorship. Yeah, question.
Yeah. I've seen the quick one down Red Bull.
Okay.
Like, it took an account that I had for my son's Minecraft.
I see. Okay. So how many of you have sons with Minecraft? Allright. Uh, probably best that is why I used a whole new browser, not my normal browser. Um, so did it already if it already granted you the pass?
I just went with, you know, it came up with my email address.
Okay.
And that, and then it takes me to my account, but there's nothing here that says anything about Azure. And when I put in Azure, um, portal.
What do you do if you don't have the American number?
If you don't have?
Uh, I don't think it makes you it shouldn't make you oh, to get a whole new Outlook account?
Once I, uh, log in to my Microsoft account, it asks me to send you a pass information. And then if I enter my normal number, it says it's not a valid pass number.
Okay. I didn't remember that step. Maybe we can just give him our number.
Okay.
So if anyone else is having trouble, let me know. I can help.
Yeah. Do we want to just get through this stage now, or how do we want to do it? Sure. Like, do you want to, like, see what other issues everyone has? Yeah.
So, um, how many folks are working on getting the Azure pass redeemedright now? Are people working on that now? Okay.
And how many of you have actually successfully gotten it? Yeah. Okay. Good. Good. Okay. Cool. Okay. So we'll just get through that.
Because in other words, I'm missing things in the process of working account.
So what should I move on at some point, or should we wait?
Like a five-minute break?
Okay. Allright. So we'll spend five minutes getting through this step and answering any questions. So, uh, so go ahead. If you haven't yet, you know, try to get your Azure pass redeemed, and, uh, we'll just make sure we have enough time to get that redeemed for everyone.
You can pretty much use any account, like your university account, personal account.
Just not work account.
Just not work account.
With the university account, still be okay? Because university tenants sometimes have restrictions too, don't they?
Can try. I've used my university account. It works fine.
Okay. With these ones?
Yeah. With this one. Yeah. With Azure one. Yeah.
Okay. No, but with these templates. Okay. It might not work.
Um, do you all remember if we need to add the address and all of that? I don't remember when the Azure pass. You, you have to. Okay. So you have to.
Wait, so how did he get over the number stage?
I just, I just, like, asked him for a 204 number.
Oh, okay.
Starting 646.
Okay. Allright.
The next step is the proxy.
Okay.
We'll see until Gabriela's out of questions.
Yeah. Uh, raise your hands if you got that pass in. We're just trying to see. Uh, okay.
Activated it on Azure?
Okay.
Okay.
Okay.
Yeah.
No.
Oh, question.
Can you get me stuck here? Can I catch her? Um, I'm on that. I'll read it.
Yeah. See, and you click get on board. You already went through this part?
Yeah.
Sorry.
Um, yeah. Let's see.
Here it is.
Okay.
So we'll, we'll show that next. Yeah. So to confirm that you have the Azure thing working, you can go to portal.azure.com, and then it's going to look super empty. You'll see nothing under resources. But if you click on subscriptions, then that's where you should see Azure pass sponsorship.
So your subscription is basically like it's kind of like a billing account sort of thing. Um, so that's, that's how you know you've got it is if you see that under your portal under subscriptions.
Allright. So let's try now let's look at the proxy. Okay. So for the proxy, that's another URL which is linked from the docs. Um, so I'll go ahead and open that. I think I've already logged in on this one.
Okay. So what you're going to see is this page that looks like this, and it has this login with GitHub. So then I log in with GitHub.
And now I'm logged in. You can see it says welcome to my GitHub username at the top. And when I scroll down, I can see an API key and an endpoint. So here's the API key, and here is the endpoint.
So this is what we're going to be using in order to interact with OpenAI, uh, Azure OpenAI models. Um, and we'll just have to specify this key and this endpoint when we're using a template. So you're going to basically, like, you just keep this open so that you can continually copy and paste these two fields here.
Just, like, create a small note where you're, like, you have your subscription ID, your endpoint, and your key copy-pasted on that.
Yeah. And to clarify, normally I don't recommend using API keys, and I've got all these, like, videos and blog posts about how you should never use API keys. Because when you're actually using Azure API, uh, OpenAI, you can do keyless authentication using this thing called managed identity.
And we have a talk coming up next week about that. Um, but in order to use this proxy and, um, you know, take care of the permission issue, we are temporarily sinning and using keys. Okay. Uh, so that's, that's a proxy.
So you just have to log in, and then you should get the info for the proxy. Okay. Allright. So now I'm going to step through one of the actual templates. So, um, you know, all the templates are open source repos, but we've put together instructions specific to this workshop in, in this, uh, README here, these three READMEs, and really specific to using that proxy.
Quick Start25:49
So we're going to ramp up in terms of complexity. So we'll start off one that's, uh, really simple. Like, this is one where you, you know, it's, uh, just to show you how things are working and, and get things going.
And then we'll move on to two different RAG applications, uh, that are more sophisticated and ending with our most sophisticated one that has been deployed like 100,000 times at this point by Azure developers. So it's a very, very popular one.
Um, so let me start with, you know, this first one. So we go to the README, and the first step is to open the project using GitHub Codespaces. Uh, has anyone here used GitHub Codespaces? Okay. A few people.
Okay. So GitHub Codespaces is very cool. Any GitHub repository you go to, you can open them up in a code space. So you can, like, start hacking on that repository immediately. Uh, and then we can also customize, like, the environment for that.
So what it's going to open is actually a VS Code in the browser that has that project loaded in. So we're going to click on here.
Do you know one very cool thing about GitHub Codespaces is, like, you know, when you share something with, like, someone and they say, "Well, it was working on my machine, but it's not working on yours,"right? That problem that we all have with all the setting, the local environments, and all of that.
Codespaces, if you use Codespaces, you don't have that problem because I will have the same environment as Ash, as Pamela. So I will not have the problem, like, it's not working on my computer, but it's working on yours,right?
So it's one of the pain points that Codespaces came to, uh, solve. And all of us, we use Codespaces on a daily basis because setting up your local environment in your computer can be very, very painful. Um, so, so yes.
Does everybody have access to their Azure subscription by now? Anybody who's not? Okay.
If you're not, you can just.
Okay.
Okay. Can we just, like, try to solve that and then we go through?
Yeah.
Okay. Cool.
Actually, Azure.
The proxy.
Well, that's a proxy to the repo.
Yes. Uh, so the question is, how do you get to the repo? So there is a link.
This one?
Yeah. This link over here. Uh, it takes you to the GitHub repo with all the everything that you need with the code and everything.
No, I can't edit. Yeah. There's, like, a blank page. Skip the blank page and then go to the not blank page.
It can blame on me.
I'll help you.
Control F hands on.
So the other question was, how do you know if you have the subscription? You go to portal.azure.com and then you find subscription.
Yeah. And then you click on subscriptions. There we go. And then
so now, um, now I'm going to step through the instructions just for the quick start one, and then we'll, we'll really set everyone loose, um, and so that we can walk around. So this one is, yeah, just a simple chat application.
So we're going to start off with running this local. Well, I'm going to call it local. I'm inside GitHub Codespaces, which is VS Code in the browser in this containerized environment, but I'm going to start a local server inside GitHub Codespaces.
So I like to start with local development first when I can so I can, like, make sure that things are working. And then once I know it's working locally, then I can deploy it to Azure. So we're going to start with a local server inside the Codespaces.
Um, so looking at the instructions here. So I've got the project open. Uh, the first step is to make a .env. So we have a sample here. So I'm just going to copy and paste the stuff from the sample.
And, you know, the one negative about using Codespaces with conference Wi-Fi is that it is an online environment. So if you want, you are also welcome to try these out, these projects locally. This is just how we can guarantee less issues with developer environment setup.
Okay. So you can see in this environment file, we need to specify the endpoint and the key and the deployment. So for the endpoint, we're going to go to the proxy page and grab that endpoint URL and put that here.
And so it looks like this HBS polite ground, blah, blah, blah, blah, blah, slash API slash V1. That's what your endpoint should look like for all of these. So literally the OpenAI SDK is going to send requests to this endpoint and, you know, get responses from it.
The next step is the key. So we go here and we're going to copy and paste that key. There's my key. Love it. Uh, and then the deployment name. Uh, so this is the name of our deployment of our GPT model.
So if any of you use openai.com, okay. So on openai.com, you just use things by their model name. You just say, "Oh, I want to use GPT-4.0. I want to use GPT-3.5 Turbo." On Azure OpenAI, you actually make deployments of models where you say, "Okay, I want a new deployment of GPT-3.5 Turbo," and then you give that a name.
So we've actually named the deployment the same as the model here to make it either more or less confusing. Um, but that is one big difference between openai.com and Azure OpenAI is that the Azure OpenAI has this notion of deployments.
Uh, but everyone's going to put in GPT-3.5 Turbo here. Uh, so, so there we go. Okay. So I've got a .env, and the next step is to run the server.
Allright. So this is running the, um, the Python backend. We can open this up. And so this will, when I click on this URL inside the Codespaces, it'll actually open up a very different URL. So this is this port running on side the Codespaces.
So it's got this, like, really funky URL. So then I can say, like, write, uh, haiku about AI Engineer World's Fair. Okay. I always do all my testing, I say to write a haiku because otherwise LLMs get really, you know, verbose and you're just sitting there waiting forever.
So there we go. That is working. That's getting back responses. So you, you know, send any, uh, message that you can. Haiku about San Francisco.
And, uh, there we go. There we go. I actually just watched a video yesterday about how the Golden Gate Bridge was built. It's fascinating. Okay. So there we go. Now it is actually running. So this is running the app that's inside the source folder.
So if you want to explore it, you can. This is a court application. Has anyone actually heard of court? It's not well known. Okay. Who's heard of Flask? Allright. Court is just the asynchronous version of Flask. Like, literally probably Flask will become court at some point or vice versa.
So court is just Flask, but async. Uh, and we always want to use an async framework when we're building applications that make, uh, calls to LLMs because we want better concurrency. So you'll see for all our samples, they're all either using court or FastAPI for the Python backends because those are the ways that we can have async.
So you'll see, uh, async. If you haven't, you know, worked with async a lot in Python, you just see asyncs and awaits all over the place, um, because that's how we build, uh, with async backends. Uh, but this is the actual code that's happening here.
We're streaming in the response from the OpenAI client, and, uh, we're getting back the responses, and we're streaming it back to the front end using something called JSON lines or newline delimited JSON. It's just a way of streaming one line at a time.
So let me do a longer one just so I can show you, uh, how the streaming works. Do do do. And streaming is also another general practice. If you're if you're making a user-facing application that's making a call to an LLM, you really want to stream that response in ideally because then it's going to appear faster to the user because the time to get the first token, as soon as you get that first token, you can start streaming in that first token.
So we are actually streaming in a token at a time. So, like, write, uh, a long essay about San Francisco. Okay. So we should see this actually stream in here.
Uh, once it so it still takes some amount of time to get that first token, but then once you get that first token in, there we go. That was so fast. I wonder if the proxy is actually making it be a bit different.
Let me see if I can see it in the stream what happened. Um,
I should see it in the response. Fascinating. Um, normally I can see the stream tokens in the response here. I can try it with another of ours though. Okay. Allright. So we'll have to trust that. But, um, but yeah, there you go.
So this is just our, our getting started experience. So this is the local server,right? So we are running the local server, but we are hitting up Azure, uh, as our, uh, you know, so we are hitting up a cloud resource, and that's because it's hard to have a local, uh, a local GPT-3.5.
Now, if you want, you can actually use these models with, like, Ollama. I don't know if any of you use Ollama, but Ollama is a really great way to run small language models. And so I add support for Ollama to all my samples when possible.
So if you want locally, you can actually run against these small local models like Phi-3 or Llama 2 or whatever. Uh, it's just not going to be the same thing because they're different models, but that is an option for local development.
Allright. So that's all set up. The next step is to actually deploy this to Azure. So that's when we are, you know, going to be using that Azure account that you set up. So the first thing we have to do is log in.
So we're going to do azd auth login. And we're going to use the device code flow when we're inside a Codespaces. So that's going to have us copy and paste some, uh, some code here. So I press enter and it opens up this new tab.
And I'm actually going to open this up in, uh, another browser. We'll just pick a random one. Okay. Here we go. And then I'm going to grab the device code here. And then I put it into here.
Sign in. Okay. So I guess I'm using my Gmail here. Allright. So I am signed in. Okay. So now I've logged in. So you want to make sure you log in with the account that you just set up for that Azure pass.
Then we're going to create a new, uh, azd environment. This is kind of like a new deployment environment. So a lot of times when I'm developing these, these samples, I've got like 20 different deployment environments where I'm trying out different configurations and stuff.
So I'll make a new one here for this one in chat. Quick start. So you just give it a give it a little name. And then the next step is that we need to set some environment, some azd environment variables.
These are like our deployment variables that's going to tell the, um, the infrastructure how to provision everything. So we're going to tell it to not make Azure OpenAI because we're using the proxy. We're going to tell it the name of our deployment.
So I can just go ahead and copy and paste those two things. So I'll just paste them here.
Okay. And then I need to tell it the key. So this is the same key that we did earlier, but now this is going to be used by the actual deployment flow. So I go and find, oh, I deleted way too much.
Come back. Okay. So just delete that part and grab the key.
Allright. So now I've set the key for deployment. And then I'm going to set the endpoint for deployment.
And here we go. Where's that proxy at? There we go. Allright. Okay. So now I've set all these azd environment variables. These are going to be used when we are configuring the infrastructure. And so then we run azd up.
So what this is actually doing is that we're using this, um, infrastructure as code. Does anybody here use Terraform or Bicep or ARM? Okay. Yeah. So Terraform is probably the more well-known one. So at Azure, we have our own version.
Um, originally it was ARM and it was JSON. Now we have Bicep, which is like a better version of ARM, stronger. Uh, but these are all our Bicep files. You can also write Terraform if you want to do that.
I just, I know Bicep more than Terraform. So we are using Bicep, which is infrastructure as code. And that Bicep describes how everything is going to be made. So how are we going to make the OpenAI? How are we going to make analytics, container apps, uh, roles, all that stuff.
Allright. So I need to select a subscription to use. So I'm going to use the sponsorship subscription. Um, for many of you, you might just have one subscription and then a location. This is going to be where, like, our container app is going to go.
So I'll just pick a random location there. Okay. And now it is packaging everything up. Uh, so it is, yeah, it's actually building a Docker image. So this one gets deployed to Azure container apps. That's like an Azure option for running containerized applications.
We also have like Azure App Service, Azure Functions, Azure Kubernetes. Uh, but for this one and the second one, we're using container apps as a, you know, a nice place if you're using Docker. How many of you like Docker?
I shouldn't say like. How many of you use Docker? Okay. It's a similar, okay. You know, I don't want to presume. Um, so if you do like Docker, you know, Dockerized, um, environments, you know, this is using a Docker file.
Uh, you can see the Docker file here. Uh, you know, installing requirements, running the server, all that sort of stuff here. Uh, so then it's provisioning the resources. So it's going to do that whole step. And then I've already got one, uh, you know, pre, uh, already deployed here.
So we, if once it's deployed, we'll have a container apps URL. And this is a URL that you can, you know, tweet, share publicly, whatever. Try not to use up all, I mean, I guess use up all your credits, whatever.
Uh,
uh, hi LLM, what's up? I don't know. I never know what to say. Um, I don't have feelings, thanks. Uh, so now it is deployed there. And then if I can, I can look at my portal and see, you know, actually see what it made.
So I can go here and look at maybe container apps. And, uh, that's actually a different one. So let me look for, let me go.
I know. I think I just have to remember which thing I, I, which I have got three different portals goingright now. Here we go. Uh, so we'll do, uh, uh, quick start.
Maybe this one. There we go. So once it's all deployed, you can go into your portal and actually find the resource group that was made, and then you can see what was made underneath it. So here we have a container app.
You need to register. These are everything we need in order to make a containerized app. So that's the flow for that one. The flow is similar for the other ones, but the other ones are a lot more, uh, a lot more sophisticated.
RAG Demos41:55
I'm saying like they're ones that people are actually using, uh, in production. So I'll just talk about RAG. Um, RAG stands for Retrieval Augmented Generation. This is our solution for the fact that LLMs like to make stuff up.
I mean, that's kind of their, it's kind of their, the way they work. They're just word prediction machines. And so if you get them to predict something that they don't know, then they'll, they'll go ahead and predict something,right?
So how do we get LLMs to give us reliable output for a particular domain? We can use Retrieval Augmented Generation. And so how this works is that we get in a user question. We use that to search some sort of database, whether it's a, you know, a search engine, a vector database, whatever you want.
We search it, we get back results, and then we send both the original user question and the search results to the LLM and say, "Hey, now please answer the user question based off the search results." And then you'll get a really good answer.
So as long as you have a very good search engine, so you really want to pick really good retrieval mechanism search engine at that step, 'cause if you get good results, then you'll get a great answer from the LLM, 'cause LLMs are incredibly good at synthesizing information, summarizing, uh, based on what they see.
Uh, so they just need to have this, you know, really good search step. So we have two different RAG options that you can try deploying today, uh, and over the next seven days. So the first one is RAG on Postgres.
And so this is if you like already had a, a database. Imagine you've got like a retail website, you've got a bunch of products that you're selling, and you wanted your customers to be able to ask questions about those products,right?
So you can, uh, you can just search off those table rows for, you know, what the user is asking about, get back the matching table rows, and then you pass those table rows to the LLM and say, "Hey, answer this user question about the table rows."
I'll show you, uh, what that actually, uh, looks like when deployed here. So here I've got, you know, a, a product table for this outdoor, uh, outdoor shoe company. And so the user puts in, uh, puts in a question here, and we get back all these results and where they cite, you know, this is just info from the rows.
And we can look at the thought process here. And so we can see, um, that actually, actually it's even I did the fancy one. Okay. I'll use the simple flow first so I can show the simple flow.
And, uh, and then we'll move on to the advanced, advanced RAG. Okay. Allright. So here now if we look at the thought process here, we get the search query, we use that to query the database. So these are the rows we get back from the database.
And we do both a vector search and a text search. Now, I'm sure you've heard lots of things about vector databases. They're great, but you need vector search and text search, and you need to combine those results together.
If you use vector search alone, you will not get good results. I've done like hundreds of evaluations of, of this sort of thing. You need to have a hybrid search, which is going to do both a vector search and a text search.
So for Postgres, we can use pgvector for vector search, and then we can use their built-in full text search for text search, and then we can combine them together. Uh, so we get back results, and then we, you know, send it to the model.
We say, "Hey, your job is to answer questions based off of sources. Here's the user question, and here's the sources." So at its simplest, this is what RAG is. This is the actual call that we make to the model is, "Please answer according to these sources.
Here's the question. Here's the sources." We get back the response. Now, uh, we can get a little fancier with that. So we go to the, um, advanced flow here. And, uh, and I say it's fancy, but I think it's actually what most people are, are doing at this point for their RAG at the least.
So in this flow here, the first thing we do is we take the user's question and we rewrite it into a better query. 'Cause user questions aren't really optimized for searching databases or searching search engines. So we first ask an LLM like, "Hey, here's a user query.
Make this into a better query." So this is what we can call the query cleanup phase or the query rewriting phase. And it's a really useful first stage to have in a RAG application. And so then we get back, uh, you know, search results.
And then, uh, well, we get back the query,right? So in this case, uh, it actually ended up giving the same query in this example. And then we get back the results, and then we send it. But this gets particularly helpful when we have multi-turn conversations like, uh, if I type in, let's see if it's going to perform for me today, more options,right?
More options on its own is a terrible query to send to a search engine,right? What is more options? More options about what? So I'm hoping that my query rewriting phase is going to clean this up. Now, yeah, they're seeing I should test this before I do it.
Allright. So I can demonstrate that more in, um, in our other example 'cause I have done that demo more. This one's repo's a little fresher. I made it like a month ago. Okay. Uh, so that's what we're going to need for it.
Now, another thing we can do in the query writing phase though, is that we can actually use OpenAI function calling in order to get the model to generate SQL filters for us. Uh, so that's what we've done here is that the user asked, "I want climbing gear cheaper than $30."
Well, we can make that into a SQL filter. So we actually ask the LLM, we use OpenAI function calling and say, "Hey, can you tell us if there, if we should do a price filter here?" And so it comes back and says, "Yeah, you should do a price.
It should be less than 30." And then we can use that to construct a SQL filter. So that's another really cool thing about having that first query rewriting phase is that then you can start doing more sophisticated things and, and having it actually, uh, come up with more structured queries, uh, and not just do just like a full text, um, you know, full text search.
Okay. So that's RAG on Postgres. And so there we saw the flow of that. The other one that we have, and this is the one that's super popular that's been deployed thousands of times, and this is what many people think of when they think of RAG is being able to do RAG on unstructured documents.
So you've got PDFs and docs and Excels and HTML or whatever,right? You've got all these documents, and you want to be able to ask questions about it. And people are really excited about being able to finally ask questions about PDFs 'cause then we don't have to open PDFs 'cause nobody wants to open a PDF,right?
So we can do, uh, we can do RAG on documents. And so that's what this demo does here. And so let me go ahead and, uh, show the deployed version of that one,right? So I asked, "What does a product manager do?"
I get back citations. I click on a citation that will load in the particular page number that it got it from,right? So this is the PDF and the page number where we got it from. Uh, and this is, you know, a question that's specific to this particular employee handbook.
And we could do this RAG with anything. Um, and so here you can see the, you know, all the citations it found. And for the thought process here, it's similar,right? We have a query rewriting phase, we have the search results phase, and we've got the prompt to generate the answer.
So a lot of this is, is really similar. The big difference is that here we have to have a data ingestion phase because we need to figure out a way to take these like, you know, 50 page long PDFs or something and store them in, you know, in a searchable way.
So we have a data ingestion that will take a PDF. Uh, we crack it using, uh, Azure Document Intelligence, which is very good extracting text from documents. Then we chunk it. Um, we do a token-based chunking. So we try to come up with chunks that are about 500 tokens large, and then we vectorize those chunks with, you know, the OpenAI embedding models, and then we store them into Azure AI search.
So that's the data ingestion phase. Uh, so you're going to need, that's why this is the most complicated architecture is because of that, uh, that data ingestion phase there,right? So data ingestion, we're using Document Intelligence, Azure storage to store them, Azure OpenAI to embed, and then Azure AI search to store that there.
Uh, but then it's really cool, and you can do it with all sorts of things. So I, uh, I've got one that has my blog in it, so I can ask questions about myself. Let me get that one open.
Uh, let's see my blog. Here we go. Boop.
And there we go. Good sleep strategies. I was just telling them how bad I am at sleep, but I have researched it a lot because I'm so bad at it. These are all from my blogs.
Okay. Uh, and that's from, this is from parsing in like an HTML site. Um,
yeah. So there you go. So those are, those are our, there's my blog. Uh, those are, those are the RAG ones. And so with, um, RAG with the Azure AI search, that's one where you could immediately get start, like get started with putting your own documents into it and seeing what it's like to be able to chat off them.
The slides we await the.
What?
The slides we will await them on something.
Oh yeah, good question. Can you put them somewhere?
Yeah. Uh, we'll put it in the same doc.
Okay.
The same Word doc. We can put the slides over there. Uh.
Okay. And I'm, I'm done.
Okay. Cool. I know like let's go for the questions and then I know that some people were a little bit behind. I want to make sure that you have something running. Uh, of course, everything that Pamela showed you, uh, you are going to be able to do in your own time.
Like we ran this over and over and over again. We tried different things. We customized the HTML of the chat. Uh, we also changed the, the, the, the, the message. Uh, like for example, one of the things that you can do is like instead of like you are AI assistants, um, I can say you only know about Nintendo.
Any other thing, just say woho. Right? So you can do things like that. So there is a lot of like customization that you can do on this, on these applications. Uh, as she was showing you, you can use your own data.
She was showing with the data from her blog,right? Uh, but I want to make sure that everybody's more or less on the same page. Or if you have any questions like, "Whoa, Pamela, what did you just do? I have no idea what is RAG or any questions so we can help you get up to speed."
Q&A52:48
Um, I have to create a document AI and Azure search basically, but I can one, uh, is it more the PDF you can parse or would you use something else for parsing?
Uh, yeah, great question. So what, um, formats can it handle? Uh, so now Azure Document Intelligence can handle quite a few. Uh, so we have the, let me find the document about it. Um, data ingestion. Okay. Uh, so these are all the ones supported by DI,right?
PDF, HTML, DOCX, PBX, XLS images are all supported from Document Intelligence. And then we built our own parsers for text.
Images means?
What is what?
What is images means?
Right. So if you put, yeah, it's a good question. You see images. So if you send images to Document Intelligence, it'll OCR them basically. It'll extract the text. Now, you different thing is that you might want to do, um, like a GPT vision, uh, which is, which is a whole different thing, which is where you're actually asking, sending an image to GPT like 4.0 and asking a question about it.
Now, this repo does actually optionally support that. So if that's something you're interested in, you can try that out. That's, that's actually different from ingesting an image. Um, slightly different process, but that's, that's also an option. So it just depends what are your images and what are you hoping to get out of them.
Yeah.
What about like PDFs with images inside that have the text that is contextual to what's going on? Can it do multi-model?
So Document Intelligence will extract as much as it possibly can. To me, that's sometimes too much, but I did have an, an incident where, uh, I did another one based off the Python Playwright documentation, and the Python Playwright docs has some images of the node Playwright.
And so it actually extracted the JavaScript out of those images, and then that messed up my whole RAG, uh, because it did extract it. So Document Intelligence will generally try to extract as much as it can. And so if there are text in the images, it'll just bring those out as text.
For search, can we do more than just like embeddings? Like can you have metadata query on top of that as well?
Yeah. Yeah. So you could do, you can add additional fields, um, where you just mark them as searchable, and then they would get searched as part of the full text search. Uh, so I think we even start off with three different searchable fields, but you can, yeah, you can add fields, you mark them as searchable, then they'll get searched as part of the full text search.
In terms of the vector, only what you vectorize will be searched. So if you do really want something to be searched in the vector search, then you'd want to do like what's called a content stuffing or content expansion, which is where you take everything and you stuff it into the same field and then you vectorize it.
Um, which is totally something you can try if you think it's going to be useful for vector. But I, I like I have to like really warn about vectors is that you, you just want to be careful with your vectors 'cause sometimes we like put too much faith in, in vectors.
Like I'll show you like the blog post I did, uh, last week where I ran, um, I ran the, the, the stats. Um, so vector search is not enough,right? But look at the stats, uh, down here. Um, the other thing you should do is evaluate.
So if I did, this is a text only search, it got a groundedness rating of 4.87, which is fairly high. And, um, I only was able to get 0.02 improvement by moving to text plus vector. So you'll hear a lot about vector, but, but please remember to use hybrid and to use good hybrid.
So if you just, if you just blindly like, you know, if you just combine vector and text and kind of just use a basic algorithm, which is the reciprocal rank fusion algorithm, then you'll actually get pretty poor results 'cause those vector results, 'cause you have to remember with, uh, when you do vectors and you do a search across vectors, you will always get results 'cause it's going to give you the most similar, even if it's really far apart,right?
So that's the danger of vector search is that you're always going to get results, and those results might be noisy, they might be distraction, and if you distract an LLM, it is very, very distractible. So, um, yeah. So, so the best, like with Azure AI search, this is the AI search, the best results is if you do hybrid with their semantic ranker model, and that's an additional machine learning model that actually re-ranks results according to the original user query.
And so that's the only way in my experience that I can use vectors and actually get, you know, get back to the results of a, a, you know, or better than a full text search. So just to, just to plea to evaluate and be careful with your vectors.
Any, any other question or do you want us to repeat like one of the apps? We can go over slowly, making sure that you are getting every little step 'cause some people were having issues with the key and not able to run it.
Uh, yeah.
Yeah. Okay. But we can also walk, like walk around,right?
Yeah. Yeah. Uh, like, okay, I have a question then for you. Uh, were any, like, were you able to get at least one thing running? Okay. Anybody didn't get anything running? Like have no idea what we are doing.
Okay. So we are going to help you. Um, you, you didn't get anything running?
It said it was running, but then, uh, you know, it was running.
Okay. And before that, I know that some people will probably walk around, uh, but I really, really want something from you all. So, uh, one of the things that I want you all, and you can say for later because this is very important for us, is like we love feedback.
So we give workshops and talks all the time, and we want to improve,right? But don't say like, "Oh, I wish the Wi-Fi was better." Sure. Yeah. But this is out of control,right? So if you can take a picture, I would really appreciate any feedback that you have for us to improve or to make this, this is like very dense, this workshop.
Like there is a lot. Uh, we try to compress so you have like a lot of like materials that you can go and work by yourself. Um.
Yes.
So I have one of the questions for people who are like either working at a startup or like want to build a startup or are currently a founder of a startup. Would love to know what are some of the use cases that you are working on?
And does it like closely relate to any of the AI templates that we showed today? If not, like tell us more because this is not the only templates. This is just like three of the examples that we showed in the, for the workshop.
There's tons of more AI templates available on the GitHub repository. But for now, like would love to learn more about like, you know, some of the use cases that you are working on or you're interested in building and maybe like, you know, we can, we can answer some specific questions.
So it's, is somebody interested in sharing a use case that they have? You can win some prizes. I'm kidding.
But yeah, like really appreciate if you have any questions. Yeah. Uh, any like examples?
Uh, so one thing that it's, it's a good practice now that I'm seeing is when you open code spaces, you are paying for it. So make sure that you either pause it or you delete it. Otherwise, your GitHub account, uh.
You're paying for it if you're past 60 hours. 60 hours?
Yeah. You have 60 hours for free, butright? Every month, I guess.
Yeah.
But make sure that, uh, if you go to github.com/codespaces, you can see everything that is running because I saw someone had like three or four or five. So pay attention to that. Uh, so you don't.
Is there like a timeout after? Was it equal to like five days?
I don't, you can, you can exp, you can set, you can, can I, can I do a, you can set up, like you can say in my case, if, if it's like I'm not using for more than 30 hours, I just shut down automatically.
Uh, but it's something that sometimes you forget. Um,
yeah. So make sure that let's see how many Pamela has. Pamela has a bunch of them running, as you can see.
Yeah. So yeah, quite a few people are actually using this in production, which, um, at first we were a little surprised, Mike, 'cause originally it was a sample, but now we've really hardened it. Uh, so people are using it for public facing stuff like, um, government websites using it 'cause governments have lots of PDFs and documents and it's hard to sort through their stuff,right?
So making it easier for citizens to interact with government data is one use case. The other big one is internal HR stuff or internal meeting transcripts. Um, like being able to look through all the transcripts and ask like, you know, what did the CEO say then?
Um, uh, internal sales training manuals. Like there's, there's like, there's so many people using it for lots of things.
Yeah. 'Cause it's being automated. You know, anytime there's like new documentation that you need to just have it like have a script or whatever.
Yeah. So in, in terms of being able to automatically update the index, like you just had it in manual ingestion. Um, but you could use instead integrated vectorization, which is an AI search, uh, option where you set up an indexer.
So you would point it at like a blob storage, your, uh, blob storage and say, "Hey, every time this updates, every five minutes, make sure you refresh the index." And then it would, uh, refresh it. Uh, so that would be one option.
Or you could use like an Azure function with a trigger. And that's, uh, there's another, another repo that does that. Uh, it's the chat with your data solution, uh, accelerator repo. And this one sets up an Azure function that has a trigger.
So, um, you know, we've got a few different options for how you could keep it updated. Uh, people just figure out what's, you know, what works out for them.
Cool. Thank you so much.
Yeah.
Thank you.
Sure.





