Welcome0:00
Hi everyone, thank you for joining us for the session on how you can accelerate your AI journey with the Azure AI model catalog. This is Shubhi, I'm one of the product managers on Azure AI.
And I'm Sharmila, I'm one of the product marketing managers on Azure AI.
Great, so let's get started. So today we'll be talking about how we offer the best collection of foundation models on Azure, why the model choice matters so much when you have such a huge collection of models, and how does the Azure AI model inference API make it standard and easy for you to switch between multiple models, swap one out for another without disturbing your rest of the code base, and then how you can build generative AI apps on top of it, and how the platform makes sure that all your enterprise readiness needs, like data privacy, security, and content filtering, are met.
And then we'll talk a bit about some of the customer success stories.
Model Choice1:02
So Shubhi, as you mentioned, there are a lot of large language models that are being released pretty much every day, every week. And I know in the Azure AI Studio and model catalog, we are bringing a lot of models, like large language models, foundational models.
I want to understand why are we bringing all these models, and why does it matter to a customer?
Yep, absolutely. That's a very valid question. So let's talk about why model choice matters in the real world,right, as Sharmila asked. So the three questions that we all try to answer when we're trying to build generative AI apps are: can AI solve my use case first?
And then what is the best model for my use case? And then how do I go about scaling this for the production workloads? So the first step is called prototyping, where you try out multiple models that are available to you.
You try to establish feasibility, build that prototype, do compare model benchmarks, and find theright one for your use case, or maybe at least shortlist some of them. And that's when you move to the next stage where you try to optimize it for your use case, where you try to optimize for cost, latency, regional nuances, etc.
And that's where the Azure AI platform comes into picture, where you can use the techniques like prompt engineering, RAG, and fine-tuning to make sure that you've optimized it. Once you're done through this loop of prototyping with the model catalog and optimizing with the platform, you can go ahead and operationalize it.
So you don't have to worry much about the capacity and cost trade-offs. You have got monitoring, you've got scalability, you've got data privacy, content filtering, and all your enterprise needs are met. So once you're through this, you have your generative AI application in production.
So another question I had here is, have you seen situations where a customer is using more than one model for a specific use case and even for multiple use cases?
Multi-Model2:45
Yeah, absolutely. Like, that's the whole point of the model catalog, that you have this wide range of models, you have your use case, and you're able to plug in any model from that model catalog. So you can swap out one for another without disturbing anything, with quick prototyping and quick comparison.
So as we mentioned, let's talk a bit about the model selections that we offer on our platform. So we offer a wide range of flagship LLMs and SLMs. Recently, we launched the Azure OpenAI models. GPT-4 is already on the platform.
Catalog Overview3:15
We launched Mistral models, Llama models, Cohere models, as well as small language models like Phi3 and the Mistral OSS models. Along with that, we make sure that your multi-model requirements, image generation requirements, and specific needs, such as embedding model requirements, are also fulfilled on this platform.
So GPT-4o, for example, has function calling and JSON support, so you can make sure you use that for your agent-centric workflows. Along with that, we also make sure that we cater to your region or language-specific needs. So for example, the Mistral language is really good with the European languages, or the Cohere embedding multilingual model is very good for your multilingual requirements.
And we also recently launched JS on a platform, which is an Arabic LLM. Along with all these flagship and premium models, we also have hundreds of open models from Hugging Face, and we've been actively partnering with Meta, Databricks, Snowflake, and NVIDIA to make sure that we get their models on our platform as soon as possible.
Okay, so now that we have seen what all we can do with the catalog, let's try to see a live example of how to actually go to the catalog and deploy your models. So once you type into your URL, ai.azure.com, it's as easy.
Live Deployment4:26
You land on this AI Studio page where you can go to the model catalog on this left nav bar, and you land on this page that has the list of all the models that we offer. It has 1,600-plus modelsright here.
And to make it easy for you to filter it for your use case, you can filter by the different deployment types, different inference tasks that you want to do, or even if just by the model collection families. So let's start out by filtering for the Cohere models.
Like, click on Cohereright here. Let's try to see Cohere Command R, for example. Once you click on the model, you land on this model card page where you have all the information about the model, how you can customize it for your own use case, tool use capabilities.
This one specifically talks about RAG capabilities because of the Cohere Command R model. All model catalog pages also have these inference samples that you can use as starter codes to get started with, for example, a LangChain SDK or a LiteLLM SDK.
As we mentioned, like, we try to standardize the APIs across all use cases. So you can just plug in your R APIs into any third-party application like the LiteLLM and have it working in no time. So once you've gone over this, check the pricing.
We go ahead and click on deploy. And this is where we make sure that we are connecting to the Azure Marketplace. So we use Azure Marketplace just for the billing side of things to make sure that you're billed correctly based on your token usage.
And this is the step where you actually subscribe to that offer. Here, I've already subscribed to that Microsoft subscription, so it's giving me the option to continue to deploy. It's as simple as choosing a deployment name, checking if you want to enable content filter or not, and clicking on deploy.
So under a minute, you'll have your URL and key ready to get started. So while this is happening, let's look at other capabilities that we have, like model benchmarks. So when you're in the Azure AI Studio, you also have the ability to check model benchmarks, which isright here on the left under model catalog.
Once you go in here, here I'm showing all the models that we have. You can see that we try to benchmark on certain common characteristics like the model accuracy, model coherence, groundedness, etc. And this is a perfect place for you to filter out which model you want to choose based on the extreme selections of models that we offer.
So let's go back to check. Oh yeah, and we see we go back to the deployment that we created, and it got created within a few seconds. We have our target URLright here and the key, ready to use in any code base that you already have.
So now you may be thinking that before I move on to using my IDE, I want to try it out a bit,right? Is there something like a playground? And that's where we also have this playground capability. So once it's also in the Azure AI Studio on the left, if you see, we'reright under the chat playground.
RAG Playground7:27
So let's see a live exampleright here. You can choose the model that you want to use in the playground. So in this deployment section, I've chosen a Mistral large deployment that I already have in this project. Let's try to chat with this model.
Right here, it's not customized on any data. I'm just directly asking the model. So I'm trying to ask, how does Microsoft promote the culture of giving? So this will, in general, give me a generic response about how it has a culture of giving through various initiatives.
It has employee match programs and some generic information that's available online. But what if you want to specialize it for our own data? So here we can go ahead and use the add your data functionality, where you can choose an available index.
So let's choose an available index that's called Microsoft Give that tells it in specific that what are the specific things that are very particularly known internally or may not be available in generic circumstances. If we send out the same questionright here, we should get a more targeted response based on the documents in that index.
So just wait for a few seconds.
So Shubhi, while this is happening, I had a quick question. It's great that we are doing all this. I'm just curious because you mentioned data and data source and everything. Are we using any of the data from our customers to train the models, or is Microsoft using it?
Are our model providers using any of the data that a customer brings in?
So that's a great question, Sharmila, because that's a very common question we get from our customers. And no, we have very strict data privacy and policy rules in place. Your prompts and your completions are not shared with the model provider, nor your data is used for training any of the models.
So yeah, looking back at the results, we see that it gave us a very specific response that says, you get 50 USD to start off with a new hire credit for the giving program. And that wasn't in the response earlier.
So with just the click of a button, we were able to link it to an index and get that response. So that's how the playground works.
Deploy Options9:27
So talking about the different ways of deployment, the one that we just saw was a serverless API option. So in the model catalog, there are two ways you can deploy a model. One's called the managed compute, and one's called the serverless API option.
With managed compute, the user is responsible for getting their own GPU. So you basically pay for the VMs per hour, and you're responsible for the quota management, capacity management, and you can use hundreds of open-source models with this.
The second way you can deploy models is by getting a serverless API. And this is available with both Azure OpenAI service and models as a service. And this is what has about 30-plus flagship models, premium models that you pay for based on your usage.
So you get ready-to-use APIs, and you only pay for the input tokens or the output tokens that you use. We've also put in a lot of effort to make sure that we standardize the schema and the APIs of these models for you.
So we've worked with the model providers to make sure that we build an SDK on top of a very standardized REST API system. And such an SDK works with common open-sourced applications, things like LangChain, as well as the model provider-specific SDKs.
So all you have to have is a different endpoint, and every endpoint has the same API structure and the same SDK structure. So you can just swap in one for another, evaluate, create multiple evaluations, compare the results, and choose the one that's perfect for your use case.
Okay. Awesome. So we've been seeing everything about the model catalog and models. Can you show us an actual use case example?
Function Calling11:08
Yeah. So let's briefly talk about how you can actually use these APIs in your IDE. Let's talk about the function calling example, and let's take the Mistral large model for that use case. So here I have a simple function calling example where I'm trying to use this model as a chatbot for an example of a shop.
The shop sells certain stationery items, it has certain specific pricing, and may have certain ongoing discounts. So if you just use a model as a black box, you will not get the specific pricing for the model or for the shop or any of the ongoing discounts.
But what I'm doing here is using the function calling capability of this Mistral large model to define a function called get bill amount that can take in the specific information that we fed to it, recognize that it needs to call this external function based on the prompt, and smartly make that call, query that result, and give you the exact information.
Soright here, I've defined that function. I've defined the tool for that model for Mistral large. We send in a prompt that says, you're a helpful assistant that helps users find how much they have to pay. And we also make sure that we tell it that you also care about the environment, and you also have to help users understand possible things they should be careful of when using these items.
So this is just to add more context to the response and see how the model can adjust based on the requirement. We go ahead, we send this response in. We can see that the model has intelligently identified that it is calling the function get bill amount with theright arguments.
So it identified that we queried for a stapler, and we tried to ask, what is the price for the 10 staplers? And if we see the chat response, it says, the cost of 10 staplers, including any ongoing discounts, is $45, which is very specific.
And it also makes sure that it reminds the users to be mindful of the environment and try to use staples when possible. And this is a result of the system message that we sent to it when we asked it to be environmentally friendly and give users theright context.
So similarly, if you seeright here, you can swap any code base with any endpoint that you have. And without putting much time into it or much effort into it, you have a running API appright here. We also make sure that we use model provider fields like the safe prompt setting to true.
So on top of, we always build on top of what the model provider capabilities are already existing.
Awesome. Thank you.
Prompt Flow13:42
So now that we've talked about how we can set it up in the IDE, let's talk a bit about how you can set it up in the UI and how you can create a generative AI app using Prompt Flow.
So when you try to, here I'm trying to create a shopping assistant chatbot using Prompt Flow, where it's a simple RAG application where we take in the user prompt, we try to get retrieval, we retrieve context-specific information from our index, and then send it to the LLM to generate an output.
Here, we've created the lookup step for it, which is basically doing the RAG part of it. The generate part is going to generate the output from the LLM. But we've added this extra step of rephrasing where we're using the query transformation technique where we take in the user prompt, which is generally very succinct, but we try to make it more verbose by rephrasing it because we've seen better results of RAG with that.
So let's look at a live example of thisright here. I have this Prompt Flow runningright here. My compute session is running. We see that the first question that we ask is, do you have any new hiking shoes? But the rephrase step rephrases it into a longer verbose output that says, I'm looking for hiking shoes available, and if so, what materials and features?
So it basically elongated that question. We check the output of the lookup step, and we see that the prompt was able to get the context-specific information from the index that we provided to it. So it identified a certain amount of information that we can now send to our LLM in the next step, and the LLM generates a response.
So based on that specific information, we were able to get this output that recommended the fleece-fit flex jacket for the women. So we can see that we are able to generate a Prompt Flow end-to-end. But you may be wondering that how do I make sure that I'm able to plug in different models into this flow?
And this is where you can try to create variants. So here, you can see that in the generate step, I'm using a connection from the cohere command R model. But you can go ahead and choose any other connection to any other model and use evaluation to try to compare the results for the same flow for different models.
So example, for the first step when we're trying to create the embeddings, here I've used the Ada model. But you can go ahead and try to see, OK, how does the command R model work with the command embed model?
So you can create these variants and try to see the evaluation results. In the interest of time, I already ran some evaluations, as you also saw in the previous demo. And here we're trying to compare the cohere command R versus Llama3 versus the Mistral large.
And the evaluation capability helps you to compare the same model for the same flow on different parameters, and you can see how one fed against another.
So this is all great, and I think you touched upon data privacy a little bit. So can you go a little bit more into the details of what else do we have in the AI Studio or model catalog to ensure customers' privacy and data security?
Safety & Privacy16:26
Yeah, absolutely. That's the key. So let's talk a bit about how we ensure that the data privacy and security compliance needs are met. So as I mentioned, there are three pillars to this. So one is the data privacy part.
Second is the security and compliance. And the third is the responsible AI and the content safety. Talking about data privacy, for both managed compute and serverless APIs, your prompts and completions are not shared with the model provider. Your prompts and completions are not used for training the models.
No data is shared for training or with the model provider. So you can be assured that the AI platform makes sure that your enterprise needs for data privacy are met. We also have this additional feature of adding a hidden layer to our model scanning.
So we make sure that we are finding the embedded malware and backdoors. It scans for common vulnerabilities and exposures and detects tampering and corruption across model layers. So for any model that's labeled curated by Azure AI, you can be assured that it's passing through the required checks.
Talking about security and compliance, in addition to the data privacy norms that we mentioned, we also offer the capabilities of adding private networking so that your data is not exposed to the internet. So you have the control over routing your ingress and egress traffic through the VNets, and you also have the ability to set up FQDN rules so you can approve outbound access to non-Azure resources.
In addition to this, you can also regulate access to models with Azure policy integration. So you can have allowlist or denylist patterns, and you can split out which model collections you want access to or not. And you can also use these different policies for separating out the dev, test, and production environments.
Awesome. So Shubhi, thank you so much for going through the model catalog and AI Studio. What are the features available in there and all these great demos? So now I just want to go into a few customer success stories.
Customer Stories18:26
And one of the main kind of underlying themes for all the customer success stories that I'm going to show is that these customers are not just using one model for their use case. They're using multiple models from the model catalog, and it could be in one use case or across multiple use cases, similar to what Shubhi has shown in the demo.
And the first customer we're going to talk about is EY. And they've been using our large language models. They've started off with the OpenAI models that were available in the Azure OpenAI service. They are doing that today. And they're also looking into Llama models, where they are looking into Llama models for really task-specific use cases, like for documentation and for summarization and all that.
And they're using our model catalog. They built eyy.ai, which is a generative AI platform for EY professionals, which addresses the need for enhanced productivity and accuracy in professional tasks. And one of the key things that we want to show is what's the result of what they've been doing.
The EYQ chat that they built has been adopted by 275,000 employees internally and allowed their employees to perform a wide range of tasks efficiently and with great accuracy. And some of the lessons that they have learned is using AI in their use cases is not like a one-time thing.
They need to do continuous evaluation of AI performance, and they want to stick to all the responsible AI practices. And that's one main reason why they've been using Azure AI model catalog and Azure AI Studio. It's because they feel that they can easily do this evaluation and make sure whatever they're putting in production is going to be really adhering to safe and responsible AI.
And then the next customer I want to talk about is CMA CGM. Again, they are a big global player in sea, land, air, and logistics solutions. And they're also building a kind of like a robotic process. They've been doing traditional robotic process automation, and they decided to use Mistral model from a model catalog.
And they have built, again, a similar chatbot-like scenario for their customer care agents. And one thing that they have seen is they have seen a reduced response latency and increased customer satisfaction in their chatbot use case. And they plan to, again, extend the application of LLMs to encompass specific products for core business activities like invoicing, customer document analysis, interpreting free client text, and writing emails, and all that.
And finally, the last customer story I want to share is Bridgestone. Again, Bridgestone is a very popular name. Their use case is they have been using the Nixla model from a model catalog. We launched Nixla time series model at build last month, and it's a time series forecasting model.
And they're using it specifically to predict monthly demand for a vast portfolio of products. And one of the things that they want to do is streamline their forecasting pipelines, enhance accuracy, and reduce operational complexity. And again, what they have seen is that they have seen that using the forecasting model like TimeGen from Nixla has helped them reduce errors by nearly 30% on average in forecasting errors, which is huge.
And again, in all these customer use cases, the time it took for them to start using LLMs in their applications to see the results and impact has been reduced significantly because they've been able to use model catalog and Azure AI Studio, where we provide, as Shubhi showed in the very first slide, we have tools for prototyping, optimizing, and operationalizing.
So whether you're just starting off with, let me try this for a prototype project, to realizing, OK, I need to put it into production, the time it takes from going from that to the last step has been reduced significantly, mostly because we have streamlined all the different foundational models.
We have provided theright tools for all our customers to kind of go through that whole LLM life cycle. I think that's pretty much it. Thank you.
Thank you.





