AIAI EngineerFeb 6, 2025· 28:56

Multi model multimodal and multi agent innovations in Azure AI: Cedric Vidal

Cedric Vidal, Principal AI Advocate at Microsoft, demonstrates Azure AI's multi-model, multimodal, and multi-agent capabilities in a session packed with live demos. He shows how GPT-4 Omni mixes text and vision to read handwritten French menus and translate them, and diagnoses infrastructure damage from photos for energy and insurance industries. A new video translation service lets him speak German, Spanish, Italian, and Japanese in his own voice while preserving tone, such as whispering or yelling. The Azure AI model catalog now offers 1,600 models, including serverless deployment options, and Phi-3 Vision, a small 3.8B parameter model running locally in a browser via WebGPU. He also demonstrates code interpreter analyzing a GPX file from a kitesurfing session, plotting a map with turn markers, and GitHub Workspaces preview generating a Java GUI from Python code.

  1. 0:00Intro
  2. 3:41Multi-modal
  3. 11:36Speech Translation
  4. 13:35Model Catalog
  5. 15:58Phi-3
  6. 19:40RAG
  7. 22:49Evaluation
  8. 23:58Code Interpreter
  9. 27:23GitHub Workspaces

Powered by PodHood

Transcript

Intro0:00

Cedric Vidal0:14

Uh, so, I'm Cedric Vidal. I'm an Principal AI Advocate at Microsoft, and today I'm going to— we're going to do quite cool stuff. We're going to do— to talk about multi, many things: multi-models, multi-modality, multilingual, multi-agents, all of this with Azure AI.

So, yeah. And also, one particularity is, apart from a few slides at the beginning, it's only demos. And to be honest, bear with me in case one of them or all of them don't work. But it's going to be fun.

We'll see how it goes.

So, as you know, you know, Azure AI is the best AI platform out there. We have a lot of AI services. We can do machine learning, and we also do all of that responsibly with the whole Responsible AI framework.

And we encapsulate all of this in the Azure AI Studio. And I'm going to do a lot of demos of Azure AI Studio today.

And

since now almost a year, a bit more than a year, we've been partnering with OpenAI, of course, and we have all of the OpenAI models available on Azure, on the Azure platform. But we're going to see that in addition to all the OpenAI models that we have and all the modalities that we can get using those, we also have many more models available on the platform for text, of course, vision, and speech.

And many organizations trust us today to use AI and build their products.

So, without further ado, I'm going to jump into the demos very quickly. But before I do, we've had many things announced at Build a couple of months ago. We've had the GA version of Azure AI Studio. We've had the latest model from OpenAI, GPT-4 Omni, which supports text, vision, and soon speech.

We've had the new small language model from Microsoft, from MS Research, called Phi-3. Now we have also announced GPT-4 Turbo Vision, Dall-E 3, and Whisper. We've announced the Assistant API that allows us to build agents, and I'm going to demo it.

We've announced fine-tuning for GPT-4, the new inference batch API. And also, another very cool thing I'm going to demo today, and you had a glimpse of it. I mean, I guess the surprise is kind of out, but the video translation service that I'm going to demo.

And Azure AI Studio. Oh, yeah. So, let's go straight to the demos now. So, apart from those slides, now it's only demos. So, the fun begins.

Okay, the first demo. So, we've had

Multi-modal3:41

Cedric Vidal3:41

Azure, I mean, like a year ago when everything started, you didn't have much, many choices. It was basically GPT or GPT. The only modality available was text. But now things have changed dramatically. Now we support also multi-modal vision, mixing text and vision.

And this opens a completely new era of use cases. For example, here, and let me zoom. So, I'm going to demo GPT-4. Oh, and actually, I selected vision here, but what I wanted to select was GPT-4. Oh. And I'm going to demonstrate a use case where, so it is kind of smallright now, but this is a menu from a restaurant.

And I'm going to ask, what's vegan on the menu today?

So, here we can see the menu a bit better. So, as you can see, we have winter chicory salad, duck, sea bass, et cetera. And what's very interesting here is that the menu that you just saw was, the font is funny, but it's printed.

So, it's a font from a computer. And GPT-4. Oh, it does a very good job at reading what's on the menu. And let me zoom here so that we can see what's. So, I asked whether there were vegan options today on the menu.

And what's interesting is that it looks, it mixes vision. So, it extracted all the text from the image. But not only that, but it resolves on it. So, it analyzed all the items on the menu today, and for each one of them, looked at which ones were vegan.

And as you can see here, the cauliflower soup and winter chicory salad. Huh, let me zoom. Ah. Up, okay. So, okay. Both mention vegetarian and vegan versions.

So, that's a very good example of how to mix text and reasoning, something that was not possible before with just OCRs, which becomes available with the new generation of multi-modal models. Another example, slightly harder, because this one

has

handwritten text. So, this menu has not been printed. It has been written by hand on a chalkboard.

So, as you can see,

oh, and it's in French. So, not only does it recognize handwritten sentences, written on a chalkboard in a picture, but it also translates it and reasons on it. So, that's three things that the model is doing all at once, thanks to multi-modality.

This is very important to understand how it differs from what we were doing before, because before, we were using image to text to extract the text and then reason on it. Now the model understands natively both pixels and text.

And in its internal representation, it has the same vectors for the same concepts, visual concepts and textual concepts. That's a very important thing to understand. And as you can see here,

it displays the answer. I mean, I asked what's on the menu today. So, it's displaying the entries of the menu in French with the English translation, because I asked the question in French. And I could also, oh, yeah.

I asked, okay, that's funny, because I asked what's good on the menu today. And so, the choice of what's good would depend on your personal taste preferences.

And the menu offers a variety of traditional French dishes that could cater anyway today. So, yeah. So, let's move on now to the next demo. So, we looked at, you know, something that you might want to do at the restaurant when you are in a foreign country and you don't understand what's in the menu, and you want to get a better understanding if you have a special diet.

So, that's very convenient. But that technology can also be used for more serious challenges or use cases. So, in this case, we're going to look, and that's actually an actual use case from a discussion I had a couple of weeks ago with a customer working in the energy industry.

And so, here we have

a picture of electric poles that fell on the ground. And I can ask a very open question. What's going on here?

By the way, you can see how fast the model replies, which is quite something. So, not only is GPT-4. Oh, understanding both images and text, but it's also much faster at answering. So, as you can see here, the image shows several power lines and utility poles that have fallen or are leaning, indicating damage to the infrastructure.

I'm not going to read everything, but what matters here is if I was working in the energy transport industry, I might want to observe continuously all the infrastructure of all the networks, like for a whole country, at the edge to make sure that the network is operational.

So, I might want to automate looking at all the video cameras of filming the infrastructure. So, I could ask, is the electricity working here?

It is highly unlikely that the electricity is working in the area shown in the image. Of course, here I ask the question in natural language, and the answer is presented to me in natural language too. But I could also ask for the output to be generated in JSON in a format that could be interpreted by code so that I can automate like dashboards and monitoring of infrastructure in real time.

Another use case is for the insurance industry. So, here we have a house that we can ask, what happened?

The image

shows a house that has collapsed and is severely tilted. Uh-huh. Natural disasters such as hurricanes, earthquakes, or landslides. So, yeah, that's also a very interesting use case for the insurance industry.

Speech Translation11:36

Cedric Vidal11:36

Next. So, the next one is not going to be a surprise because there was a spoiler. But, so, we talked about

multi-modal models. But now we're going to talk about another modality, speech. The Azure AI team, product team, has released an amazing new feature, which allows you to translate videos. So, here I'm going to play that video. So, disclaimer, that's me on the video.

This is our new video translation service. With this, I can translate videos into other languages in my own voice. Jetzt kann ich Deutsch sprechen, wie ich es immer wollte. Me hubiera gustado saber hablar español, pero ahora puedo hablarlo sin haber aprendido el idioma.

Posso anche sussurrare in italiano. そして 、 日本語で大きな声で話してください 。 This will make the world more inclusive.

So, what's really impressive about that video is not only the fact that now I can speak German, but

it took into consideration the intonation of what I was saying. So, when I was whispering, it was whispering too. When I was yelling, it was yelling too. So, it takes into account the language and the tone. A disclaimer, I stitched the different videos together myself using post-processing.

But apart from that, I didn't do anything. Like, the service did that all by itself.

Now, yes.

Guest13:22

So, as I understand, these are native models. Do these exist as embedding models as well?

Cedric Vidal13:30

Okay, I'm going to talk about embedding models in a minute.

Model Catalog13:35

Cedric Vidal13:35

So, where was I? Model catalog. And thank you. That's a good segue, actually, because, so, like I was saying, like a year ago when GPT-4, ChatGPT was released, basically you had almost no choice. Now, the amount of models available on the Azure AI model catalog is extraordinary.

So, here, I'm going to remove that filter here so that we display all of them. So, as you can see here,

we have 1,600 models available in the model catalogright now. And what I like, there is one specific feature that I really like is deployment options. Here, you can select serverless API. So, you have two ways to deploy models on Azure AI at the moment.

You can deploy them serverless, or you can deploy them using your own infrastructure. Bring your own infrastructure means basically that you rent for GPUs, that you pay for GPUs, whether you use the endpoint or not. Serverless means that you pay for the token and that the infrastructure is managed for you by the vendor.

And paying by the token is nothing new. You've been doing that with OpenAI's GPT ever since it was released. But now, you can do it for many vendors on the marketplace. And as you can see here, those are all the vendors and all the models that are available on the catalogright now, serverless.

So, you pay by the token and you have nothing to manage yourself. Not to mention the fact that getting GPUsright now is not the easiest. So, being able to use those models serverless makes it much easier.

And because we have so many models now, it's kind of hard to know which one to use. So, now we also have the model benchmarks where we compare, not all, but we compare many of the models that are available in the catalog.

And you can look at the accuracy as well as a bunch of other metrics to figure out which model you want to use for your use case.

Phi-315:58

Cedric Vidal15:58

One, okay. Let's hope that the Wi-Fi is not dead. Okay. One of those models that I want to focus on today, because we talked about GPT-4. Oh, which is a very big model, able to do text and visual analysis.

But Microsoft Research also came up with our own text and visual multi-modal model called Phi3 Vision 128K. This model is very, very interesting. It's a family of models. We have a vision version. We have many sizes. That one specifically is very interesting because

I'm going to upload one of the use cases, for example, randomly, the same use case as the one I talked about before with GPT-4. Oh, so, is

electricity working here?

And so, what's really interesting here

is that the model, despite being much smaller, and the explanation is a bit simpler, but still, Phi3 Vision is able to analyze the image and give a very good answer about the fact that the electricity most likely is not working here because of an outage or disruption.

So, now you have the possibility, in addition to GPT-4. Oh, to use Phi3 Vision for that.

And not only can you use Phi3 Vision as a service on Azure AI, we also have in the Phi3 family a very, very small model of 3.8 billion parameters, which weighs under roughly two gigabytes. And here, if I refresh the window here, and I hope I didn't make a mistake.

No, it's okay. So, model size two gigabytes. And here, Phi3, 3.8 gigabytes, quantized with four bits, is downloading in the browser and running locally on the edge using WebGPU, which is a new HTML5 specification allowing applications and browsers to get access to the GPU.

And so, here I can ask,

what do you know about the Rivian R2? And you know, this is going to be a segue to one of the next subjects I'm going to be talking about after.

And so, as you can see, you get an answer generated in the browser at, I mean, look at how fast it is for a model running locally. So, okay. I have a Mac, which has an Apple Silicon and some kind of GPU support, but it's by no means a horsepower machine.

It's a MacBook Pro M2, I believe. And as you can see, it's running very, very well in the browser. That means that I can disconnect the Wi-Fi. So, I'm not going to do it now, but you could disconnect the Wi-Fi and it would still run locally.

It's, which is also very interesting for,

RAG19:40

Cedric Vidal19:46

you know, to process sensitive data. Next thing, chat. So, here, if I ask the question, what are the different Rivian models?

So, I'm asking GPT-4 Turbo, which was last updated in April 2023. It knows about the Rivian R1T, the R1S.

Also, Amazon's Rivian Truck. And that's all. Now, I can select an index to do RAG, Retrieval-Augmented Generation, where I can ground my model into my own documents. And I'm not going to show that now, but what I did before to prepare the demo, I just took a bunch of Wikipedia pages of Rivian models that I uploaded to the index.

And now I am querying. So, if I ask again the same question, what are

the different Rivian models?

Huh, is it bugging? That's a problem with live demos. Let me clear. I'm going to copy that. I'm going to copy, clear, and rerun.

Okay. I'm going to refresh.

Okay, I have the index still selected. Now I should be able to, huh,

is it the model? Let's try with 35 Turbo.

Yeah. So, now, so apparently we have a bug with the other model, but with 35 Turbo, it works. And here you can see that we have the R1T, the R1S, the newly released R3, and we should have, it doesn't mention the R2,

but it does mention the R3. So, if I ask, what about the R2?

Yeah. So, it does know about it. So, this is very interesting because it allows us to use an LLM on up-to-date information. Now, let's move on to the next demo.

So, I'm going to skip that one. But what I can show here is one thing which is very important is evaluation of models. Because when you build an application, you want to be able to evaluate how good they are.

Evaluation22:49

Cedric Vidal23:02

And when you make modifications to the system prompt or you change your models on your application, you want to make sure that it continues to work as expected. So, inside Azure AI Studio, you have a feature called evaluation, which allows you to run a bunch of metrics.

Here, I show coherence, groundedness, and relevance. Groundedness is something that is very important for RAG applications because you want to make sure that the answer is grounded into the documents. So, this system allows you to do that moderately easily.

And as you can see here, so I use a very simple data set that I prepared for the demo. So, I have only one entry in my evaluation data set, but still, it shows that for that question, which is what are the different Rivian models, the coherence was four.

It was well grounded because it's evaluated between one and five.

Next, something very cool that I absolutely want to show you before we run out of time. So, here, this is one example of how to build an agent with code interpreter. Last Sunday, and that's real data from last Sunday, I went kitesurfing in the bay.

Code Interpreter23:58

Cedric Vidal24:17

And I did a pretty good session that I recorded using my watch. And I exported the GPX file from my session from my watch. And now I can ask a question, which I uploaded the file to code interpreter.

And now I can ask a question about it. So, I can say, hey, how long was I on the water?

And so, here, the LLM is going to automatically generate Python code and execute it in a sandbox to analyze the file that I uploaded, which is a GPX XML file. So, it's not a CSV. Usually, when you see those demos, they use CSVs,right?

But here, I'm using an XML file, which is more complicated. And you're going to see that it should work. Yeah, 42 minutes. That's roughly how much time I was on the water. And now I can ask, how many turns, how many attacks did I do?

So, attack in sailing is basically a turn. And here, this is a much harder question because asking how many attacks I did requires not only analyzing the GPS coordinates, but also analyzing the angular difference between each point and applying a threshold to decide which of the points on the path are actually turns.

And as you can see here, it replies with 217. And now we can ask, can you draw my session

on a map?

Because I want to see visually what it looks like,right? It's more fun.

So, why generate? What's very interesting here is that I have no expertise in GPX. Like, I don't know the file format. I could, like, if I show you what it looks like, it's pretty, you know, technical and it's pretty hard to parse.

So, here, without any explanation of what the file format is, code interpreter figured that out by itself. And here,

here's the map that it drew. And I'm going to skip. I could have asked, hey, can you please, and I did that to prepare the demo. Can you please add, draw red crosses for each one of my turns?

And here's the result. Can you imagine how powerful that is? Like, I didn't code a single line.

And 57, can I show you the last thing?

Well, real quick. Here's something pretty cool that I want to show you too. This is GitHub Workspaces. And here I can go to that repository and I can, so it's a preview, pre-access. And I can ask, can you add a Java GUI

GitHub Workspaces27:23

Cedric Vidal27:48

front-end?

So, this is a demo. I don't know if you noticed, but that was Python code. And I ask, can you generate a Java GUI front-end? And it's going to automatically, using an LLM, figure out what is the state of the code repository, figure out what it contains, what it does not.

So, it's going to write specifications. And based on the specification, we can ask to generate a plan. And every step of the way, if you make a mistake, we can ask it to make corrections.

This is preview. This is not yet available, but this is incredibly powerful. And that's upcoming.

And I'm over.