AIAI EngineerDec 6, 2025· 1:23:52

VoiceVision RAG - Integrating Visual Document Intelligence with Voice Response — Suman Debnath, AWS

Suman Debnath, a Principal ML Advocate at AWS, demonstrates how Colpali—a vision-based retrieval model that treats each document page as an image and generates multi-vector embeddings via patch-based late interaction—can be combined with voice synthesis for a more intuitive RAG system. He explains that Colpali bypasses traditional OCR and preprocessing by directly embedding document images, then scoring query-page relevance through dot-product similarity across patches. The workshop shows how to embed pages, store them in Qdrant with multivector support, and retrieve top pages. Debnath then wraps this retrieval pipeline using the Strands Agent framework, adding a speak tool to output answers in natural voice. A live demo answers a textbook question about trophic levels, first using Bedrock to generate text and then speaking the answer in a female voice, all without a predefined system prompt. The talk positions Colpali as a complementary technique for complex visual documents like IKEA instructions or scanned forms, not a full replacement for traditional RAG.

  1. 0:00Intro
  2. 7:49Multimodal RAG
  3. 15:25Colpali
  4. 31:11Late Interaction
  5. 35:36Demo
  6. 49:17Answer Generation
  7. 52:55Agentic RAG
  8. 1:04:31Voice Output
  9. 1:12:08Q&A
  10. 1:23:35Closing

Powered by PodHood

Transcript

Intro0:00

Suman Debnath0:16

Allright, so we are almost on time. Uh, firstly, thank you so much for your time, uh, for joining us. And what we are going to do is, for the next an hour or so, is, uh, we will try to explore something around which is which I found, uh, pretty interesting when I started working on this.

Uh, and I'll tell you some background about that, how I end up into this, uh, on Vision based Retrieval. Uh, but the idea of, uh, that I had was just to share a few of my learning on this particular approach of retrieval.

And there are a bunch of things that we have here. Uh, I'm going to share one of the latest research papers around retrieval, which is a, uh, Vision based Retrieval. And also, uh, I just thought to wrap this around with an agent.

Uh, without agent, we cannot talk about anything these days, so, uh,right. So, uh, it's funny. Without agent, I had this, but then, uh, the organizer said that, you know, we need to have some agent. I'm like, okay, that's not a big deal.

Right. So, yeah. So, allright. So, uh, we'll focus mostly on the, uh,

science side of this, like how that, uh, uh, Vision based Retrieval works. And then we will switch gears and wrap it around with an agent. I mean, that's a very simple task. And, uh, I'm going to use one of the open source, uh, uh, frameworks that, uh, uh, we launched recently.

I think it was two weeks back, called Strands Agent, uh, which is kind of a, a framework, lightweight framework, uh, to build agentic applications. I'll talk about that a little later. And I have a session on that tomorrow, uh, day after tomorrow.

Um, but that's the premise. And, uh, before we get started, uh, how many of you are, are from, uh,

science side of things? Like, who how many of you have worked on transformers? Okay, perfect. How many of you have worked on RAG in general? Fantastic. Okay. And, uh, how many of you have worked on AWS? Okay, great.

So there's nothing about AWS here. Okay. So, uh, so the last question was sponsored by my manager. Okay. So what so what we are going to do is, uh, uh, we are going to, uh, share one, uh, notebook.

You can just, uh, clone that repository. And, uh, there is a lot more, uh, there inside that. Uh, but we are just going to use one part of that, uh, repository. Okay. And I'm going to share a few of the, I think, there's some $25 credit code, uh, which I was given.

So I think, uh, you may like to use that. So, uh, let's get this logistics, uh, sorted, uh, first. Okay. So first thing first, uh, can we just switch the screen, please?

Yeah. Uh, can you just, uh, uh, take a moment and see that if the URL is working? Uh, you may if you are on laptop, you may like to, uh,

open the URL, or you can just take an image, uh, on your cell. You can have a look later on. Is it working?

Guest3:55

Yeah.

Suman Debnath3:56

Okay, perfect. Okay. So now, this is something, uh, you can take an image now, or you can, uh, do that survey later on. I mean, I don't like this, but again, this was given by my manager, so. So, so this is just it might ask you a few questions.

I have no idea what question they will ask, but, uh, you will get some $25 credit. And, uh, if you don't want to do it, don't do it. I'll give you $25 credit. I have it. So,

okay. So, yeah. And, uh, I don't know why this slide is next, but so this is my oh, I, I actually forgot to introduce myself. So I work with AWS, uh, as a principal, uh, machine learning advocate. I'm with this company for the last six months.

I focus mostly on, uh, natural language and, uh, RAG and fine-tuning. And if you have any questions around the talk that we are going to discuss, uh, or anything around, uh, machine learning or generative AI, feel free to, uh, ping me.

It's not just about this session, but, uh, my takeaway at whenever I go and speak at any conference at this scale is, uh, just to make, uh, a few, few connections with whom I can work with, uh, you know, after this conference.

Uh, because as long as learning is concerned, we can learn everything at home,right? And so you don't have to come to a conference. Uh, so feel free to, uh, connect. So with that, I will just switch to, uh, the GitHub repository.

Okay. So, and I'll just, uh, uh, walk you through the notebook. So my idea is not to have any presentation because, uh, first, uh, I'm lazy, and second, it's a little complicated. I thought that taking images and embedding it in the notebook is, uh, much easier.

So, uh, you will find this, uh, GitHub repository. Uh, and in that, there are, uh, uh, many things there. But what we are going to focus on is, if you come to this, uh, section eight and come to this first one, agentic voice based RAG.

Uh, so I just added that agent thing yesterday, so that's why. So I had no idea of that. So, uh, so what we are going to do is, uh, these two notebooks are exactly the same. One is without output.

One is with output. I find it, uh, useful to have both copies because if you are doing it for the first time, um, you may like to start with this. Don't see the output and run through. And at the same time, if you want to see what it is, you know, what is the expected output and all that, you can go into this.

Allright. So for the purpose of today's, uh, uh, workshop, I will start with an introduction, and then I will, uh, come here. Okay. And if you feel that, uh, this is not something that you are interested in, or this is not something that, uh, you are looking for, you know, feel free to, uh, uh, you know, uh, go to some other place because I don't want to waste your time.

Uh, uh, but I want to make sure that if you are here for the next one hour, you learn something new, uh, with respect to what you already know at this point in time. Okay. And if you have any questions, uh, feel free to ask.

And so that's, that's the other thing. Okay. Let me just expand this. It's a little too big, uh, here. So, uh, okay. So I noticed that most of you are aware of RAG, uh, but we are going to talk about, uh, multimodal RAG for a moment, uh, just to set the premise.

Multimodal RAG7:49

Suman Debnath7:49

And then we will, uh, get into the Vision based Retrieval. Okay. So if you think about, uh, multimodal, uh, RAG, uh,

what we essentially do is, and this is by no means is the only architecture. This is just one of the architectures. There are many different ways that you can do multimodal RAG. But in general, this is what, uh, we have been doing, and still today we do.

You take a data, and that data will contain images, text, and, uh, tables. The first thing that we do is we use some framework or is that so bad?

Okay. Thank you. So, so we can use any framework of our choice, or you can write your own custom script, or you can use any managed service like TextTrack OCR based technique. The idea is you extract the images, tables, and, uh, uh, text out, uh, you know, separately.

You can have some metadata, uh, to make an, uh, hash, uh, which tells you that this image is coming from which page and all that. But essentially, you divide all these three separately. Then you can use one multimodal embedding model.

And this multimodal embedding model can take any of these three entities because it's multimodal. It can take any of these three. And when I say multimodal, you can think of it like input can be multimodal. Okay. And then it will generate some vectors.

For images, it will generate some vectors, tables, text, and all that. Then you go to the database, any vector database, and store all these embeddings. So what you are essentially storing here are the actual embeddings of text, tables, and images.

And then comes the retrieval part. When you ask a question, any raw question, like raw text, it goes through the same embedding model. Then it is first searched here. It will get some relevant chunk, which could be, again, image, text, or table.

And then you take those chunks along with your text and use a multimodal LLM. Why multimodal? Because your relevant chunk can be images, text, or table. Right. And then you get an answer. So this is one approach. The second approach is you do the same thing.

Like this part is common. After that, you used a model which will just generate a summary of all this separately. So it will use, uh, you can think of it like a summary of an image is nothing but, uh, image captioning.

Right. It will generate a summary of this image, summary of the table, summary of the text. Now all you have is the summary. That means it's all text. So now you can use any text based embedding model to generate the embedding of the summary.

And then you store this embeddings here. So what you are storing here are only the embeddings of the summary, not the actual data. Right. And then when the question comes, now we are talking about the option number two.

When the question comes, we do a semantic search with the database. And what we get as a chunk are some summary. Now that summary could be a summary of an image, table, or text. We don't know. Whatever it is.

But whatever we get, both of them are of text format. So that's why we can use a general text based LLM to generate the output. Okay. So that's, uh, option number two. The option number three is exactly the same as option number two, uh, with a slight change here.

When you store the summary, you also think of it like this. You have a hash here, let's say a dictionary, which says that, uh, this image number one, uh, summary is this. Image number two, summary is this. Table number one, summary is this.

You create a hash file or hash, uh, uh, any data structure of your choice so that you can, you know, come back later on, uh, from a certain summary, and you can figure out this is summary of what entity.

Okay. But you store only the summary, just like before. But the difference here, uh, with respect to option number two is when you ask a question, you get some relevant chunk, which is a summary. Then you go back to that hash and find out the actual data, not the summary, the actual data which is mapped against those summary.

And then you take those actual data and then pass it on here. So what you're doing is summary you are using here just to reduce the search space, just for semantic search. Once you get that, uh, relevant chunk, you don't care about that summary.

You just take the original data from that hash. And then you take those relevant chunks and your question. And since relevant chunk can be, again, text, image, or, uh, uh, table, you need a multimodal LLM. Okay. And you generate the answer.

Are you with oh, yeah.

Guest12:47

Uh, let's say you have a table, but the table is an image.

Suman Debnath12:51

Mm-hmm.

Guest12:51

Which one would you prefer to use in that case?

Suman Debnath12:54

Yeah. So that's a good question. So the question is, if you have a table which is an image,right? So now it's all whatever you are saying is all at this level. So you don't you it's all up to you how you segregate these three entities.

So let's say you use an OCR based, uh, uh, uh, technique, and let's say it, it's identified a table as an image. So it will be treated as an image because the model till this point, the model has no idea from where these three are coming because these are like prerequisites for this particular pipeline.

Okay. So are you with me with all these three approach? Yes? Yeah.

Guest13:35

Uh, what is the brand name on, like, the website?

Suman Debnath13:39

This one?

Guest13:40

Uh, no, like the brand symbols on the website.

Suman Debnath13:42

Oh, okay. So these are nothing but, uh, uh, models, basically.

Guest13:46

Oh.

Suman Debnath13:47

It, it doesn't resonate. It's what I understand now. But.

Guest13:49

So it's not a language model.

Suman Debnath13:51

Yeah, yeah, yeah. Right. That's correct. Yes. So it could be this is any so here, we have a multimodal embedding model. Right. Here, it's actually first, you need a model which will generate the summary. Then you can think of it like another model which will generate the embeddings.

So, uh, it's just, uh, one icon, but think of it like there are two things happening, uh, in sequence. Okay. Are you with me so far? Yes? Okay. So do you see any problem in this, uh, when you have a multimodal data?

There is no problem as such, but, uh, there are few scenarios when this may not work. Okay. So scenarios like, like, uh, what you mentioned, uh, there are few documents or data which we have seen where the PDF was created using images.

So basically, you can think of like this. Um, let's say, uh, uh, toll that we, uh, cross in a highway. All are images,right? They just take the images of our, uh, number plate and all that. Uh, similarly, you can think of, uh, any government organization where forms are just, uh, they, they just keep on, uh, taking images.

And later on, all those images are converted into a PDF. So in that case, not always, uh, these techniques of extracting images, tables, and text works nicely. It's not like, uh, it never works. It's all about how your data behaves with your technique that you are implementing.

Colpali15:25

Suman Debnath15:25

Okay. So now the next technique that we are going to, uh, discuss today is using a Vision based, uh, Retrieval Model. And we will see that why we are using this. But the premise is this. If you use if, if, if your data with your data, any of these, uh, three options works, you just go with this.

What we are going to discuss in the next one hour, you it's not relevant for you. But this, uh, you know, what we are going to discuss is an option number four, which is, uh, a smarter technique which is based on a Vision based model to, uh, perform the retrieval.

You don't have to extract all these three entities in the first place because think of it like this. The moment you have your data and you first in the first place, you segregate these three things. It's just like you have a family.

You just, uh, you know, let your kid go somewhere. You go somewhere, and, you know, your partner goes somewhere else. It's a good thing, but, uh, you know, uh, if all of them goes, uh, you know, separate, and you expect somebody else to identify that they all are part of one family, it's a it's a task,right, for that external person.

Uh, so that is what we are going to solve, that can we can we come up with a technique where we don't do all this? Okay. So before we go, uh, I think, uh, you have a question.

Guest16:46

Yeah. I think you kind of answered my question.

Suman Debnath16:48

Yeah.

Guest16:48

Because you were explaining the case about, uh, scanning all the PDFs.

Suman Debnath16:52

Mm-hmm.

Guest16:53

And it wouldn't quite work. And I was a little bit confused as to why these approaches wouldn't work.

Suman Debnath16:59

Yeah.

Guest17:00

But then I think you are going towards the notion that we need to establish relationships between.

Suman Debnath17:04

Exactly.

Guest17:05

Things on the on the PDF.

Suman Debnath17:06

Absolutely. Yeah. So I think, uh, I'll give you one more, uh, one more on, uh, example. Uh, if you go to IKEA, you buy something from IKEA. If you have seen the IKEA, uh, ins, uh, you know, instructions, uh, you know, we don't I personally never looked into the instruction.

But while I was reading that research paper, they said that refer to that, uh, uh, because we generally go to YouTube and search what are the instruction steps and all that. Right. But if you look at the IKEA instruction set, they just have emoji kind of, uh, human, uh, and, uh, they are just assembling something.

There is no text there. There is nothing there. So unless and until you have a visual understanding of, uh, what it is, you will not have any idea what they are talking about. Okay. So there are some data sets, and I will show you, uh, a few of the data sets, uh, where the, uh, there are some text embedded within the image, and there are just an image.

They don't have any text. So you need some model or some, uh, technique which can help us to understand, uh, what is the semantics of the data. Okay. So let's see how, uh, we are going to solve this.

So this is, uh, again, the text might be small. Uh, you can just leave that. Okay. You can open it in your laptop, uh, or, uh, I'll try to explain it as much as I can. So this is the traditional technique that we discussed.

Right. You first place, you, uh, divide all these three entities, uh, uh, separately. Uh, but this is not very helpful because if you look, uh, you know, think about it, let's say you were given a book, and, uh, you were asked to answer a particular question.

Let's say I, I give you this book. By the way, this is a fantastic book from Simon. Have you heard of this book? Yeah. If you are getting started with machine learning, uh, deep learning, uh, you know, you can, uh, read this book.

This is a really fantastic book. It's a recently, uh, published book. And Professor Simon is very reachable. So it's, it's, uh, it's a fantastic book. So let's say if I give you this book, and if I ask you some question, and let's say you are not aware of this particular topic, you will not, uh, go and scan the entire book.

What you will do is you will try to first find the structure of the book. Maybe you will find the index. Where is the index? Where is the appendix? And all that. And then you will try to, uh, figure out, uh, which chapter this book might, uh, this question, uh, can be answered from.

And then you will go to that specific chapter and then read through those chapters. Right.

Guest19:36

What would.

Suman Debnath19:37

A human will do. That is exactly the philosophy of, uh, Colpali. Okay. So

when you get a question, what you do is first, you will first scan through the appendix and all that, and then you will figure out where exactly, uh, uh, the portion where your question can be answered. And then you will accumulate all those relevant chunks or relevant information, and then, uh, finally, you will come up with a response.

Right. So this is where, uh, uh, or rather, this was the motivation of this Vision based Retrieval Model called Colpali. Have you heard of this model, Colpali? Yeah? Okay. Yeah. Few of you. So Colpali was introduced, I think, in, uh, July 2024, just, uh, less than a year back.

Uh, and the motivation is we will treat every page as an image. So assume that you have a PDF document of, let's say, 100 pages. Your data set is not one PDF, but 100 images. Okay. There is no concept of retrieving, uh, images, text, and tables from there.

So how it works. So it first creates patches of every page. So now let's consider one page. One page is nothing but one image. And the same will apply on all the pages that you have in your document.

The first thing that it will do is it will create some patches. Uh, in the paper, I think it was, uh, the model was trained with, uh, 32 patches, like 32, cross 32. Here, how many patches we have?

We have 1, 2, 3, 4, 5, uh, 15 patches. Right. Now what we do is after that, once you have those patches, you will use this Colpali model, embedding model, and it will generate one vector per patch. Okay.

So in this document, how many vectors we will have?

Guest21:52

15.

Suman Debnath21:53

15. So now if I if my document is having 10 pages, how many vectors I will have in total?

Guest21:59

150.

Suman Debnath22:00

150. Okay. So now what we are going to do is we are going to see this middle part, how it generates the embedding, and then at later, uh, and in last section, we will see that how it does the retrieval.

And then we will go to the code. Okay. Okay. So before we get into that, um, uh, this embedding process, let's take a, uh, detour of, uh, Vision based Language Model. Uh, have you worked on Vision based Language Model, any of you?

Okay. Okay. Few of you. So ultimately, if you think about it, uh, we had, uh, language models, uh, I'm not talking about, uh, uh, large language model, but text based, uh, models, uh, since we had Transformer based architecture.

Right. And then at that time, we also had, uh, models which can work pretty well with images which are based on, uh, CNNs. Now what researchers thought was, uh, now that we have a language model, why can't we just, uh, make use of that to work with, uh, vision?

It could be images or videos. Videos are nothing but, uh, images with but with a timestamp. Right. Another dimension you can think of it. Uh, and then

what people did was they took some Vision based model, and then they took some, uh, text based model. They both are separate. Right. At, at this point, what we are talking about is, uh, before the training. Right. Like the these two models are completely separate.

They all have, uh, you know, they, they, they are in different space, basically. Right. And the idea is at the end of the day, come up with a model where if you send an image of a dog, the vector that you will get, and if you send a text about dog, the vector that you will get at the end, those two vectors will be very close to each other.

Initially, it will not be close because when, when I say a dog is sitting on a field, and if I use any text based model, it will generate a vector. And for the sake of simplicity, let's assume the vector dimension is 10.

So it will generate a array of 10 numbers. Similarly, when you pass an image of a dog, it will generate an im a vector, final vector with 10 numbers. Let's say the embedding vector size is 10. Now those 10 numbers and these 10 numbers for the text, they will be anywhere in the space because they don't have any correlation at the before, uh, the training.

Now what happens is at the time of training, we take a lot of samples, positive samples where the text is there, uh, text is there which replicates the image. And there are a lot of, uh, negative samples where image is there, but text is something random.

And the idea is the loss function that we use is if they are similar, we want to make sure that the loss is, uh, uh, less. But if they areorthogonal or very separate, we will say that, okay, the loss is high.

And during this loss, you know, this training process, we kind of optimize. And at the end, we, we see that when you send an image or a text, the embedding that we get at the end are very close to each other.

Okay. So we are not going to deep dive into, uh, Vision based model, but there is something called, uh, contrastive learning where if you send an image and a relevant, uh, positive tag, and if the vectors are very, uh, uh, very sparse, like very, uh, uh, uh, very much apart from each other, then the loss will be very high because we want these two vectors to be close.

Right. So that way, uh, during the ba uh, back, uh, uh, back propagation, we update the weights accordingly. So this is one of the reason if you think about it, uh, you might have seen in language models or when you use any, uh, let's say any, um, foundational model, they say that always your prompt should be, uh, about what you want, not about what you don't want.

Have you seen this if you are into prompt engineering? Why they say this? Just think about it. Let's say if I say, uh, uh, okay, le-let, let me give you an, uh, analogy. Right. If you are going for a dinner,right, with your wife, and if you ask your wife, what would you like to have?

She will say that I, I, I let's say, uh, I don't like this. I don't like that. But that was not my question. My question was what you want. That is always difficult to, uh, answer. Right. People will say that, uh, uh, okay, would you like to have this?

No, I don't like this. But when you ask that, okay, you tell me what you like, it's very hard. So that's why when you give a prompt that, uh, I want a dog sitting on this chair, it's a very nice prompt.

But if you say that dog should not sit on the floor, it can generate any image because you are not saying that it should sit on the chair. It might be sitting on a desk or somewhere else. Right.

So that is the reason we always say that the prompt should be, uh, very much positive, what you want, not the negative because, uh, you know, data set doesn't have that much, uh, of negative samples. Now let's come back to Colpali, how it works.

So you give an image. So now this image is, uh, you can think of it like, uh, one of the patch. Okay. It goes through, uh, the, uh, Vision based encoder, and it will generate an, uh, embedding. And then we have a linear projection.

And the reason that we have the linear projection is because at the end of the day, uh, when you ask a question, that will also be generating some vector. We want to make sure that these vectors are compatible to each other.

They are of same size. And that's why we have added a new projection layer. You can simply, uh, think of it as a fully connected layer. And ultimately, you will have a standard Transformer, and then you will get the output token.

Okay. So now if you think about

let me just scroll down. Yeah. So if you think about Colpali, when you give an image, it will have, let's say, in this case, there are 15, uh, patches. Just think of this patch. Okay. This patch will go through this.

It will generate an, uh, uh, vector, and this will be the final representation of the first patch. Similarly, when you give the whole image, you will not give the, uh, you know, uh, single, uh, patch, uh, in the batch.

You will give, give one full image or let's say page number one of the document. This model will do all that patching, and it will finally generate one embedding vector. Now at the time of, uh, and if you, if you see here, in this case, this is grayed out because now we are talking about after training.

Once that model is trained, after training, you will create the embeddings of your document. So while you are creating the embeddings, there is no question here. Right. So it will just use this path. Now once those embeddings are done, like once you get all these final vectors for your entire document, in the query time, you will just use your text based query.

So Colpali doesn't say that, uh, you can query with your as an image. Like in ChatGPT or any GPT based model, you just upload an image. We are so lazy. We don't even ask the question these days. Right.

We just upload the image and, you know, model just generates something. So here the question should be always in text. That's the prerequisite for this model. And then this goes through the same, uh, model, and then it finally gives you a response.

Now this response, this vector, you will do a semantic search with the vectors that you have stored in your vector database using those, uh, image patches. With me so far? Yes. Okay. So if you think about it, uh, both for query as well as, uh, your embedding, there is a certain amount of, uh, uh, uh, preprocessing that is needed because, uh, your images can be of different size.

Right. So let's say you have an, uh, a PDF document, uh, and, uh, the tool that you use to convert that into an image, uh, it created an image of 800 by 800. But let's say somebody else have used another, uh, technique, and the image was of 50 cross 50.

So we need to make sure that the images are of standard size. Right. So that's why when we look into the code next, you will see that always before it actually generates the embedding, there is a preprocessing, uh, that we do.

Late Interaction31:11

Suman Debnath31:11

Okay. So let's, uh, go to the code and see that. But before that, let's, uh, let's share, I mean, uh, let's talk about how it generates, uh, the similar chunks. So this is the most important part of Colpali.

Okay. Now imagine that your page now just consider page number one of your document. And the page and the patch size that you use is, let's say, two cross two. That is total four patches. Okay. Now let's say this is page number one, and this is the embedding of your fir uh, fir, uh, first patch.

This is the embedding of the second patch. This is the embedding of the third patch. This is the embedding of the fourth patch. Okay.

And you ask some question. Let's say, uh, what is AI? Just for the sake of simplicity, what is AI? And you have used through, uh, it went through the tokenizer, and it generated three embedding vectors. Right. Three tokens basically.

So now what we do is we do a dot product between each vector and each vector of all the patches. Okay. And then for every row, we try to find which is the maximum number. What this number signifies, this 89 signifies, 0.89 signifies that the first part of your question has the maximum similarity with the second patch of the image.

Right. Similarly, if this is 97, that means the second part of your question has the maximum similarity with the third patch of your image. Right. And at the end, what we do is we just take the addition, I mean, we just take a sum the maximum numbers of each rows.

And if let's say that is 2.58, that means this query has a score of 2.58 for page number one. Similarly, we will do it for all the pages. And then in RAG, what we do at the end when we do a semantic search, we say top five chunks or top 10 chunks.

So in this case, chunk is nothing but pages. So if I say top five, then in that case, it will show us the top five pages based on this score. Getting it? So this is the most important thing.

So this is called late interaction. Have you heard of late interaction embeddings, all that? And the reason that we say late interaction is because these token embeddings are already stored. We have already done that. It's there in your vector database.

All we have to do is we need to just do the dot product and then use this matrix to generate the top five or top three, uh, pages. Okay. Uh, with me so far? Yes. Okay. Now this functionality is not supported in all the, uh, all the vector databases.

We are going to use one of the vector data database called Quadrant. Have you heard of that? But there are a few other databases. I have not done enough research which are the databases that it supports. Uh, but, uh, this, uh, maxim calculation is not supported by all the database.

Okay. There are some open source contribution that we have for a few of the vector databases. I, I, I, I tried with OpenSearch. Uh, it did not have, but I think there is a, uh, extension, uh, which you can use to make this functionality.

Okay. So now we are going to get into, uh, the demo. So just like what I said, uh, once you have those, uh, scores, like in this case, 2.58, uh, like this, you will have for all the pages in your document.

And then at the end, uh, you can pick the top three or top four pages of your choice. Okay. So now see this. So far we are not talking about agents. Okay. Because that's a very simple task. Uh, we will just wrap this with an agent at later point in time.

Allright. So let's try to do this. Okay. So this is an, uh, uh, I'll just come to this, uh, image later on. So now let me just increase, uh, the font. Can you see this? Yeah. Okay. You don't have to read all that, but just, uh, you should have an idea what we are doing.

So first, we are just importing a few of the libraries. Uh, I have no idea what I'm importing, but, uh, there are a few. It's, I think it's, uh, it's a Colpali. Yeah. So this is the Colpali model that we have.

Demo35:36

Suman Debnath35:46

Okay. And this is the Quadrant database. And this Quadrant database, we are going to run locally in a Docker container. Okay. So if you are planning to run this, uh, make sure that you have Docker installed in your, uh, in your laptop.

Okay. I, I, I think the README have all the information. Okay. Uh, so first we need some data. So I have used one data set. Uh, I'm basically it's a small textbook. And if you see this textbook, uh, this is a science textbook, uh, chapter number 13.

So we have, let's say, see, one of the thing that, uh, which is interesting here is if you see this image, there is no text here. Right. So it's, uh, if you ask anything about this image, uh, and use a traditional technique, it might not answer properly.

Uh, like this. This is also another image along with some text. And, uh, you know, you can pick any data set of your choice, but, uh, this is the data set that I have. Okay. And, uh, feel free to use any data set of your choice.

But, uh, for the purpose of this, uh, demo, you may like to download one of these, uh, PDF from this URL and, uh, play around with this. And then you need to have a Hugging Face, uh, uh, uh, token, uh, because we are going to download this model from Hugging Face.

Right. So you should not do this. Right. So this is, uh, you know, I was just trying this because without creating a .env file, but you should have an env file inside that your token should exit. Exist. Okay.

So this is, uh, a token. Like this is not my token. If you see this, uh, this is just a dummy one. Right. This is not my token. Okay. So it's, but, uh, this is, this is just, uh, the, the Hugging Face token that you should have.

So here we are just loading and logging into our Hugging Face account. And next, we are trying to check whether we have a CPU, GPU, or, uh, MPS. In this case, it's a MacBook. So I'm just using MPS here, uh, as a device.

Since it's a Vision based model, it's better to run it on a GPU. It will be faster, but you can very well run it on CPU. That's fine. I'll tell you, uh, uh, you know, you should be a little cautious about this if you're running with a, within your laptop, uh, on CPU.

Uh, if it's an office laptop, no one cares. But if it's your personal laptop, make sure that the batch size is very small. Otherwise, it will, it will crash. In fact, I, when I first ran this, I did not check, uh, the processing time and all that.

It just, uh, went on, uh, you know, crashing, and it, uh, rebooted my laptop. Uh, and, uh, I, I did not even read through all this, and I raised a IT ticket, and I actually got a laptop, new laptop.

Uh, so, but it was my fault here. So they thought that my work, my work needs a laptop with more memory. So, so if you are finding out tricks to get a new laptop from a company, so this is the cell.

Okay. You can try that. I'll tell you what you have to change to get a new laptop. Okay. So just increase the batch size to batch size to 12. It should work fine. Yeah. Yeah. Okay. So this is the model that we are going to use.

It's a Colpali, uh, version 1.3. There might be a new version, but just have a look. I, I checked the last one. It was still 1.3. And I'm having a model and the preprocessor. Remember that we discussed that we need to have a preprocessor first.

Uh, we will process our data, and then we will use the model to generate the embeddings. Okay. The same model, but there is a preprocessor and the model. And these all are coming from Hugging Face. And we are using a cache directory so that we can load this model locally in our, uh, local directory, uh, so that every time you run, it doesn't download from the internet.

Okay. And once that is done, uh, you have to have a vector database. So if you have a Docker installed, you can just copy and paste it. It, it is nothing but it just created a, a container with a port forwarding, and, uh, there is a folder which gets created locally, uh, as a storage.

So all your vectors will be stored locally in your laptop. That's all. And if you click on this dashboard, you should be able to see, uh, that, uh, UI of that. And if you come to console, uh, sorry, um, here, collection, initially you, uh, since I've executed that code, that's why we see this.

But, uh, you should not see anything here. And as you run through the notebook, you will see the, uh, collection here. So collection is how many of you are aware of databases? Okay. Okay. Many of you. I have no idea about what it is, but, uh, I just asked.

So collection is basically you can think of it like a database and where you will just store all the, uh, schema and all that. So I'm creating a, a Quadrant client. And this is something that I imported, uh, earlier.

And this is the local host and port number this. And this just, uh, we are just creating the setup. Right. So now we have a vector database, and we have the data. So, and we have also, uh, downloaded the model.

So now the second thing that we need to do is we need to create a collection. Right. And if you see this, uh, we have a collection called, uh, class 10 science. Uh, so you can give any collection name.

Here we are mentioning what should be the vector size. Right. So, uh, this is the, uh, embedding length. So here it is 128. And in this, in this code, what we are essentially doing is if there is a collection already exists, it will not create any new collection, or else it will create a new connection.

Yeah. You have a question? Uh, yeah. Do you have, like, do you have to have, like, the best size to choose? Yeah. So there is, uh, let, let, let me ask you this. What do you feel if I increase the embedding size from 128 to 256?

What do you feel? How, how, how it would behave? Just a guess. It's got more content in each fragment. Mm-hmm. So, uh, the retrieval might have irrelevant information. Mm-hmm. But it might have some more as well. Okay. Okay.

Let, let me, let me give you an example. Okay. Let's say I, I, I, I have just come here, and I mentioned you two things about me. I work for Amazon, and I'm married. That's all. Okay. These are the two information that you have.

Now, if he asked you some question about me that, uh, okay, uh, tell me is Suman plays cricket or not, will you be able to give some answer? No. You will be giving some answer based on these two information, but it will be random.

Right. But now, let's say if I give you more information. I'm Suman. I work for Amazon. I'm married. I have one wife, uh, as of now. I, let's say I have one kid. Okay. And a few other things.

So if I keep on giving more features about me, you are having a richer, uh, information about me. So now if he asks you a question, it's more likely that you will be able to give a, uh, you know, you are, you will be able to give a more accurate answer.

Same with this. The moment you increase the embedding length, you are, it's not about chunk and all that. It's just about how much granular information that you are having about a specific thing. Okay. So you can always embed, uh, any entity with just one number, a vector of size one, but it will not have much of an information.

As you increase the length, it will, it will be more richer. Okay. Okay. So coming back to your question, uh, in the documentation, I think they, uh, they said that 120 is a good number, uh, but you can always use 256.

Right. If the vector database also supports that or the embedding model. Okay. So, so this is where we are just creating that collection. We have not even, uh, started creating the embeddings and all that. And see this here.

Uh, this is what, uh, I was referring to when I said Quadrant supports that matrix multiplication, that, uh, uh, late interaction thing. Right. So it says, uh, I'm setting some configuration that it's, it should have multivector configuration and multivector comparator as max sim.

So max sim is what, uh, helps us to get those three numbers from that matrix and then add those three numbers and give us the final value of your query and each page. So that at the end of the day, what do we want?

We want the relevant pages, uh, based on our question. Right. Okay. So now once this is done, it's now pretty simple. I have to, uh, first create the embedding. But before that, I need to create, convert my data into images.

And that's what, uh, this, uh, this function does. So what it does is it, it takes you, it takes an, uh, directory, and you can have hundreds of PDF files. And it will go through all the PDF files, and it will create, uh, images of each pages.

Okay. And not only that, it will also add all of that into a list called all images. And this is just for my own housekeeping with some metadata like document ID, page number, and the actual image in the form of RGB.

And it will store it in a local directory called PDF data. And if you just see the first two entries, you will see that, okay, this is document number zero. That is, let's say, I just have one PDF.

So all the entries will have document idea zero, page number zero, and this is the image. Page number one, and this is the image. Okay. So this data set contains everything with me so far. Yes. Okay. Great. Now that I have this, uh, uh, images, I can use the embedding model to generate the embedding.

And, and this is where, uh, you know, I just crashed my laptop. I initially used a batch size of 10, uh, 12. So it took a lot of memory, and I had just, uh, I think 16 gig of memory.

So it, it actually crashed. But, uh, if you're trying in your laptop, make sure that you start with two or three. So it basically means how many images you want to process. And now here we are generating the embeddings.

And first we are going through this Colpali preprocessor, which will just, uh, preprocess the image in a standard, uh, size. And then I'm passing it through the Colpali model, which actually generates the embedding. So this will have my embeddings.

And once I have all those, uh, embeddings, what I want to do is I want to store it in the vector database. And that is what I'm doing it here. I'm just inserting into the collection that I've created for all the points.

Each point is nothing but you can think of it like, uh, each, uh, vectors. Okay. And in this case, I have just 10 pages. So it will just generate the amount of number of embeddings for those 10 pages, and it will store it here.

Now is the final thing, you know, how we can retrieve. See this. I've just asked this question. What are the different, uh, tropical levels? Because this is there in the book. And this question also need to be, uh, need to go through that embedding model just like images.

So I will do, I'll make that through the preprocessor and the model. And once that is done, I will do a semantic search from the vector database. And that is what we are doing. We are just querying, uh, the vector database with our query token.

And I'm saying that the limit is five. What this limit five means? That means I need the top five pages, which is relevant to this, uh, question. And at the end, you will find some five pages. And if you want to see, uh, those five pages, how it, uh, uh, you know, how it looks like, uh, you can actually visualize.

So this is just a wrapper, uh, Python function, which will just take all the images, and it will just generate, uh, the images in a pictorial format. Okay. And in fact, if you see, uh, this image, I think this was, uh, this was the image, I guess.

Uh, yeah. So this is where the tropical levels are mentioned. It actually identified based on the question and the, uh, Colpali embeddings. Right. So this is the page. And also there are other pages which we got. Now comes, so retrieval is done.

Answer Generation49:17

Suman Debnath49:17

Right. So Colpali just talks about retrieval. Its, its job ends here. Okay. And, uh, if you, if you think about it, uh, with respect to, uh, sorry, with respect to this, uh, sorry, here, I guess, in the traditional technique,

we came to this point. Right. Uh, we came to this point. Sorry.

We came to this point when we got the retrieved images, and the question is already there. Now we can use any multimodal LLM to generate the answer. Right. But we have skipped everything here. Right. So now when we, when we use any generative model, you can use any generative model of your choice.

Uh, if you don't have any AWS account or if you have any other, uh, model access, you can always use that. Uh, but let's say, uh, you don't have Bedrock access. So we can use Ollama. Have you used Ollama?

Just a local model. The response may not be that great, but you can work it out. Right. So this is again a wrapper function just to, uh, convert all the images, uh, into the format that the model expects because we are, we need a multimodal LLM.

Right. So we will take some multimodal LLM from Ollama. But, uh, depending on what model you are using, the model will ask you to have the input in a certain, uh, format. Right. So it needs the data to be in base 64.

That's what this small tiny function does. Right. Uh, and then we just say Ollama generate, and this is the model that I'm using, and I'm sending the query and the image. That's all. Right. So see this. Now I'm sending the full query, not the embedding of the query because Ollama has nothing to do with that embedding of the query.

That embedding was needed just for semantic search. Right. And then, uh, we get some response. If you want to use Bedrock, then you should have Bedrock access. How many of you know about Bedrock? Okay. Perfect. So it's just a managed service on AWS through which you can access any different, I mean, different kinds of model.

And the way that Bedrock expects you to give the input, uh, multimodal input is a little different. And that's why we have some wrapper functions, uh, which will, which will make your prompt, uh, you know, according to the multimodal, uh, models requirement.

Right. And you can go through these two functions. It's standard, uh, uh, you know, uh, Converse API that we have used. So nothing fancy here. So I don't want to go there because that is not the purpose of this, uh, problem.

But ultimately, you give the images and the query, and you mention the model ID. So in this case, I'm using Sonet, uh, Claude Sonet 3.7. You can very well use, uh, Sonet 4 if you would like to. And you generate the image.

Uh, sorry, the final response. Okay. With me so far? Okay. Now comes the agent thing. How we can make this agentic. So it's very simple. Right. You don't have to go through all these things because what we have done is ultimately what we want when somebody is asking a question, we want an agent to, to retrieve the shortlisted images and give it to me.

That's all. Right. So we have seen how to shortlist those images. Right. What we are going to do is, uh, here,

Agentic RAG52:55

Suman Debnath53:07

um,

I'll just go through that later. Yeah. What we are going to do is we are going to create a function called retrieve from Quadrant, which will just take your query. And if you see the, uh, uh, return for this are the matched image paths.

That is what we want. Nothing else. Right. And the code here in this function are the exact same code which we have gone through in multiple cells, you know, uh, previously. It just, it does the same thing. And now to make it agentic, I have used a framework called Strands.

Have you heard of Strands? Right. Okay. So Strands is a new agentic framework. Let me just show you this. It's StrandsAgent.com. This is a, uh, SDK which was launched by AWS. Uh, I worked with, uh, uh, at the launch.

There are some YouTube video as well. Uh, you can just go over and just search for Strands Agents. You will find a launch blog as well. Okay. But basically, it is very, very simple. Just to give you an example how to get started with Strands, uh, you just pip install.

And, uh, do you want to see a quick demo of Strands before we go to that part? Will that help? Yes. Okay. So let me just show you. I think I have that. Um, okay. So let me quickly spend four minutes on that.

Four, five minutes. I have a good demo actually. If you want, how many of you have heard of, uh, three blue, one brown? Okay. Perfect. Okay. So then let me show you that. You might, you might, uh, it might be interesting.

So Strands is an, uh, framework. Very simple. It's a model first framework. So we are just taking a pause on that. Okay. We are just, we will, whatever we learn here, we will just use this framework to make our workflow, whatever we have done, agentic.

And we will add a voice part of that as well. So here there's an open source framework which is model first. That means, uh, now the models are so strong, we expect that the model should reason rather than we telling, uh, the agent with a lot of backstory, goals, prompting, and all that.

We don't want all of that. We throw a question. We expect the model to generate the response and do the reasoning, uh, on, on the model side. That's, that's why this is very, very lightweight. And it has an integration with, uh, different models.

And you can use model from Bedrock. You can use directly from Anthropic. You can use LightLLM. Have you heard of LightLLM? Yeah. So when you have an access to LightLLM, you can access any model that LightLLM supports. Right.

Now this is what it is. So Strands, by the definition of Strands, uh, it's a DNA structure and it just have two Strands. And that two Strands stands for model and tool. That's all. So you make an agent with one model, uh, uh, and few tools.

And simply you just ask the question. That's all. It's as simple as that. Right. And let me show you, uh, one quick demo. Let's see if it, uh, if it works. Okay.

So this is, is it visible from, you know, last row? No. Not that much. Right. Um, okay. I'll just read it out. So we are just importing, uh, agent and we are importing the tools. Okay. And okay. This is, I think this is the MCP one.

No, no. This is not the one I wanted to show you. Um,

okay. Let's, uh, see this.

Okay. Just, uh, uh, I think it's a video. It should work fine.

Yeah.

Let's see. So we first install Strands Agent and Strands Tool. Pip install. Simple pip install.

And it's open source. Okay. So you don't have to, uh, and it supports Ollama as well. Uh, so you don't have to have an AWS account or anything of that sort.

So what we are going to do is we are going to create a file, create a summary, write the summary into the, into our file, and, uh, also add a voice part of it. Let's see. I think it would be pretty quick.

So we are importing Strands. Is it visible or should I show it? It's pretty straightforward. So we are importing agent and we are importing the Bedrock model. By default, it uses Bedrock model. Uh, it actually uses Claude 3.7, but you can use any other model.

And I've used some built-in tool called read file, write file, and speak. And this is the model ID. And this is the prompt. You can have a prompt. You can skip a prompt. Doesn't matter. And lastly, you have to create the agent.

So

this agent contains the model ID, system prompt, and the tools. And all these tools, you have not written the code for this. This is by default. Right. And I'm just asking, uh, a particular question. And see this. I'm, uh, in the prompt, I'm saying that this is a textbook, uh, in my local directory.

Read that, create a summary, and write it into the local directory and also speak out the final answer. And see this. It is using the tools for read the file. Second, it is creating.

Guest59:20

You have functions like a camera.

Suman Debnath59:21

See, now it is speaking.

Guest59:22

It's like entering through the cornea and focus by a lens onto the.

Suman Debnath59:24

So we have not done anything. Just pip install.

Guest59:26

The eye controls pupil size to regulate incoming light.

Suman Debnath59:29

Okay.

Guest59:29

The eye can adjust focal length through accommodation. See.

Suman Debnath59:34

Allright. So now I, I will share one more thing. Now I'll not show you the code. That is not the purpose of this. But have you heard of, uh, uh, of course you have heard of MCP. Yeah. So see this.

I'll, I'll not tell you. I think you should be able to. I've created an MCP server called, uh, created an MCP server with Manim. Okay. So Manim, have you heard of Manim? Okay. Just, just see that. So idea is, so let me just show you what we are doing.

We are creating a Manim server and this is the MCP server. Now this is the client on nothing but, uh, our Strand agent. And this will call this MCP server. Okay. And I can give any question. So question is, create a Manim screen which draws a cubic function like 2x power 3 minus and blah, blah, blah.

Okay. And see what happens.

Now it is executing the code, calling this, uh, MCP server. It is working, uh, here. And then it should give you some response.

So it generated this video. Okay.

And now you will get some familiarity.

Looks similar. Right. I have not done anything. All I have used is a Manim, uh, uh, SDK and created that, uh, MCP server which can generate videos like, uh, what three blue, one brown have created. So this is just a small demo of how you can make use of Strands with an MCP and write simple code and, you know, do wonderful things.

Okay. Allright. So this is about Strand. The core idea of Strand is, uh, just pip install and use it with the by default tools and, uh, your model of your choice. That's all. There's nothing, uh, no scaffolding, uh, beyond this.

Okay. So it's just like this. You pip install, create an instance, and just ask question. Here we have not mentioned any model. That means it will by default use Bedrock model. But in the demo, we have seen that you can define your Bedrock models here.

Okay. So now let's come back to our, our, uh, our problem. In this case, our tool is not the default one, but the tool that we have defined. And what is that tool? The retrieval tool. And how I can create a, uh, custom tool is just by importing tool and just use that as a decorator on top of your function.

That's all. Now this becomes a tool for me. Just like read file, write file, speak. This is just a tool for me. Okay. And we can define, we can use, make use of Bedrock model or Ollama. Up to you.

And now look at this. We are also importing an image reader. Why we are importing this image reader, I will, uh, uh, tell you a little later. But, uh, you, you remember that when we use this Bedrock model for final answer, when we use this Bedrock model to generate the final answer, we created some custom functions which are nothing but, uh, contains the information about how to, uh, create the prompt for your images for Bedrock models.

Right. So I don't have to do all these things. And, uh, I can simply make use of this image reader which just takes an image and generates, uh, the, uh, prompt for us. And now I have a system prompt.

System prompt says that you are a RAG based system and all that. And it also says, uh, these are the two functions that you have, uh, or the tools that you have to use and all that. And that's all.

And now you create an agent. Again, just like before, you define the model, system prompt. And in this case, we use two tools. One is the retrieve from Quadrant, which is our tool, and the image reader for the generation part.

Okay. And then we ask this question. What is the difference, uh, different tropical levels? And now it just agents, uh, uh, generates the response. Just like before, but now everything is done by the agent. And the beauty is, let's say now you want to add the voice feature.

I don't only want the answer, but also the final response in the form of voice. So far I have done this. I'm just reducing the image so that everything fits in. So far we have done this. We ask a question.

It goes to a Strands agent. It uses this retrieval tool, custom tool that we have created. It gets the relevant chunk, which are nothing but the shortlisted pages. And then it uses any of these models, let's say Bedrock, Ollama, whatever, to generate the final response.

Voice Output1:04:31

Suman Debnath1:04:51

And to generate this, it uses this image reader tool. Now what we have to do is to add voice functionality. I will just use the speak, uh, tool. That's all. Just one, uh, import. Okay. And

that is what we are doing. We are just adding speak here. And again, the system prompt remains the same. And I'm querying the same thing. And now when I ask this question, so let's say, let's ask this question.

Okay. So let me run this.

And let me, let me just ask in the question itself. Explain the answer over a female voice in a natural way.

And let's see.

And now let's run this.

I hope I'm connected with the internet, but let's see.

So when you run this, uh, code in your environment, you can simply remove the system prompt. You will still get theright answer. In fact, try this prompt. I have not tried, but try this. Change this prompt and say that a male voice or something like that.

Right. Robotic, uh, uh, way of, uh, you know, not a natural way. Maybe robotic way. Something like that. The idea is see that whether Strands is able to, you know, forward that information to the model or not. Right.

So you don't need a system prompt. Uh, it may be because of my internet, uh, but it's, it doesn't take that much of time. You just give it a shot. It should work fine. Okay. So that's what I, uh, I had, uh, uh, for, for, uh, this particular, uh, workshop.

Uh, I would, okay, it's now running. So it's a little slow. But let me. So it is now able to generate the images. I mean, shortlisted the images. And now it should speak in a female voice.

So while that happens, okay.

Guest1:07:19

Tropic levels are the different feeding positions in a food chain representing the flow of energy through an ecosystem. They are typically for.

Suman Debnath1:07:27

Okay. So let me just stop this. And let's say if I, let's try this. Okay. Um, let me delete this system prompt. And let me just have this model and the tools. There is no speak, nothing. No system prompt.

There is nothing there. And here I will change this to a male voice. Okay.

And, uh, I'll give it a shot. Let this, I don't want to interrupt.

Guest1:07:58

Which are small carnivores that eat herbivores. These might include frogs, small birds, or foxes. The fourth trophic level is occupied by tertiary consumers or top carnivores.

Suman Debnath1:08:10

You can in fact say something like summarize in, uh, you know, 50 words or 100 words rather than waiting for this to complete. So.

Guest1:08:18

That energy transfer.

Suman Debnath1:08:19

It's still going on.

Just give it a second.

And before I forget, if you want to know about that multimodal, that the traditional technique, in this GitHub repo there is the part three. And here you will find the details of, uh, that architecture. Like, uh, uh, this architecture.

Right. And, you know, this notebook is about how you can do the same thing, but, uh, preprocessing this image text and table. Okay. So just play around this GitHub repo.

Okay. It's done. So now I will quickly, uh,

I just created this agent now, but, uh, without any, uh, system prompt. Now I just executed this. And now let's run this. So now I'm letting, uh, the agent know only about the models. Nothing else. There is no system prompt.

I just hope we get a male voice at least.

Guest1:09:49

Let me explain trophic levels, which are essentially the different feeding positions in a food chain or ecosystem. Think of them as the levels in nature's dining hierarchy. Starting at the base, we have the producers. These are main.

Suman Debnath1:10:03

Okay. Let's see if I have.

Voice.

Let's try this.

Guest1:10:49

Tropic levels are essentially the different feeding positions in a food chain, showing how energy flows through an ecosystem. Let me walk you through the main trophic levels. At the very bottom, we have the producers.

Suman Debnath1:11:01

You can try this out and see what you can mention so that you can augment the tool. In fact, this is not theright way to do this because by default it is a, uh, female voice. You can actually change the behavior of this, uh, speak tool.

Okay. So the way that you can do is, uh, you can go to the documentation. And, uh, if you see this documentation, uh, here we have tools. And, uh, you can see the overview. And if you see this here,

there is a tool spec. Yeah. Here's a tool spec for different tools. And you can mention what persona that you want. So that is a more deterministic way, uh, to do that. Or else you can put that in the, uh, system prompt.

Okay. So that's all I have. Uh, uh, if you have any questions, feel free to ask or, uh, you know, uh, you know, feel free to connect. And, uh, you know, would be more than happy, uh, to connect offline.

Q&A1:12:08

Suman Debnath1:12:08

Yeah.

Guest 21:12:10

So, uh, I have a couple of questions.

Suman Debnath1:12:12

Yeah.

Guest 21:12:13

Yeah. So have you seen any, uh, uh, companies already using this in production? And what type of upscaling, uh, do you have to do?

Suman Debnath1:12:22

Yeah. Yeah. Yeah. That's a good question. So we have used this in, uh, one of the insurance, uh, leading insurance company where they had, uh, the images of driver, uh, licenses. And they have the images of insurance policies and all that.

And, uh, we tried with different techniques. One of the techniques that we used was OCR, which worked fine. Uh, but Colpali was working pretty well. And it was the only drawback which I have seen with this Colpali, uh, uh, model is it is very heavy.

But, uh, that heaviness comes only at the time of data ingestion. So when you create the embeddings. Once that is done, at the query time, it is pretty fast. Okay. Uh, but when you are putting the data, at that time it's a little heavy.

Okay. And, uh, uh, I guess if you are thinking that if you have 1,000 documents, each has 1,000 pages, you will do a search among all those images. That is not how it works. Because imagine if I ask you the same question, forget about all this.

If you use a text based embedding model and if you have a book of 1 million pages, you have 100 million vectors. And when you ask a question, does the vector database search for all the vectors? No. There is a different indexing techniques that all the database uses.

Same indexing technique, uh, uh, are used here as well. It's just that now the vectors represent different thing. Now the vector represent patches. In the previous case, the vector represent images or, uh, sorry, uh, a chunk of text.

But, uh, that semantic search happens, uh, very efficiently, uh, using a different indexing technique. One of the techniques that we use is, uh, I think hierarchical, uh, small world navigation. Uh, so where it uses a tree based, uh, you know, structure.

Uh, it just finds, uh, the root node. Uh, I mean, it, it starts on the top layer. It finds one of the closest node. And whichever node is closest, then it goes down and finds its neighbor. So you are just, you can think of it like, uh, uh, you know, in, uh, in computer science we have tree pruning.

Right. So that's what we do. So it reduces the, uh, search space. Yeah.

Guest 21:14:36

So a quick follow up.

Suman Debnath1:14:37

Yeah.

Guest 21:14:38

Um, so can we see, uh, like, uh, more companies, uh, implement this? And then can we see this as a replacement for.

Suman Debnath1:14:46

No.

Guest 21:14:46

Traditional RAG?

Suman Debnath1:14:47

Yeah. That's a good question. No. I don't think this is a replacement. This is just another technique. And this is also a, you know, uh, uh, you know, it's a space where things are changing very fast. Right. Um, I personally feel if we get a vision based model which is more efficient in terms of computation, this might be a good model.

Uh, but again, this may work for your data, may not work for your data. So it's all about your data. What I would do and what I do generally is whenever I get some problem, I try to solve with the least, uh, I mean, the most cost effective way or most efficient way basically.

More than the cost, first we have to find out which architecture works fine for my data. If that is working fine, I don't why to complicate things and create images and all that. I will go to this only when my data set is very much complicated and where you as a human, you feel that I can read this data only if I look at it.

Imagine that you have a PDF file. For that, simple text file. For that, you can get the answer from that PDF file even if somebody converts that into a text file and give it to you. But let's say you have a PDF file which contains mostly images and embedded text on top of that.

Then you will say that, okay, okay, don't give only text. You give me the book. I will figure out because I need to see what is the context of that. So it just like it replicates humans, uh, uh, you know, uh, behavior to understand any data.

Uh, so I would recommend not to start with this. Start with the traditional technique because that is more effective, um, cost effective. And also it, it is less heavy. Because here we are storing a lot of, uh, vectors for each page.

Right. So, uh, but use this when you have a very convoluted data. Yeah. Okay. Perfect. Yes, sir.

Guest 21:16:38

So, uh, I'm trying to get a sense for when it's good and when it's not. And I'm, I'm trying to wonder whether when you chunk the image up into these little squares, is there an issue where when the chunks don't overlap, you know, let's say you're in the middle of a paragraph and you kind of chunk that into two, uh, two different segments.

Suman Debnath1:16:56

Mm-hmm.

Guest 21:16:57

Does that cause problems in practice?

Suman Debnath1:16:59

Yeah. That's a good question. But here the model doesn't know that, that there is any chunking or anything. That we are understanding it that way. But to the model, it's just an image. And the way that it creates the embeddings for that image is by, uh, doing that, those patches.

And why the model knows this? Because when the model was trained, it's a vision based model. So when the model was trained, it used to chunk all the, uh, training data set like that. And that's how the, uh, you know, it has optimized, uh, for that data.

For example, during the training time of Colpali, not at the inference time. During the training time when it was given an image of a cat and the text about a cat, the cat image was also chopped into those many, uh, uh, patches.

Similarly, when there was an image of a, uh, of a PDF page, it was chopped with the same, uh, patches. So that was inherited during the training process itself. So we don't have to question that, okay, hey model, how you are doing this.

You, the model will say that I have been doing this. Don't give me advice. I have been doing this with, uh, you know, the plethora of data. So if we, if we just look at it blindly from outside, I also had the same thought.

How the model is going to create an embedding when it splits a table into multiple chunks? What is the relationship between one chunk and the other chunk? How the model is, uh, you know, uh, doing that? Later on I realized that this has been incorporated during the training process itself.

Initially it was not able to do that. Right. But when during the training process, the loss must have been very high. Right. So and that's how it has been optimized. So once that is optimized, you don't have to worry about that.

And this is basically if you think about this, this patching and embedding, it's, it's not a new technique. Uh, uh, you know, it was there in a lot of, uh, vision based model. Now we are using it for retrieval.

So that's, that's how it works. In fact, if you are curious, I would recommend, I, I'll try to do this later on, but I would recommend that, uh, try to fine tune this model or train it from scratch if you have some resource, um, you know, for a smaller data set.

Um, and use a different patch size. Uh, uh, so let's say start with a patch size of four. Right. And, uh, uh, you know, try to see that how it works. Uh, I have a lot of assumptions, uh, on that.

But this will give you a lot of clarity of how, uh, the semantic search things work and why that matrix max matrix multiplication that we have done. Right. Why that is a good technique. Uh, uh, because imagine you have uploaded your data set is the attention you all you need paper.

And you ask a question about what is positional embedding. Now this positional embedding, this text is there in a lot of pages. Almost all the pages. It should not give me all the pages. Right. So it should give me the page where there is an actual information of positional embedding is there.

Right. And when you, when you think through that, you will find out that the ma max multiplication that, that we have done. Right. That actually takes care of that. That, uh, it will just show you the page where all the tokens of your query has the maximum similarity with the particular page.

Not just, uh, one chunk of your question, uh, with just one patch of your page. You're getting what I'm saying. Otherwise, you know, when you say top five, it will give you any five random pages where this positional embedding is written.

So just give it a shot. Yes, sir. Yeah.

Guest 21:20:47

Is there any sort of hyper approach where you can process, you know, long PDF and only send an image heavy still to Colpali?

Suman Debnath1:20:56

There is something that, uh, one of my teammate started to work on, uh, where we are trying to use, uh, Colpali along with the traditional technique. And the way that we are trying to do this is based on the question that we are getting.

And while we are doing the preprocessing and, uh, storing the embeddings, we are trying to store, uh, in a different way. Like not for all the data that we are using Colpali. Just for few data we are using Colpali.

For the rest of the data we are just using the traditional technique. But for a particular data set, we just use one single model. We cannot just go into that. Okay. First five pages of this document we will use Colpali.

The next five pages we will use the traditional technique. That's not how, uh, uh, you know, uh, uh, we are exploring. Uh, but we are kind of trying to use two different approach in the same, uh, uh, architecture.

But this is we are using because the data set that we got from the customer, they started off with a requirement, certain requirement. Then it changed. It changed means it appended. And now when the new request came, the data set is completely different.

But they want a one unified system. So that's why we are just checking the question is coming from where. And we are storing some metadata to identify this question should go from this space or that space. Uh, but, uh, nothing beyond that that I have seen.

I have seen either this or that. Yeah. Yes, sir.

Guest 21:22:15

Did you have to fine tune the Colpali model to for it to work well?

Suman Debnath1:22:20

No. I have not done that. So this is, uh, these are all, uh, uh, fine tuned models. You can just make use of this. I forgot the data set that they have used. Uh, you can read the research paper on that.

The link is there. But you don't have to fine tune that. Can you do that? Yes, of course you can do a fine tuning. That's what I was referring to him. I myself have not done that. But I will certainly try this out, uh, to fine tune that.

That's a good exercise.

Guest 21:22:42

So it worked well for your use case?

Suman Debnath1:22:44

Yeah. It just worked fine. Yeah, yeah, yeah. Yeah. Yeah. Because I used a standard textbook, uh, which are publicly available. Uh, but convoluted data. Try to do that with, uh, uh, IKEA data set. IKEA data set is good because you cannot use an OCR based techniques in that, uh, data set.

And because that's a very strange sparse data set. And that will give you a good intuition that, okay, this is, you know, you can understand. Only you can answer those questions. If you, if somebody asked you that question from that, uh, IKEA, uh, manual, you can do that.

Not, um, uh, a computer if you use a traditional technique. So that actually a good data point to make use of this. Okay. Allright. Thank you so much everyone for, uh, uh, coming. I really appreciate it. Thank you.

And, and, and one last thing is if you need, uh, uh, any AWS credit for any of your project, uh, just ping me on LinkedIn. I'll share a few credits. Okay. Even if you need more, I can give you more.