The Myth0:00
Hey folks, I'm Anushrut, I lead the applied research team here at PromptQL. Uh, PromptQL you might have seen as the sponsor for the reliability track here at the AI Engineers Booth Fair. So today I'll talk about, um, that data readiness is a myth.
How many of you are trying to deploy some kind of AI system on some kind of data in a production environment? Okay, awesome. Who is trying to work with, uh, more than, like, documents and vector databases? Okay. Is whose data is perfect?
Clean, annotated, perfect column names, table names? Anyone? See, no hands. Okay, so, and how much time does everyone—like, okay, I'm not going to ask—make this a question. You all spend a lot of time, uh, making your data ready,right?
So that the AI can understand it. So that the AI understands the meanings and the relationships in your data. But it's just a pipe dream that we're all chasing,right? We are all chasing this perfect data dream so that our AI can finally work reliably on it.
And that's never going to happen. So how do we still make AI reliable?
Only no matter how messy data we have.
Okay. So tell me if this is a fact. Is this what your data looks like? Your, uh, that's how you name your tables. That's what you—how you name your columns. Sometimes they have null values. Sometimes you have old values, old column names.
Sometimes you have shorthand, CST_NM. Does that customer name? Is it custom nomenclature? I don't know what that is. Uh, rev_amount_ust, that's how you name your revenue. Is active? Now, is this, uh, binary? Is it Boolean field? Is it 0, 1, true, false, null, not null?
I don't know what it is,right? Then you have other systems. That's just one system, one table in one system. You have other systems which have similar things,right? It has organization name, it has total revenue. How does that map to your other systems,right?
Is the revenue in cents? Is it in dollars? Decimal value, floating points, what is it,right? You have no idea. Okay, so in 2019 everyone was saying, "Let's standardize everything. Move everything in Snowflake. Let's move everything in Databricks and finally your problems will be fixed."
40% complete. Uh, MDM will fix this. Master Data Management team will fix this. That's their responsibility. They're still implementing their solution. So in 2023, with the rise of AI, rise of agents, we'll, we'll add create semantic layers that understand our, uh, data domain.
I mean, it breaks every quarter. Your data domain changes every quarter,right? You change your tables, you change your schemas, you change your, uh, workflows. So in 2025 we are saying AI needs perfect data to work. And it's still waiting.
And it's never going to happen,right? And McKinsey, uh, said that on an average a Fortune 500 company loses $250 million because of poor data quality. So how do we fix this? Uh, who has tried playing with semantic layers?
Semantic Layers3:16
Mem0, Atlins, Semantic Kernel, you have tried playing with it,right? How has your experience been?
We actually use JAGOR.
Okay.
To create our kind of flexible model.
Okay, gotcha.
Uh, yeah, it needs a lot of pruning.
A lot of pruning. Okay. And I'm assuming you're manually adding information to it, maintaining it, and stuff like that,right?
We do it automatically, but yeah.
Sure. Yeah. Uh, okay. So let's say you've added a definition, like customer acquisition cost means marketing spend divided by new customers. Okay. Now this is some information your AI needs to answer your questions. But that's not enough,right? Like, what which marketing spend?
Coming from the brand team, from the performance team, what does a new customer even mean,right? First purchaser, uh, purchase, uh, customer, reactivated customer, for what time period are we talking about? Does it include failed trials or not? Uh, does it is it accounting for seasonality?
There are so many things that you need to do. And you can never capture all of that in a semantic layer just by if you think you can manually add everything, you can't,right? So you can't predefine, uh, every edge case.
Knowledge graphs, who's played with knowledge graphs, Graph RAG? Heard of it at least? Okay. A bunch of people. Okay, awesome. So, uh, let's take a very simple example. Assume a customer's, uh, assume a sales data set, uh, where you have defined this graph that deals map to stage, a date, and an owner.
Okay. And a very simple question I ask. Show me deals at risk,right? Very simple questions. The graph knows that deals map to stages, stages map to close dates. But what does at risk mean? Is it mean that the champion has just left?
Is it mean it has been stuck in that stage for two months? Like, what does at risk mean to my business,right? How do you capture that in a graph database,right? How do you capture a billion rows of Snowflake data table in a, uh, graph database?
You can't. So graph is also not the solution. Knowledge graph is also not the solution. So the real problem here is not that we need, we, we need a better semantic layer solution. We need a better Graph RAG solution.
Business Language5:24
We need a, like, better, uh, named database systems. No. The problem is that the AI does not speak your business language. Like a GM in a finance domain can mean gross margin, but in a HR domain might mean general manager,right?
What does conversion mean to you? What does quarter mean to you? What is the definition of your quarter? I'll show you an example, uh, here with an AI system not working with that. And, uh, what is, what is, what is an active customer?
Every team has their own definition,right? So how does an AI speak this tribal knowledge, this tacit knowledge that you have developed while being in your company for so many years,right? Your AI does not know that. Your vanilla LLMs don't know that.
They're super smart, incredible at doing so many cool stuff, uh, so much cool stuff, but they don't understand your business, your domain,right? So traditionally we had these analysts, these engineers,right? Whenever a, a business user or a customer had a problem, had a task, they had a question, uh, they would go to this analyst or an engineer who knew about the business, who had this tribal knowledge in their head.
They knew how to write code, SQL, whatever. They can talk to your underlying data systems. SQL, NoSQL, doesn't matter. Any kind of data source,right? And they have this tribal knowledge which they use to answer your question with 100% reliability.
They explain what they're doing,right? And that's how you have all built trust in your colleagues, in your, uh, peers,right? That's what is missing with AI,right? This, this tribal knowledge piece, that doesn't exist today with AI. And that's the problem.
So the solution, the same semantic layer, but let's make it agentic. What that means is let's not try to improve it. Let's not try to, uh, manually add context to there continuously. No. How about we make an AI system that behaves like the analyst you just hired today?
Agentic Layer6:56
Day one, day zero, the analyst comes to your company, super smart, can do a bunch of things, doesn't know a lot about your business yet. They start working with you. They mess up somewhere, you tell them, "No, no, that's not what you should have done.
This is what I mean when I say this." It learns, learns, learns, learns. Now this analyst, 10 years later, is an experienced analyst in your company. They know everything about your business,right? Let's make an AI like that. An AI that keeps improving, keeps learning as you use it more and more, as you course correct it, as you steer it.
But assumption is your AI needs to be correctable, explainable, steerable, already accurate in what it knows,right? So let's see how you build such an AI,right? So we are trying to replace this human part of the AI,right? So that's what we have been trying to do with PromptQL.
It's like a day zero smart analyst,right? So we take a foundational LLM, and that's the whatever LLM you bring. We make it create PromptQL plans. PromptQL is basically a domain-specific language which can do three tasks: data retrieval, data compute aggregation, and semantics.
And this is a deterministic domain language. And vanilla LLMs are incredible at generating. We have don't have to fine-tune them,right? Now, within this DSL, I can ask the LLM to create this DSL whenever I ask, usually ask the question.
Now I can execute this DSL in a deterministic runtime. I do not involve the LLM in actual execution, actual generation of an answer. Because if I let the LLM generate the answer, it's by default hallucinating. And I'm just hoping the hallucination is correct.
That's how LLMs work,right? So I'm saying decouple it. Let the LLM generate the plan and we will execute this plan in a deterministic runtime. And let it work on a distributed query engine which will, uh, talk to different data sources, pull out the data, do whatever composition was required inside that DSL, show the answer directly to the user.
Don't give it back to the LLM. Let's not do RAG,right? Let's not give the LLM, uh, data back to the LLM and make it generate the answer,right? That's what, uh, the PromptQL design is. Let me show you PromptQL working in action.
Uh, I have five minutes, so let's make this quick. Simple question. Uh, who are my top five customers by revenue? You'll be like, any AI system can answer that question, dude. It's a simple text-to-SQL question. Okay. So PromptQL is like, first of all, it understands what revenue means.
Live Demo9:19
Revenue means your invoice items. Okay, cool. Uh, I'll do, uh, do the math, execute the math, and here are your top five customers. Oops, nope. I didn't notice we did not get any results,right? That's what a smart analyst does.
In their first attempt, they realize they messed up somewhere. Okay, I see the issue now. We're looking for succeeded status, but the actual statuses are paid and pending. See, your data is messy. It did not know what was happening,right?
It figured that out. And now these are your top five based on the actual data, uh, that's under the hood. So that's possible. Let's run this query now. Okay. Find the unique customers we serve. The cost org ID data is messed up, so don't use that.
Find unique orgs based on the email domains of the individual users. Then find the org with the third highest revenue. I'll, I'll let it run because, uh, it'll take time. Uh, find the unique orgs based on the email domains of the individual users.
Then for the org with the third highest revenue, take a look at the latest 30 support tickets. So multiple database, then your Zendesk support, uh, system, uh, including the comments on those tickets. Then summarize each ticket. Then use those summaries and extract their feelings towards our product.
Create five categories from bad to great. And then tell me what is going well, what, what can be improved. And then issue up to $5,000 to this org proj, uh, org project as, like, credits with the highest usage, uh, which means 5,000 if their feeling is bad towards us and 1,000 if it's great towards us.
Now, how do you think an AI can do this? Spread across your databases, your, uh, SaaS application like Zendesk, your, uh, API, uh, like Stripe to issue the credits, stuff like that. So they'll be cool. First, I'll get all the users, extract their email domains.
See, this is an analyst explaining their thought process. You see that tiny, uh, pencil icon? I can edit their brain. I can tell them, "No, this is not, uh, what I wanted you to do. Don't do step three.
Instead of that, do these three steps instead,"right? I can, I can be in charge of my AI, but still, every single time if I have to nudge my AI, the AI is going to learn. The AI will understand, "Okay, that's what you wanted me to do.
Makes sense. That's how you do your business. I didn't know that. Sorry. I'll learn now,"right? So, okay, cool. I should have said, "Just show me intermediate results step by step." But anyway, um, it says I'm about to issue $3,000 refund.
Probably it got a neutral sentiment from it. Um, and see, it's exactly what's happening. It's saying, "I got the top five domains by revenue. Uh, this is the sentiment. I summarized a bunch of tickets. I classified, extracted the sentiment out of that.
And then finally, I'm, uh, figuring out what is happening there." Um, okay. So ready to show details. Perfect. Based on analysis, uh, Peter's Thomson dot Bizzard's third highest revenue customer with this much in revenue, sentiment analysis, decent support tickets, blah, blah, blah.
See, this AI just woke that an analyst as an analyst with such a complicated prompt. So now let's look back at the learning,right? Uh, just because we are at two and a half minutes. Okay. So that's day zero.
Day zero, it had to figure stuff out, figure it out, did, did well. Day X, once it's become a veteran analyst,right? It has learned a bunch. Let me make this bigger for you. Uh, learned a bunch,right? So as it's using, uh, as you are working with it,right, it learns from it.
There's a PromptQL learning layer which basically improves the semantic graph and starts creating your company's business language. This AcmeQL, Acme, assume it's a name of a company,right? And now suddenly PromptQL becomes AcmeQL, becomes Google QL, Microsoft QL, Apple QL, Cisco QL, whatever company you come from,right?
Self-Improving13:04
And I'll show you that learning process in action. Okay. So this is an example, uh, where I have purposefully named my tables extremely bad,right? And I ask the question, which employees are working in departments with more than US, uh, dollars, $10,000 budget?
Okay. It says, "I have no idea what you're talking about. Your data says there are three tables called Morc, Plug, and Zorp. I have no idea what that means. I can't answer your question." I'm like, "No worries. Can you sample a few rows from each table and figure out what table con contains the employees?"
Like, cool. Like, that's what I tell man. Let's go figure it out,right? It's like, okay, now I see Zorp contains employee information, Plug contains, uh, department information, and Morc is like a junction table that you have. Okay. So now this is your answer that you were asking for.
And I'm like, "Okay, but the data is in cents, not in dollars. The budget is in cents. So can you divide by 100, please, and give me theright answer now? There shouldn't be five, uh, employees." It says, "Cool.
There are two employees." Like, perfect. Now this is, um, uh, this is manual, but this also runs agentically in the background where all I have to do is suggest metadata improvements based on the recent threads,right? Look at what have we spoken about, how much I had to guide you, whatever, uh, hints I had to give you.
Learn from it and improve your semantic layer,right? It's like, cool. So based on the, uh, interaction that we have, now I have this, um, improved semantic layer where it's like, okay, Zorp and Plug are two tables. I need to add a lot more context to it in my own, uh, semantic layer,right?
And the department budget is in cents. Like, cool. Apply the suggestion. Every single instance of your semantic layer is version controlled. So when new build is created, you can always fall back to a previous build. And now the next time, as a generated by autograph, that's what you call the feature.
And the next time I ask the same question, which employees are working in departments with more than US $10,000 in budget? It creates theright plan and I get theright answer,right? So same. If I had to say something like this, uh, find accounts, um, with the maximum suspicious anti-money laundering, uh, outgoing amounts for the first quarter for each, uh, print the account ID and name,right?
If I let my AI do this, okay, the internet is a little bad, but, um, I have this thread preloaded here. Okay. So it gives me the answer. And then I'm like, no, the account, uh, my quarter starts in February, not in January.
So you should know that,right? So I just tell it this time. Next time, the semantic layer learns from it exactly the same way as it inferred the meanings of these tables. It'll find the relationships across your tables,right? And finally, with all of this, what you have is
day zero. Your AI does not know what an enterprise customer means. It does not know how to match customer IDs across systems. It does not know what your financial quar when your financial quarter starts. Day 30, it has figured out 47 business terms.
Outcome15:46
It has mapped relationships across the six systems. It has discovered 12 calculation variants and 100% accurate on your complex tasks. That's what an agentic semantic layer allows you to do. So reduces, uh, months of work into immediate start.
Like, just deploy your AI today. Let your AI start working on your data. Let it improve itself. No wait time, no lag time with your AI deployments,right? It's self-improving and gets to 100% accuracy. Uh, that's what we've been hearing from our customers.
Conclusion16:13
It's a Fortune 500 food chain company. They evaluated 100, uh, vendors. They realized, no, none of them work. Finally, they saw PromptQL and it worked 100% reliable AI. Same with a high, high growth fintech company. The on their hardest questions, we were able to demonstrate 100% accuracy.
So try out, um, reach out to us if you have these big problems that you want to solve and you want 100% accurate AI on top of that. Uh, we're there for you. Thank you.





