Intro0:00
Doesn't this look like something's going to drop from the ceiling? Like a ground zero type thing. Be honest. Like, who has a buzzer that if I'm "I really suck," they press it and everything falls down through the trap door?
No? Yeah? Okay, who was it? Okay. You tell me if I'm doing okay or if I should take a couple steps back,right? So hi everyone, I'm Asaf, um, and I'm here to talk about GenBI and kind of— first disclaimer, this presentation was not created with GenAI.
To be honest, I actually started doing it with GPT-03 back in August, and then I did kind of a first draft. And then a couple weeks back I wanted to come in and refresh it before the conference, and then GPT-5 took over, completely messed up my slides, so I ended up doing it manually, kind of old-fashioned.
So if I'm missing like an em dash somewhere in the middle, let me know after, okay? Uh, so first of all, a bit of housekeeping. What's GenBI? So it's a fusion of GenAI and BI. It's basically an agent that helps people answer business questions with data, like a business intelligence person would do in real life.
The Setting1:14
The reason that we're pursuing GenBI is really because of the data democratization that it can bring,right? So having access to data at your fingertips without having to be reliant on a BI team that helps you find a report, figure out what it means, understand your world before they can even give you any kind of input.
So that's GenBI. A bit about Northwestern Mutual, that's where I work. So we're a financial services, life insurance, and wealth management. Been around for 160 years. Some very impressive numbers there. But first of all, I want to say, why is Northwestern Mutual a great place to do GenAI?
We got a lot of data, we got a lot of money, we got a lot of use cases, and we got access to some of the best talent anyone can dream of, really, truly humbled by the people that I get to work with.
But on the flip side, why is it hard to do GenAI at Northwestern Mutual? Because it is a very risk-averse company,right? If you think about it, our main motto is generational responsibility. I call it "don't f-shit up." Because what we end up selling to people is a decades-long commitment,right?
You buy life insurance now, if you stay with us until it comes to term, so to speak, that can be 20, 40, 80 years down the line, depending on when you buy it and how long you get to live.
And so stability is something that's very important for us because it's important for our clients. So how do we balance stability with innovation? That's what I want to talk about today.
Challenges3:14
And really, the four main challenges that we had when we even came up with the idea, kind of a pie-in-the-sky GenBI concept. First of all, no one's done it before,right? Truly. No one's done GenBI in this fashion in the past.
Secondly, and this was really a preference for us, we wanted to use actual data that's messy because we knew that those were— that's where the real challenges are going to be,right? Understanding actual messy data for a 160-year-old company, and how can we perform well within that ecosystem.
The third was kind of a blind trust bias. So the bias— the trust that you had to build was both with the users but also with the leadership of the company,right? How can we bring accurate information, accurate answers to people when all of these things that we know about and everyone's talked about is just out there,right?
No one's blind to the trust barriers, no one's blind to the accuracy barriers. So how do we convince that this is actually something that we can trust in the company? And lastly, but really firstly, when we go to approach this from an enterprise perspective, budget impact,right?
How do we convince someone in a leadership organization where risk aversion is ingrained in the DNA to even invest in something like this that no one's done before? We don't really know how we would do it. We're not even sure how it would look like when it comes to term.
Real Data4:30
So I'll start kind of one by one. And first of all, really talk about why we chose to use actual data and not synthesized data or cleansed data. So really, it's about making sure that we understand the actual complexities that we will have to face when we eventually want to go to production,right?
We know that, you know, building POCs and demos is so easy, but the gap from POC to production is so broad, especially in this GenAI space, especially because we don't know upfront how to design the system, what we would expect it to behave like.
So making sure that we operate with real data just gave us that extra confidence that when something works in the lab, it's very likely to also work in reality. But also, and maybe not in the least, less important, is that we got to work with actual people who work with the data day in and day out.
And that gave us two things, okay? First of all, subject matter expertise, which was super critical for us to be able to validate that the system is actually working, gave us a lot of real-life examples of what people are actually asking in a corporate and what people have answered to them.
So basically the evals,right, and all the testing and stuff. But at the end of the day, it also brought the business to be a part of the research project itself, and they became kind of bought into the idea as part of the process.
So we didn't just test something in the lab and then had to convince someone to go ahead and use it. The end users were part of the research process itself. And so when eventually it matured enough so we can take some of that to production, they were already there, and they actually were pulling that.
They told us, "We want to take this. How can we wrap it? How can we package it quickly enough so we can put it into practice?"
Building Trust6:39
And the next part was really about building trust. So this is about building trust, first of all, with our management team,right? Now, I don't know about you, but last time that I got a million dollars to do a research project that I wanted in a pie-in-the-sky idea, I woke up from the dream and I realized that this is not how things work in reality.
You don't just get a million dollars and go ahead and try something out. You had to show that you know what you're doing. And part of what we did, it's kind of listed out here, but obviously, you know, we did all the regular stuff,right?
We worked in a sandbox environment. We made sure that we're not using actual client data. We made sure to put in all the security risks aside. But one of the first approaches that we said we're going to take is we're not just going to build a tool that's going to be released to everyone,right?
We understood very quickly that
how people interact with the tool, their ability to verify that what they're getting isright, and also give us feedback, changes dramatically depending on their expertise and understanding of the data. So we took that crawl-walk-run approach that basically said, "We're first going to release it to actual BI experts,"right?
People that would be able to do it on their own and know what good looks like when they get it. And we're just going to expedite the process for them, kind of like a GitHub Copilot. The next phase would be to bring it to business managers.
And again, people who are closer to the BI team, but when they see a mistake, they can pretty much figure out that what they're seeing is wrong because they're used to seeing that on a day-to-day basis. And they might be less sensitive to these types of mistakes and be more inclined to give us that feedback instead of just, you know, dumping it aside and never using it again.
Giving this type of tool to executives in the company, I don't even know when we're going to get there,right? Like an executive, they want clear, concise answers that they know they can trust. We're definitely not there yet. I think that's the vision at some point in time, but the system is not accurate enough for us to get there.
Maybe it never will be.
Another way that we— another lever that we kind of used to build inherent trust in the system is that we said, "Well, in the get-go, we're not going to even try to build SQLs,"right? This is very complex. This is very hard, even for a person.
So we said, "Step number one, let's just bring information that is already in the ecosystem that's already verified,"right? We have a lot of certified reports and dashboards. And actually, in the conversations we had with some of the BI teams that we worked with, they told us, "Guys, like 80% of the work that we do is basically sending people to theright report and helping them figure out how to use it."
So the report is already there. And that again built some inherent trust into how we architected the system because we said, "We're not going to make up information. We're just going to deliver you the same asset that you would have gotten anyway, just in a much faster, much more interactive way."
And that was the alignment of expectations that we did very upfront with the users and also with the management team. Now, the biggest
process, or kind of the most important approach that we took when approaching our leadership team and convincing them that we want to do this, was to create a very gradual, incremental process that gave them a lot of visibility and control.
The Roadmap9:58
And it was very important for us to build incremental deliveries throughout that process so that not only did they have the visibility into what are we funding now, what do we get out of it, they actually had business deliverables they could realize potential from throughout the process.
And at any point in time, they could pull the plug,right, and say, "Okay, like, it's not working well," or, "We got enough out of it," or, you know, "The next phase is so, you know, unknown and long that we don't want to further invest in it."
And this is how we basically broke it down. So phase one was just pure research,right? We kind of did the shift from natural language to SQL. We figured out how to write responses. We figured out how to understand questions that are coming in, just kind of setting the stage.
Phase two was about really understanding, "Okay, so what does good metadata and good context look like in the perspective of a BI agent,"right? It looks very different if you're just chatting with something or if you're trying to do a RAG with, you know, unstructured data like documents and business knowledge and stuff like that.
And this phase on its own already had impact on the business because when we defined what good metadata looks like for an LLM, we could immediately apply that also to just the ecosystem of data users across the enterprise.
And by understanding how to extract LLM from the information, we could also— how to extract metadata, sorry, here's where the trapdoor comes into play,right? We could also project that on how or what good metadata looks like for humans interacting with the data.
We have another initiative around semantic layer going on, which tries to model exactly that, and this provided a very valuable input to that initiative as well. But the immediate next step was basically just doing this kind of multi-context semantic search,right?
People coming in, asking different questions, and having the system figure out what's theright context, what's theright information we need to bring them. And this is something that could already be packaged as its own product and delivered, and basically just do kind of a data finder and data owner finder, which is something that could take anywhere between two to maybe four weeks in an enterprise like Northwestern Mutual, just finding what data exists and who owns it so I can start the conversation with them.
And the next layer was really about pulling in information and trying to do some light pivoting around the data. Each one of these steps, as you can see, also created an input to the following step so that the research itself was kind of self-propelling, and there were incremental outcomes coming out of each one of these phases.
The next one is more kind of setting it up for enterprise-level usage, so understanding roles of different users coming in, what they may be asking about, what type of access we want to give them, etc. And eventually, and this is still some ways to go ahead, building kind of a fully-fledged GenBI agent, which doesn't only quote information from existing reports, but can actually run SQL queries on its own, pull in more data, do more sophisticated joins between different data so it can answer more complex questions.
So that's the roadmap,right? That's kind of the high-level plan. Now, why did that work well? Kind of quickly summarizing, we talked about, so we get value early and we get value often. Each one of these was a six-week sprint, at the end of which we had a very tangible deliverable coming back to the business that we could decide to productize.
And at any point in time, we could decide how we want to move forward. There was transparent progress. There was incremental business value. Each one of these steps allowed us to learn something that helped feed the next step.
And maybe the most important part, and that's the bottom line here, and that's the part that executives really look at, how do we control the risk in continuing to invest in this type of research project? And this is really about eliminating things like sunk cost bias,right?
We already paid, you know, whatever, a million dollars. Let's just get through the project, see what we get at the end. This eliminates the fear of competitors coming in, and maybe we don't need to continue investing in this,right?
So everyone in the industry is researching GenBI, and there are solutions like Databricks Genie that are coming up, and they're getting better and better. Maybe at some point in time, it's better for us as an organization to actually adopt Databricks Genie.
But at that point, again, first, it's much easier for us to pull the plug and the funding, but we already have a good understanding of what good looks like. We have benchmarks that we used for ourselves when testing our own system that we can test a third-party solution with, and we know what to expect,right?
We know what works. We know what doesn't. We know what a kind of fluffy demo from a vendor would look like, and we know where to drill in to ask the tough questions.
So let's see kind of what it looks like under the hood and how we productize different elements of this architecture. And maybe kind of very quickly, why can't we just do it with ChatGPT? So, you know, just dumping a schema into ChatGPT doesn't work.
Architecture15:19
Usually, schemas are very messy. It's not easy to understand the context and the meaning of things. And eventually, governance is super important. So there was a lot of governance built into the architecture that was very hard to apply on ChatGPT from the outside.
But even solutions like, you know, Databricks Genie is third-party, much harder to govern from the outside than from the inside, but still TBD.
So the stack kind of looks like this. We have a data and metadata layer that we produced. We have four different agents that are running across the pipeline: a metadata agent that understands the context, a RAG agent that finds the different reports, an SQL agent that can pull more data if we need that, and then eventually what we call a BI agent that takes all that information and delivers an answer to the question that was asked.
On top of that, we slap governance and trust and orchestration, and eventually some kind of a contextual UI.
And this is how the flow goes. So when a business question comes in, we push it into the orchestrator and basically decide how to facilitate the process. The first thing that we do is understanding the context. So that's where that metadata agent comes in, works with the catalog, works with all the documentation that we have across the system to understand what we're being asked about and what's the relevant information to share.
Then we go to the RAG agent, which tries to find an existing report, again, out of a list of certified reports that we know are allowed for people to use, and people have spent a lot of time fine-tuning them and making them as accurate as possible.
If we can't find the report or if it's not exactly what we need to use, that's where we go to the SQL agent that basically tries to create a more exact query or a more elaborate query. And even if the report that we have is not usable as is, it gives us that initial seed of a query that we can then expand on rather than having to build one from scratch.
So it's kind of like a few-shot example, but in this case, the example that we give is very, very close to the actual result that we're expecting to get. We then execute it against the database, pull it, and push it into the BI agent, which
then translates that to a business answer and not just dumping data back on the user. And this is what goes into the final answer. Now, there's obviously some kind of a loop that says, "If I'm in the same conversation, I'm probably talking about the same data, so we don't have to talk about this or do this again and again."
Impact18:04
Now, each one of these three components, each one of these three agents can be packaged as its own product and delivered to production with a very tangible and actual impact on business metrics, okay? And that's the kind of beauty of this approach, that after we productize each one of these, we could have basically said, "Stop," or, "Let's move forward."
And just some giving bottom-line numbers around some of these. So just the RAG agent that pulls theright report allowed us to take about 20% of the overall capacity of the BI team that basically said, "All we do is just share theright report with theright person."
So we were able to automate around 80% out of those 20%, and we're talking about a team of 10 people. So roughly two people, full-time job, all they do is find theright report and send it to theright person.
The metadata understandings that we got from learning how to interact with the data through an LLM allowed us to run A/B tests in the semantic layer project that we did, and that allowed us to prove back again to the senior leadership in the company that there is value, intangible value, measurable value in enriching metadata.
And we did that basically by running a battery of questions against a database that had good metadata and one that didn't have good metadata, and we showed how much better an LLM performs when having theright metadata in place.
So basically proving the value of something that can be very fluffy like, "Hey, let's bring in more documentation into the code." Right now, we're experimenting with the data pivoting bot. So once you have a dashboard or a report, be able to change the time horizon, some of the views, some of the segmentations, and the groupings of the data, again, kind of real-time without having a person do that for a business stakeholder.
And some of the next steps is really evaluating the tools that are out there for GenBI, like Databricks Genie, for example. And we're going to go into a much more rigorous process of enriching our catalog with metadata and documentation.
Next Steps20:12
And that's also going to come out of a lot of the learnings that we got from the research that we've done. So even if we don't end up writing a GenBI agent full-fledged end-to-end, we already got a lot of value back from this.
And this is really what allowed our senior leadership team to continuously invest in this project quarter over quarter.
One thing that I want to wrap up with is just a couple of thoughts I had about the future. So I think we talk a lot about how to prepare data. I think that's going to be a huge area in the market, and there are going to be probably a lot of companies and tools that are going to help us with that, building very specific, task-specific models and applications.
I think a lot of startups and companies are going to come up from that area. Copilot is really making sure that we meet the users where they are and securing of models, obviously a very big thing. The last thing that I want to focus on the most, because that's kind of a recent thought that came to me a couple of weeks ago, how we do pricing of SaaS in the GenAI era.
This is really about the fact that one individual person today can be 10x more effective than they used to be in the past. And then do we price software based on seats? Or do we price software based on how much they used it?
Or do we price software based on the value that they got out of it? Salesforce is already experimenting with that. So the data cloud product at Salesforce is starting to be usage-priced and not seat-priced. And I think this is going to have a big impact on just the kind of SaaS economics worldwide.
And it doesn't even matter if the product itself is GenAI. It's really about what does the person using the product can do, and what can they do in their other time, and whether it still makes sense to price it by how many employees you have or how much work you get done with the employees that you have.
That is me, and thank you very much for listening, and thanks for not opening the door on me.





