AIAI EngineerApr 7, 2025· 20:59

Anchoring Enterprise GenAI with Knowledge Graphs: Jonathan Lowe (Pfizer), Stephen Chin (Neo4j)

Stephen Chin (Neo4j) and Jonathan Lowe (Pfizer) explain how Pfizer uses knowledge graphs with GraphRAG to accelerate drug manufacturing technology transfer, cutting time from years to weeks. They argue graph databases provide superior accuracy and explainability over vector-only RAG, crucial for life-saving drugs. Jonathan details navigating organizational silos in a 100,000-person company, from C-suite taglines to client partners demanding cost savings. Gartner's prediction of 30% GenAI project failure is addressed with a concrete business case: manufacturing worker tenure plunged from 20 to 3 years, making AI essential to capture lost expertise. The architecture combines vector and graph retrieval to deliver contextually relevant, auditable answers, reducing data consolidation from three months to three weeks.

Transcript

Intro0:00

Stephen Chin0:17

Hey, so it's so great to be back in New York City. Actually, I grew up nearby here, and pleased to be co-speaking with Jonathan.

Jonathan Lowe0:25

Thank you, Stephen. Good to be here.

Stephen Chin0:27

And, you know, we're here to kind of talk about leadership, talk about how you can actually do a bunch of the things you've been hearing in practice. We're going to talk about strategy, we're going to talk about technology, but let's start with analysts.

Failure stats0:41

Stephen Chin0:41

Who here trusts Gartner when Gartner says something? They're predicting the AI wave, they're predicting— okay, nobody does. Nobody— no hands went up in the room, for the record. But when they are predicting failure and catastrophes, I try not to trust that,right?

So last year they predicted 30% of generative AI projects are going to be abandoned by the end of 2025. Now, anybody in the room— this is a real, real honest check— has anyone been on a failing GenAI project?

Okay, now, brave souls, amazing. Give those guys a round of— that took a lot of courage. Now, to make them feel a little better, who hasn't yet got to production on their GenAI app? Okay, so the rest of the hands went up,right?

So this— this is the challenge. So we all want to be successful with GenAI. We all want to do amazing things. We're getting asked to do amazing things, but we— we need to have theright way of approaching this in our organizations, with leadership, to sell this internally, to build it on technologies which they can understand.

And

Executive pressure1:51

Stephen Chin1:54

the vision— it's hard to get a vision that's technically achievable when the guy at the head of the table is this— is this guy. He's— he's the executive who's heard about GenAI. His kids are using it for their school courses, and he's like, "Oh, yeah, yeah, it just— it solves all the problems.

Insert here success. I want it in production in two months." Now, I think the great thing about having— having Jonathan as my co-presenter is that he's actually done this in a big life sciences company, and he's had to navigate all of this leadership challenges, organizational challenges, silos, to build a system which actually is something we can take to production.

So tell us a little bit more about that, Jonathan.

Pharma challenge2:37

Jonathan Lowe2:37

Thanks, Stephen. Now, as I've been introduced, Jonathan Lowe. You may know me as Jonathan when we're out in the hallway or when I give you a bit more information about my experience launching GenAI-based capabilities in business. You may think of me as Debbie Downer.

AI is so exciting.

Guest2:59

Until the singularity.

Jonathan Lowe3:04

So that's how I actually approached the problem I'm about to explain to you, but it actually worked. So the business case was technology transfer, which means in biopharma, scaling up from lab bench— think beakers and human-scale drug development— to industrial scale, making a million doses a day.

And to get from that lab bench level to multiple factories around the world making lots and lots of product very quickly takes years because the industrial people that build the factories and build the equipment need to sift through hundreds of thousands of documents and notes and test outcomes that were created at the science level.

Another challenge with doing that is I'll go to a statistic. In 2019, a study said that the average tenure of manufacturing workers— tenure being how many years they had spent in their companies— was about 20. Twenty years of average tenure.

What do you think the average tenure is in manufacturing companies today?

The study said three years. So we've gone from 20 down to three. And all that expertise has— has or will soon be retiring because the boomers are growing old. So we really need generative AI. We need a machine to take a lot of the intelligence that's captured in documents or even in tacit people's heads and get it to the new people showing up to do this technology transfer.

So we take all these millions of documents and we've loaded them into a graph. Now, we haven't necessarily loaded the document into the graph, we loaded the chunks into the graph. And one of the things that we really liked using the graph to accomplish was we structured the chunks: the document, the block, the paragraph, the line, because we wanted to understand when we searched for those chunks with similarity search which ones really returned the results that people wanted the most.

Graph chunking4:53

Jonathan Lowe5:30

We wanted to really refine how we stored and managed the chunks. So at this— at this point, it was— it was a totally new space. And because we were able to structure in the graph that level of chunking, we were able to eventually learn and get better and better at how we chunked the documents in the first place.

Stephen Chin5:49

Yeah, so I— I think what's really amazing for me about this is we were talking about business challenges and, like, like projects failing. And in the study that Gartner did, the biggest failure mode was not having a business use case which would actually solve real problems and then be monetizable.

And this is not only, like, a great business use case, but it's also something which potentially is saving lives because you're— you're getting life-saving drugs to folks faster, you're able to accomplish this quicker. But the problem is always the humans in the middle,right?

Internal hurdles6:23

Stephen Chin6:23

So the teams you work with probably have a little bit of GenAI not invented here syndrome. Where you come along with this great solution, like, "I'm going to use GraphRAG, I'm going to load all these documents into my— my, you know, my big store," and they're like, "No, no, no, we've heard— seen this research paper, we watched this talk, there's some other platform we want to use, there's some other framework."

Or maybe it's— maybe it's too expensive. I mean, compared to classic computing and cloud computing, GenAI architectures have the potential to be much more expensive if they're not well architected and, in general, are going to increase the cost of the organization.

So how do you convince people to go from a system which is— is working but not working well enough to a much more expensive system which is R&D, investment, redevelopment, to go towards a GenAI architecture? So what are— what are some of the challenges you hit internally, and how did you address that at Pfizer?

Jonathan Lowe7:23

Great. So for this one, it's more of an entrepreneurial use case within a big organization. I wonder how many of you have worked in organizations with 50,000 or more people. A lot of hands going up. My current organization has over 100,000 people.

Org navigation7:23

Jonathan Lowe7:42

I've also worked at IBM, Deloitte, big organizations. And if you are like me in these organizations, you'll be that little red guy going like this with the— with the light bulb over his head saying, "I have an idea that might help the company, and I have a team of X number of data scientists and developers and SREs, and we can bring that value, that capability to the company."

If you're like me, if you're that red guy, who's the first group of people on this slide that you're most interested in connecting with?

Guest 28:29

CEO.

Jonathan Lowe8:30

You go for the top? You're better than I am. I— I joined this whole profession because I love building applications that delight the people that use them. So my instinct has always been to go to the bottom first and say to those users, "Hey, do you really want this tool?"

And what are those users going to tell you? They'll like your tool if.

Guest 28:55

It's good.

Jonathan Lowe8:55

That'sright. What makes it good? It takes away boring stuff that they don't want to do. But it can't just take away boring stuff, it also has to give them accurate results. It also has to work in a performant way.

They can't push the button, go get coffee, and come back. And I feel like that's the easy part,right? More and more these days, you can build accurate, fast applications quickly. So where's the real challenge? So somebody said you go to the top first.

What's the likelihood in a company of 50,000 to 100,000 people that you're going to meet the CEO if you're the guy with the idea at the level four of the hierarchy? The likelihood is pretty small. Did anyone here ever see the movie Dirty Dancing?

Dirty Dancing? Maybe? Do you remember the part in Dirty Dancing when Baby, the lead— leading woman in the— in the movie, meets Johnny, the amazing dancer, for the first time? And she— she's— she's unable to speak, she's so flustered, and finally she blurts out, "I carried a watermelon."

And then off he goes, and she goes, "I carried a watermelon? Two weeks ago I stepped into the elevator on the seventh floor of— of the headquarters of— of my company, and there was my CEO in the elevator.

And I felt like Baby in Dirty Dancing. I couldn't think of what to say. I locked up. Fortunately, he's a good guy. He broke theice. I'm just back from vacation, rolling up my sleeve, can't wait to get to work.

What are you up to?" And then, thank God, ding, we got to his floor, the doors opened, and— and out he went. And as he went out the door, I blurted out, not "I carried a watermelon," but "I'm working with LLMs."

Off he went. So when you're trying to promote your work within a big company like this, it would help to know what that executive is trying to accomplish. And the way that he gets to that point is he talks to consultants who say, "Let us tell you how to be a leader in your industry and not fall behind the competition."

So an example of something that an executive at that level might do is create a purpose blueprint or something like that name, and the number one message has to be a few words and convey something that the whole company can follow.

So an example of that might be change a billion lives a year. In life sciences, a big aspiration. Now, why do you have to care about that? In the elevator, maybe you'll reference it. I'm changing a billion lives a year with the most amazing AI search engine.

Ding, and off he goes. But that— that message that he gives trickles down to the next level: the chief digital officer, the chief scientific officer, the chief supply officer. What do you think they're going to say? They're going to try to take his message and turn it into their specific flavor.

So the digital officer will say, "I want to lead the industry in AI," and the scientific officer will say, "I want to take on the world's biggest diseases," and the supply officer will say, "I want to accelerate supply."

Still very high level. And you probably won't meet these people either. Who will you meet, though? You'll meet their level twos and their level threes. And what are they going to say?

At this point, they don't really say taglines. Instead, they say, "I want cost savings, I want cost avoidance, I want earlier realized revenue, or I want more balanced headcount."

So when you're talking to these people, your slides have to have numbers and times and your promises about how your tool or your capability or report or whatever is going to meet— meet those times and those numbers. Now, you may not get to meet them either.

If your big company has a— a form of a role called the client partner where your digital people talk to the client partner and the client partner talks to the business, then that's the other person you have to convince.

And the problem with this is that the client partners tend to stay within their particular departments. There might be a client partner who works exclusively in R&D or one who works exclusively in supply. What would they say? Sometimes they don't say the same thing.

One of them might say, "R&D already has five or six or ten search engines. Why build another?" Or they might say, "Search engine is a great idea. Why don't you incorporate that capability into every tool in the supply organization?"

So either your scope goes to nothing or it goes to everything, and you need to be able to negotiate and navigate that. Are you done? If you can satisfy all those people and cross through all those gauntlets, well, no.

Because as you're starting to build, the vendor comes to you and says, "Why build in-house when you can buy our tools?" And they've been talking to the chief digital officer about build versus buy and which one is more economically realistic and appropriate.

Well, maybe you get through that and then you're done,right?

Who else would possibly stand in the way of your incredible AI search tool? Friendly fire is the answer. Your own colleagues, either a level above or at the same level, may say, "Dude, I was here first. AI search is my turf."

Or they might just say, "Hey, that client partner over in supply isright. Can you please integrate with the stuff I've built?" So I guess my message is we've heard a lot of talks about failure and challenge and Gartner not liking this.

It's an incredible time to be in this— in this amazing industry, in this amazing change in both, for me, life sciences and more generally for the information technology industry. And I love that we're hearing all this concern about failure because it just means we're at the beginning of a really exciting time.

But as representatives of that, my advice to you is know your audience, personalize for all of them, and get your human wetwear chatbot speaking theright language at theright level.

Stephen Chin15:26

No, that's— that's amazing. We've chatted about a bunch of these challenges,right? So we've chatted about getting a good business use case that can actually provide value. The organization, how to navigate, like, like peoples and— and different failure modes within the organization where the organization has a huge quantity of people who can be your allies or can work against you, depending upon how you work with them.

Why graphs?15:26

Stephen Chin15:50

But it's also a technology problem. You have to have theright technology to solve your use case. Now, one of the— the biggest challenges, I think a lot of us who have been building RAG and enterprise applications, has been the LLMs themselves fighting us with— with hallucinations.

This is getting better with newer models. It's getting easier to feed theright sort of information in with vector databases. But I think that you've chosen a rather unique approach using graph databases. Why did— why did you choose to use a graph database for your implementation at Pfizer?

Jonathan Lowe16:26

Well,

there are a lot of things that graphs aren't good at. Things like genealogic sequences of recipes or social networks or hierarchies or time series, and all of those applications were prevalent opportunities within Pfizer. So that was the— that was the original impetus for using a graph.

But I also discovered that the more data we consolidated in the graph, the faster my data scientists and engineers and developers and SREs were able to understand the data landscape. What used to take three months to consolidate, understand, clean up, took three weeks or less for— for a new project.

So I know the reason a lot of people take on graph is because traversal becomes so much easier in terms of data search and— and performance gets better. But I found that team performance also took a really big boost from using that tech.

Stephen Chin17:25

Cool. And for folks who aren't familiar with— with knowledge graphs and LLMs or— or GraphRAG, this isn't a new idea, although I— I would put you guys on the early adopter where you're actually in production now with something that uses this.

GraphRAG 10117:25

Stephen Chin17:38

But Microsoft kind of wrote the seminal paper on GraphRAG and used it, basically taking existing documents, LLMs, to chunk it into a graph and then showed superior results coming out of it. It— on the spectrum of technologies, using LLMs directly, you can get good results, but it lacks that context, it lacks that enterprise knowledge.

Using a vector database or— or baseline RAG, you— you can get better results where now it's actually pulling in organizational knowledge, but the answers tend to be a little bit generic. There's a lot of hallucinations. GraphRAG kind of pulls us to the end of the spectrum where now you're— you're getting answers from that— that knowledge graph you built.

You can evolve over time and much more precise answers, which actually get to the heart of— of real problems in— in life sciences and manufacturing and business-critical industries where you can't afford to be wrong.

Jonathan Lowe18:37

And also where in— in industries that are complicated, if there are a lot of connections that might not appear in a relational database because no one bothered to make the joins permanent, whereas in a graph, those joins are there to begin with.

So if you search for one thing, suddenly the neighborhood of related stuff becomes available to you to share with an LLM for better contextual knowledge.

Architecture18:58

Stephen Chin18:58

Yeah. And, you know, I think just if folks are implementing this or folks are thinking about how to— how to think about architectures for GraphRAG, this is a really simple way of thinking about it. So basically what you're doing is you're taking your GenAI application and you're doing both a vector and a knowledge graph representation of the data.

So you're both asking the vector for the answer. You're getting relationally close nodes from the graph database where you're getting additional context and passing that into the LLM. And then this gives you more contextually relevant results coming out of your— your expert system.

Explainability19:37

Stephen Chin19:37

So I think this is a great way to use a knowledge graph, either that you built up over time or that you have the LLM construct to kind of get those superior results where you can do better governance, you can put controls and properties on the graph nodes to control who has access to the information.

You can get better explainability now because when you're getting an answer from the LLM, you're no longer looking at statistical probabilities in the vector space. You're actually looking at graphs and nodes and edges, which we can reason about and we can start to understand the relationship between things, understand, like, what— what things are related to manufacturing, which things are unrelated to that, they're just, you know, general terms.

Outro20:20

Stephen Chin20:20

And for theright application, maybe we're saving lives, getting drugs to people more quickly, and using GenAI for a good cause. So thanks so much for joining us for our presentation, AI Engineering Summit, and appreciate everybody.

Jonathan Lowe20:39

Thank you.