AIAI EngineerApr 1, 2025· 19:38

Scaling Agents for Gen AI Products - Anju Kambadur, Bloomberg Head of AI Engineering

Bloomberg's Head of AI Engineering, Anju Kambadur, details the company's shift from building proprietary LLMs to leveraging open-source models for agentic products, emphasizing that scaling agents requires accepting inherent fragility and building resilient guardrails rather than seeking perfect upstream systems. Drawing on Bloomberg's daily scale—400 billion structured data ticks, over a billion unstructured messages, and 40+ years of history—Anju explains how their 400-person AI team organized into 50 teams across London, New York, Princeton, and Toronto. He argues that agents must be semi-autonomous with non-optional guardrails (e.g., preventing financial advice, ensuring factuality) because compounding errors from evolving APIs and LLMs demand self-contained safety checks. The talk uses the example of a research analyst agent that factors query understanding, answer generation, and guardrails into separate components, reflecting an org structure that collapses vertically for fast iteration on single agents then later introduces horizontal teams (like guardrails) for optimization and cost reduction. Anju stresses that readiness to build horizontal teams comes after multiple…

  1. 0:00Intro
  2. 1:21AI Organization
  3. 2:20Agentic Landscape
  4. 3:39Data Scale
  5. 5:02Research Analyst
  6. 6:14Non-Negotiables
  7. 6:50Earnings Summaries
  8. 8:57Semi-Agentic
  9. 9:37Robust Agents
  10. 15:19Org Design
  11. 18:19Agent Architecture

Powered by PodHood

Transcript

Intro0:00

Anju Kambadur0:17

Thank you so much for inviting me. Um, as I was trying to think what would be a good topic to present at this talk, the organizers were really nice and— So a lot of things that you'll hear today were influenced by what the organizers thought was important, because there are really so many things happening that are exciting to talk about in the agentic landscape.

So let's get started. The first thing was, um, late 2021, I think LLMs really were starting to capture the imagination. As a company, we've been investing in AI for almost 15, 16 years, so we decided we'll build our own, um, we'll build our own large language model.

Took all of 2022 to do that, and 2023 we wrote a paper about it. We had learned a lot about how do you build these models, how do you organize, uh, datasets for these, how does evaluation work, how do you coax performance in certain zones, sort of this.

But then ChatGPT happened. I think the open weight and the open source community has come up so, uh, beautifully along. So while we continue to do very similar work, as a strategy we pivoted to say, let's build on top of, uh, whatever is available out there.

We have many, many different use cases, so I think we, we pretty much pivoted to say we'll build on top. Uh, if it helps you in any way on how we are doing things, so there you go. Uh, the other was, uh, I think there was a curiosity on how exactly, uh, does a company like Bloomberg organize its AI efforts.

AI Organization1:21

Anju Kambadur1:40

So, um, I reported to, I report to the, uh, Global Head of Engineering, and we are organized somewhat as a special group, if you will. We work a lot with our data counterpart. Bloomberg has a really strong, large data organization that you can appreciate now helps us out a lot.

Uh, we work with the product, the CTO, in, in cross-functional settings. About 400 people, 50 teams, London, New York, uh, Princeton, and Toronto. So that's a little bit about our, our group. Okay. So, um, we've been, uh, building products using generative AI, um, starting with tools, more agentic for 12 to 16 months now.

Agentic Landscape2:20

Anju Kambadur2:20

I think the effort has been really, really serious. And so there have been so many things we've had to solve in order to build something today using what's available today. Uh, and then I decided somebody must cover all of these topics, so I'm not gonna talk about these at all,right?

Uh, I think there are some wonderful speakers talking about this.

Uh, I'll try to hang around a bit after this. And, uh, I'm really, um, I'm really bullish on what the developments are in any one of those challenges that we need to solve. I think it gets easier and easier to solve those challenges.

So please don't read these as being pessimistic. It's just realistic,right? I need to build and ship things today, and that means these are the things I need to deal with today. Uh, again, we won't be touching on any of these topics today.

Um, so internally, it was really hard to say what's an agent and what's a tool because everyone kind of had their own vocabulary, and then this really nice paper came out. So when I'm talking today, when I say a tool, I mean on the left-hand side of that, uh, it's Cognitive Architectures for La for Language Agents.

If you haven't read it, you should, uh, try to read that paper. And then an agent is really like more autonomous, has memory, can evolve. So whenever I say agentic, it's on theright-hand side of the spectrum, and the other one is the left-hand side.

Um, so that's what my vocabulary will be. Finally, to set the stage for the talk, um, I don't know how many of you know about Bloomberg. I certainly did not know as much as I do today when I joined.

Data Scale3:39

Anju Kambadur3:51

So, um, we are a fintech company, as you can imagine from my nice, uh, jacket or jumper. And our clients are in finance, but finance is a very diverse field. So, uh, I've listed here 10 different archetypes of people who are in finance, and they do very different activities, but they also do a lot of similar activities.

And so, um, what is like a short form of thinking what Bloomberg does, we have, we both generate and accumulate a lot of data. This is unstructured and structured. So news, research, uh, documents, slides, we, uh, also provide access to websites.

There's a lot of reference data, uh, market data coming in. So if you just want to know the scale, every day we get 400 billion ticks of structured data information, about a billion plus, uh, unstructured messages, millions of well-written documents which include news.

And this is just every day, and we have over 40 years of history on it. So when we say we offer information as one of the, uh, things to our clients, this is the scale at which we are working.

Research Analyst5:02

Anju Kambadur5:02

Uh, the rest of this talk, I will, uh, as you can imagine, we are building a very broad set of products. So to focus the talk, I'll talk about one particular, uh, archetype, uh, research analyst. If you didn't know what a research analyst done, here is a, uh, does, here is a short course.

So, uh, there's a research analyst, they are typically an expert in a particular area. Think like, you know, I'm a research analyst in AI or semiconductors or technology or electric vehicles. And the kinds of things they need to do on a daily basis are written at the bottom.

So they are doing a lot of work with search and discovery and summarization, a lot of things with unstructured data on the left-hand side. They are doing a lot of work in, uh, in data and analytics, structured data and analytics in the middle part of the segment.

They are reaching out to their colleagues both to disperse and gather information. So there's a lot of communication. And then they're also, uh, some of them are also building models, uh, which means they need to normalize data. They need to actually program and generate models as well.

So this is a, a research analyst in a, uh, in a nutshell.

Uh, the other bit is, because we are in finance and we've been here for, uh, we've been in finance for like since founding 40 years ago, there are some aspects of our products that are non-negotiable. And, uh, those include things like precision, comprehensiveness, speed, throughput availability, um, some principles like protecting our contributor and client data, making sure that whatever we build, there is transparency throughout.

Non-Negotiables6:14

Anju Kambadur6:38

These are non-negotiables. It doesn't matter whether you're using AI or not. So these should ground you in the kinds of challenges we face when we use what's available today to build agents. Okay. So what was the first thing we did?

Earnings Summaries6:50

Anju Kambadur6:50

Uh, again, 2023 is when I think we got serious. So the first thing we did was for the research in the, in the zone of helping the research analyst community, um, companies, public companies in particular, they have scheduled quarterly calls that discuss the health of their company.

They talk about their future. It's a conference call. A lot of analysts attend the call. Uh, there's a presentation by the company's executives, and then there's a Q&A segment. And during earning season, it happens that on any given day, many, many of these things are happening.

So I told you that a research analyst has to stay on top of what's happening every single day. So

transcripts of these calls need to be generated. Again, AI is used. And in 2023, we saw an opportunity to say, well, we know what for every company which is a, which is operating in a particular sector, we know what are the kinds of questions are of interest, and maybe we can try to answer them for the analyst to take a look at, and that way they can be informed on whether they wanted a deeper dive or not,right?

Seems like a simple product. And again, I'm talking about work that started in '23. So where the technology was, we still needed to do a lot to bring it to the market, keeping our principles and features in place.

So what does it mean? Just focus on theright-hand side if you will. Um, performance out of the box was not great. Like precision, accuracy, uh, factuality, things like that. Um, and for those of you who are interested in ML Ops, I think there was a lot of work done in order to just build remediation workflows and circuit breakers.

Because remember, these summaries are not somebody just chatting with a transcript. It's actually published, and everyone gets to see the same summary. And anything that is an error has an outsized impact for us. So we constantly monitor performance, remediate, and then the summaries get more and more accurate.

So a lot of, um, I think a lot of monitoring goes in behind it. A lot of CI/CD goes in behind it as well. Okay. So today, how are the products that we are building, how does the agentic architecture look like?

Semi-Agentic8:57

Anju Kambadur8:57

Well, first of all, it's semi-agentic because I don't, this is an opinion, we don't yet fully have the trust that everything can be autonomous. So there are some pieces that are autonomous, the other pieces that are not autonomous.

Guardrails is a classic example of, for example, Bloomberg doesn't offer financial advice. So if someone starts with, "Hey, should I invest in?" then, you know, you need to catch it. We need to be factual. That's again a guardrail.

So like, those are not optional pieces for any agent. Those are coded in as you must, uh, you must do this check. So just take this, keep this image in mind. It'll come back. Okay. So this is about, this is a talk about scaling.

Robust Agents9:37

Anju Kambadur9:37

So with that long runway, let's get to scaling. So I just wanted to cover two aspects of scaling. I'm hoping that both these aspects will be more of a confirmation and not a surprise to any of you. Um, so let's see.

So the first thing is, if you wanna build agents and you want each agent to evolve really quickly, because when you build the first time, unless you're a magician, it's gonna suck a bit. And then it needs to improve and improve and improve,right?

So how do you get there? Well, let's go back to how some really good software is built. When I was a grad student, I used matrix multiplication a lot. And this is a snapshot of the generalized matrix matrix product.

And if you read the API documentation, it lays out every aspect of the input, every error code, how long it will take is also available in documentation. It's just, it just works,right? And when you build software on top of such really well-documented, well-written software, your software also tends to be robust.

Your products tend to be robust. Even from 20 years ago when we started using machine learning to build products like, you know, there are tools like APIs that use models or pipelines of models behind them. You as a caller or a person downstream of such APIs, there is a bit of stochasticity, stochasticity, if I can pronounce it correct, uh, involved,right?

You don't quite know what the result will be, and you don't quite know if it'll work for you or not. And this is despite best intentions of establishing, you know, what the input distributions are and what the output distributions are.

There's always a bit of stochasticity. It was still okay to work with them, and I'll tell you why it was okay to work with these. But when you enter using LLMs and agents, which are really compositions of LLMs, the errors multiply a lot.

And that is something that causes a lot of fragile behavior. And I, and we'll just take a look at it. And, and I, I hope my answer is mildly surprising to you on how to avoid the fragility. Um, in 2009, we built, uh, a news sentiment product.

It was basically to detect if a piece of news for a given company would be beneficial for that company or not. So the input distribution, we knew which newswires we were monitoring. We knew which language it was in.

Newswires also have editorial guidelines on how they write things. So while it's, while this, while the API that sits in front of the model is not as clean as like matrix matrix multiply, you still have a very decent handle on, okay, what is coming into my system.

And the outputs are obviously just like, you know, it's minus one to plus one pretty much. So like the output space is also very easy. Training data, we built it from scratch, so we know the training data. We could have really nice held out in time and space, um, test sets, and then we could establish the risk of deploying this.

We could monitor it. So despite all of this guardrail being present, we still ended up having a lot of out-of-band communication on anyone who's downstream of us. So for example, if you were consuming our stream of output on sentiment, we would give you a heads up.

We would tell you that, "Hey, the model version is changing. If you have a downstream application using this as a signal, you wanna test it out." Things like that. This was the landscape that's changed a lot when you think about building agentic architectures.

Like you want to make improvements to your agents every single day. You don't want to have a release cycle where there is a, you know, a purely batch regression test-based release cycle because there are so many customers who are downstream of you who are also making independent improvements to your model.

So I'll give you like one small example,right? So, uh, one of the, one of the workflows that we have agents for is, um, for a research analyst is, uh, I told you that structured data is something that they look at.

The question here is US CPI for the last five quarters. Q is just a quarter. There's an agent that deeply understands the query, uh, figures out what domain it should dispatch to, and then uses a tool. It's, there's an NLP front end to the tool, but uses a tool to basically fetch the data,right?

Um,

turns out that the data is wrong, and which is why you need the guardrails. The data is wrong because of one character that was missed. Uh, it fetched monthly data as opposed to quarterly data. And if you're actually building a downstream workflow where you're not even exposing the table, a, a good research analyst would catch it.

But if you're not even exposing the table and you're just looking at an answer that says, "Well, it looks like the answer is 42," it's really hard to catch these compounding errors, which is why it is easier to not count on the upstream systems to be accurate, but rather factor in that they will be fragile and they'll be evolving and just do your own safety checks.

Even in like, I'm talking about within my own org, people are independently operating every version of the data and analytics, analytics API tool that's coming out is better and better. But being better means being better on average. It doesn't mean it'll be better for you as a downstream consumer.

So building in some of this, um, guardrail, I just think is good sense. And that almost makes you go faster as you factor out individual agents, and each agent can evolve without having these handshake signals of, "Well, every downstream caller I have, I have to make sure that they understand what's changed and they sign off that I can actually release my, I can promote my, um, new agent to like beta or production."

I think we just need to like change that mindset and be more resilient. So that's one. The second thing is, as much as I used to code one, one, one fine day long, long ago, I'm a manager now, so I thought I'd talk about org structure and I don't know how many of you will, um, resonate with it.

Org Design15:19

Anju Kambadur15:37

Bloomberg, like I said, we've been building these things for like 15 years. And traditional machine learning, um, it has a particular factorization of software, and that software factorization is then reflected in the org structure. If you are lucky, you have the reverse Conway, uh, law of design.

But you, but you really need to rethink that as you start using different tech stacks and start building different kinds of products. Um, what do, what do I mean? How many agents do you wanna build? And what should each agent do?

And should agents have overlapping functionality or not? These are some basic questions. And typically, it's very tempting to just say, "Let's just keep our current software stack and see if we can build on top of that," or, "Let's keep our current org structure and build on top of that."

And so what I've learned is

on the columns here, you can see, you know, the first two columns are vertically aligned teams. The next two columns are horizontally aligned teams. And there are some properties in the rows. And what we've learned, and we've actually done some reorgs, what we've learned are in the beginning, you don't really know much on what the product design's gonna be, and you wanna iterate fast.

It's just easier to like collapse the org, collapse the software stack, and just say, "Here's a team, go build what needs to be built and figure things out." And that's where you want like, you know, really fast iteration.

You want sharing of code, data, models, things like that. The more you have understood this for a single product or a single agent, the more you understand what its use is and what it's good at and what it's not.

And you actually build many, many of these agents, and that's when you start thinking, "Okay, I can go back to the foundations of building good software and good orgs, and I wanna have things like optimization on it. So I wanna increase the performance, reduce the cost, make it more testable, make it more transparent."

And that's where you move into the bottomright corner of the segment where you do have some horizontal. So in our case, like guardrails are horizontal. We don't want every team, uh, every one of those 50 teams like trying to figure out what does it mean for me to not accept user inputs that are thinly veiled financial advice inputs,right?

Like it's something that you wanna do horizontally, but you also don't wanna, you want to figure out for yourself what is theright time, um, for you and your organization to start creating horizontals, to also start breaking out some of these monolithic agents, which are reflected again in your org structure and start creating smaller and smaller pieces.

So all that said and done, like, you know, just again for the, uh, running example of a research agent, this is how it looks like today. So, you know, I think taking in the usual, user world and, and session context and deeply understanding what is the question and then figuring out what kinds of information are needed to answer that question, uh, it's factorized as its own agent, uh, reflected in the org structure.

Agent Architecture18:19

Anju Kambadur18:43

Similarly, for answer generation, we have a lot of, uh, rigor around what constitutes a well-formed answer. Again, that's factored out. I call it semi-agentic, like I alluded to before, because we do have guardrails that are non-optional. There is no autonomy there.

You have to call it at multiple points. Um, and then, yeah, like we build on top of like years of traditional, uh, and more and more modern forms of data munching, like, you know, your sparse indices have become dense and hybrid indices now.

So yeah, that's a little bit, and I think I'mright at time. So have a nice day. Thank you.