AIAI EngineerMay 10, 2026· 10:11

Feedback Loops are All You Need — Mehedi Hassan, Granola

Mehedi Hassan, a product engineer at Granola, argues that shipping AI features into production requires building feedback loops rather than one-shotting better prompts. Granola's chat feature for meeting notes revealed problems with web search — token costs ballooning to 10p per chat, overnight provider updates silently degrading results — and with prompt personalization, as a single prompt cannot serve salespeople, engineers, and HR managers equally. To close the gap, Granola built custom internal tracing exposing tool calls, search trails, reasoning, and cost in a UI accessible to non-engineers, not just CloudWatch queries. They also refactored their Electron app's renderer to run as a web app, enabling preview links on every PR and allowing Cursor to automatically test changes and upload screenshots. The result is faster iteration and confidence that shipped features actually work for customers.

Transcript

Intro0:00

Mehedi Hassan0:16

Cool. How's it going, guys? I'm Mehedi. We're going to talk about some product engineering stuff that we've been doing at Granola. This is not going to go deep into AI engineering stuff, so if you're here for me to, like, go into LLMs and stuff, it's not going to happen.

I'm warning youright now so you know what's coming. Cool. So I'm a product engineer at Granola. I've been, you know, coding since Jay Career was cool. I've seen React kind of change front-end engineering, and obviously now experiencing LLMs change engineering and everything else, just like many of you.

For those of you who don't know, Granola is an app for getting your work done. Essentially, we're a meeting notes app where we sit on your doc like it is doingright now, and it has access to your system transcription, system audio, as well as your microphone audio, which means we have real-time transcription, and then at the end of your meeting we can give you really awesome notes.

So I'm just going to give you a quick demo. So I was recording the previous talkright here, and you can see it picked up literally everything the presenter said. And the cool thing about Granola is that you can also write your own notes on top of what the transcription is saying, so the final result is more aligned to, like, what you'd normally actually write on a notepad.

Demo1:08

Mehedi Hassan1:27

So I'll go ahead and generate the notes here, and you'll see that this will go ahead and write a really good summary. And as you can see, like, I wrote down this 20% overlap thing, and it focused more on the output,right?

So this is Granola. We have the best-in-class meeting notes no matter what role you're in, and it doesn't get in your way, and that's been, like, our product philosophy since day one. So we ship a lot of AI features in Granola, and our product is known to be, again, not to get in your way.

AI Challenges1:59

Mehedi Hassan1:59

So let's see what happens when you put a simple AI feature into prod. I'm going to kind of give you an example with this chat feature that we have. This is a feature that already exists in Granola. You can ask questions about a meeting that you just had and across a bunch of different meetings, or, like, shared context as well, and Granola will try to answer it to the best of your ability.

So let's say I built, you know, like, a one-shot list chat system. It's very easy to do, and I put it into production in my fake Granola app. And then as soon as users hit it, you know, it's like, it can't give me a list of to-dos.

Web search is too slow. It's not writing follow-up emails how I normally write my emails. I asked it to coach me about my meetings, and it's telling me meetings about my football coach. Obviously, these are very, very common problems that you're going to run into when you make a generic chatbot.

Web Search2:47

Mehedi Hassan2:47

So how do we get around this,right? So what we've seen is, like, molding the LLM to work to your specific use case can be super hard. And one of the examples is web search. So web search for most LLM providers looks like a line of code.

You simply add the web search tool, and you expect it to just work. That's what the labs want you to believe, but once you get into it, there's lots of other complications. So, for example, the token usage and token costs can bubble up quite a lot, especially for complex queries.

It's going to blow up your context, and each chat could be costing you, like, 10 pence. Obviously, at scale, when you have millions of users, this is not really feasible. And then, you know, like, the web search providers are also, like, completely up to the labs as well.

So, for example, in our development, what we see was, like, we were using a model for a good amount of time, and then overnight they shipped an update, and for some reason web search degraded, and it was completely out of our control.

And we generally had no idea, like, what was going on apart from just, like, switching providers. But we want to have more control over that because it affects our user experience. And, you know, like, there's literally billion-dollar companies who do web search, so that kind of tells you that it's much more than just adding a web search tool to your LLM pipeline.

The other thing that's super important for apps like Granola is the output. So the summary that you saw was pretty good for what I would expect, but someone in sales might expect more of a deal focus. Someone in engineering might expect, like, action items, blockers, or, like, linear tickets.

Output Diversity4:02

Mehedi Hassan4:17

HR might want something completely different. And the thing is that one prompt can't generally serve everyone, and, you know, LLMs are stubborn, and we need to figure out how to get inside them and make them work how we want it to work.

And, yeah, as you know, like, LLM behavior is largely seen as, like, a black box, but we want to kind of go very deep into the details and figure out exactly what's going on. So what we did recently at Granola is we started building our own tracing tools, and obviously, thanks to LLMs, you can actually one-shot these things.

Tracing Tools4:40

Mehedi Hassan4:47

And this is where one-shotting is kind of nice. And so we built our tooling tracing tools here where we basically have complete visibility on the tool calls straight from the beginning to the end. So we have full visibility over the individual tool calls, why it's making those tool calls, the search trails, the reasoning trails, the cost, structured exactly how we want it.

And the most useful part of this is that we structured the data exactly how we want it, and the UI is built to, like, serve our employees internally, not just, like, engineers, but also product data and, like, CX and everyone.

So you don't have to, like, you know, go into CloudWatch and do, like, very complex queries to figure out why something failed. And that's been, like, the key for us to, like, figuring out this black box. And previously, obviously, building this kind of tools would be, like, up to using a SaaS provider, and it simply wouldn't you simply wouldn't have the time.

But now you actually can spend time building this tracing tool that actually serves what you need. And this is obviously a very basic example, but obviously you can use OpenTelemetry or, like, other providers, but we essentially just, like, save things to a DB, wrap around, like, AI SDK, and then the front end is, like, kind of the most important part because that's what people are going to use to figure out what breaks and what doesn't.

And we literally have, like, our founder literally goes into, like, the details, like, following the agent loop completely front to back to figure out exactly what went wrong. So then at the end of this, you can actually figure out, like, you know, this output fills off to, like, exactly what failed, and then when you iterate, you can improve on those things.

But as I said earlier, this is going to be more than just, like, basic LLM stuff. And LLM behavior is obviously part of the picture. The how users interact and experience your product is also very important. So with LLMs, you can one-shot more things, and you can have more variants, which we like because we can experiment with different features.

Desktop Limits6:39

Mehedi Hassan6:39

We can experiment with, like, one feature looking very different for different users. But the problem for us specifically at Granola was that we are a desktop app, which means you can only run one instance of the app at a time, and there was a lot of friction when it came to, like, testing new features, different variants, and actually testing those in parallel.

So before, you know, before you'd have to, like, run the Electron app locally, install the dependencies, and test things, if you wanted a coworker to test those changes, you'd have to get them to do those things as well.

Electron Refactor7:11

Mehedi Hassan7:11

We didn't have the same luxuries as, like, web apps do. So essentially what we did is we took our Electron app, and we turned the front end of the Electron app into a web show, and this was deployed online.

So now our CI, whenever we open a PR, we get a preview link, and we can go and test those things. And this generally sped up our development time so much more. And, like, the cooler part of this is that because LLMs can now self-verify their work, these guys are now, like, once we open a PR, Cursor goes and tests it, uploads a screenshot into our PRs, which speeds up the testing so much more.

And again, this is, like, you might think that this is a lot of work, but it's actually quite simple. So what we did is, for those of you who are not familiar with Electron, there's obviously a main process and a renderer process.

The main process works with the system APIs, and the renderer process is basically your front end. And essentially we abstracted our IPC APIs, which is the system APIs, to fall back to web standards when we're in the web environment.

And similarly with React APIs as well, like routers, sessions, and query layer, we moved those to the web standards. And essentially this just made the renderer agnostic of Electron, and we could just simply run it as a web app.

So essentially this has helped us on top of the LLM improvements was, like, we were able to just, like, change, like, test, like, one feature in, like, multiple different variants. So, like, whatever the end product is actually feels super good because we know that we've tried so many different variants, and we actually felt those products in, like, in practice rather than just, like, seeing it in Figma.

Feedback Loop8:45

Mehedi Hassan8:45

So essentially this is basically a long talk to tell you that the answer isn't to one-shot better. It's about figuring out how you can make that feedback loop where it kind of feels like playing a tennis game with LLM so the end product feels more like magic rather than just, like, a black box and hoping that the feature that you're going to release works well with customers and having that conviction that what you are shipping is actually going to connect to the users.

Thank you. Any questions?

Q&A9:19

Guest9:19

Are you thinking about replatforming from Electron to Tory?

Mehedi Hassan9:23

We've thought about moving to Tory a couple of times. I think the way Electron serves usright now has been super nice. Like, the API is changing quite a lot. We've tried Tory before as well, and we didn't really see massive performance gains, which is what we care about the most.

So, yeah, it's been discussed before. We've played around with it but haven't shipped it.

Cool. Thank you, guys.