Introduction0:00
Hi everyone, my name is Gabe De Mesa. I'm an engineer here at OpenGov, and today we're going to be talking about agents in production, specifically how OpenGov built and scaled OG Assist. Uh, so, um, this presentation is going to be jam-packed with just so much good stuff.
We're going to talk about AI agents, we're going to talk about our harness, we're going to talk about
evals, observability, traces. We're going to talk about tools and skills. It's—there's going to be a lot of good stuff in here. We're going to talk to you guys about what we do at OpenGov and how we operate at the scale that we operate at in production, so you'll be able to see a real use case and workload with AI agents.
So, without further ado, let's get started.
Agenda0:56
Okay, agenda. So, just really quickly going to go through, uh, high level what we're going to talk about today. I'm going to tell you guys a little bit about OG Assist and what OpenGov is. I'm going to tell you guys the origin story of how this all kind of came to be.
We're going to talk about OG Assist's big bet on Effect, a little bit into our core agent loop. We're going to talk about the A2A protocol, evals, and sandboxing. We're going to talk about how we manage long context.
We're going to talk about monitoring, observability, how we collect feedback, and how we iterate on that feedback. We're going to, lastly, also talk about tools and skills, and how at OpenGov we use AI not only externally, that we serve to customers, but also internally to improve our development workflows.
Just a little bit about me before we go any further. My name is Gabe. I'm a software engineer here at OpenGov. I work on the AI Agents team, and I'm one of the folks that helped build OG Assist and some of the systems that you guys will be seeing today.
OG Assist2:08
So, a little bit about OpenGov: OpenGov is a software company on a mission to power more effective and accountable government. So, OpenGov sells ERP software. That's things like budgeting, procurement, asset management, and permitting. And, um, we were founded about 14 years ago.
And what's cool is
we have this thing called OG Assist, and OG Assist is this little button on the top of all of our products in the navigation bar. And what's cool is all of our product suites and product teams have built tools and skills in order to power this button.
So, for example, if I open up this—if I click this button and I open up OG Assist, it says, "Hey, I'm going to ask about rate codes," which is very specific to utility billing, the current product that I'm in.
And you can see that inside of this kind of chat interface, I'm able to speak to an agent, and the agent is able to make tool calls in order to look up information against data inside of that suite.
So, it's really cool to be able to kind of first-party create these experiences through the capability that we've built called OG Assist.
Okay, so, just a quick story about how this all came to be. So, a little while back, we saw that AI was really starting to take off, and a principal spun up this new team called the AI Agents team and asked me to join.
And, um, instantly I said yes, and OG Assist started to grow. And we started to integrate OG Assist into all our products, and not only our backend capabilities, but also our frontend capabilities as well. So, you'll see that one of the capabilities that we give the agent is it's able to see what's on the screen and see and take action on what's on the page.
So, you could see that I'm asking the agent here, "Hey, hey, what's on the screen? Can you maybe highlight some of the next steps that I could take?" So, you can see that the agent here is thinking. It's saying, "Okay, what tools do I have available to use?"
And, "Hey, let me go and highlight something that you could actually click on and tell you more about it." So, just another capability of OG Assist, and just a little short story about how this all came to be.
Effect4:35
So, the big bet on Effect. So, I really wanted to include this slide because, um, here on the agents team we made a huge bet to—to bet on Effect. And suffice to say, it's paid off in dividends. We write Effect.
So, Effect is this library for TypeScript. It's open source, and it helps you write better TypeScript code. You know, it's got a lot of stuff baked in it, like a schema similar to, like, Xod, if you've ever used that.
It's also got things for error handling, for logging, for traces, for—it's just got so much in there. It really helps write better code and structure your code better, and helps with architecture, spinning up new services for—and for us on the agents team, really helping design and build the core agent loop.
So, you'll see throughout this presentation sprinkled in how Effect on our team has paid off in dividends. So, we really love Effect here at OpenGov, and we encourage other folks to try it out. And, um, yeah, let's keep going.
The Effect native loop. So, originally we were on LangGraph, and that was fine until the team really started to scale and our use cases started to evolve. So, we decided to move over to our own kind of Effect native agent loop to have full regency over this agent loop, such that if we have complex use cases or features that we need to build, we could kind of get in—we had full control of the agent loop.
And not only that, but now we're fully on Effect. So, all the cool things you get with Effect is now propagated throughout the entire agent loop, like the tracing, structured concurrency, the logging. Everything is more fine-grained control, and it really allows us to really unlock the full potential having our own agent loop from the ground up.
So, another thing I wanted to mention is, on the left side, you'll see a code example. This is really the basics of the Effect loop that we're using. We're using this thing called the Effect AI package, and in that package there's this thing called—there's a chat and a language model.
So, with the chat, you can instantiate, like, a chat, for example, and then you could stream text using that kind of stream text function. You could pass in a prompt. And what's cool is, with a language model under the hood of—since we're kind of doing dependency injection—we could pass in a different language model if we were to hot-swap to another one, for example.
So, really just having full control of our own agent loop just kind of gives us all the levers, and it really just unlocks the full capabilities of the model and for the team as well, to have full agency over this loop.
A2A Protocol7:46
Another thing I wanted to mention is the agent-to-agent protocol. So, here on the agents team, we've had a lot of success with this protocol. So, this protocol being the protocol that Google created, kind of an open protocol for agents to intercommunicate.
But we found this very useful for defining our agent routes, for example, in the backend, and our model and our schema to follow this kind of agent protocol. So, we modeled—so, for example, there's this thing called an agent card, which you see here, and it's got the name of the agent, a description, et cetera,right?
And having this kind of rigorous protocol, this rigorous spec, really helped drive our development and drive alignment, because, you know, all we had to do was align with this spec and follow this spec, and we knew that this was kind of the contract that our frontend and backend would both consume and produce.
So, this, I would say, also has been very helpful for us. And what's really cool is A2A has a lot of extensions,right? So, you could extend the protocol, add in, like, metadata. There's also A2UI. So, lots of fun stuff with A2A protocol, but this is kind of what's worked for us.
Evals & Feedback9:12
So, sharing that with you folks. Feedback and evals. So, here the quote is: "Shipping is the start, not the finish." So, what we do here on the agents team is we have kind of multiple ways we do evals and collect feedback.
Obviously, you know, we'll have folks call in or email us or just let us know and tell us. But the main way is we have this thumbs-up and thumbs-down mechanism. And here, someone is able to tell us, "Hey, this worked really well.
This was a great response," or, "That wasn't a great response." And that signal we take and we're able to iterate on, and we could take it back and help improve, you know, the response in the future. We also have automated evals.
So, in our CI, we have evals that run against real completions. So, we could test a prompt against, "Hey, did it hit some tools? Did it do what it's supposed to do?" And that also helps with our accuracy.
So, those automated evals in conjunction with collecting feedback really help us improve our
Safety & Sandbox10:22
tools, our skills, our harness. And that's really how we're able to iterate so fast and so quickly. Humans in the loop. So, this is a really cool feature we built where we deterministically interrupt the agent loop if there is a tool call approval required.
So, if an agent tries to make a tool call that it needs human approval for, it'll show this UI, and the human can click accept or reject. So, explicitly rejecting or explicitly accepting the action that the agent is trying to make.
And this ensures that, you know, we're building trust and also ensuring that, you know, we're being safe, especially when the agent is trying to do a mutating operation. And always, always, always making sure that humans are in the driver's seat.
Sandboxing. So, another thing that we worked on, kind of similar to the safety slide we just saw, was whenever an agent tries to execute code or tries to create files, it does so in a sandbox. So, we gave our agents sandboxes such that it could spin up these sandboxes on demand, and it could use those sandboxes to, honestly, write code, execute code, create files.
And it's kind of this safe, ephemeral, isolated space such that the agent can take action in there and not—and we don't have to worry about any risk to, you know, our production systems. And it's really cool because they also get tied, tethered down at the end.
So, in this example, I said, "Hey, create a PDF for the folks of the AI Engineer Conference 2026, and allow me to download it so I can share it with them." And you could see that the agent created this really cool PDF inside of that sandbox.
So, just really wanted to cover this sandbox feature and just give you a brief kind of overview of sandboxing.
Long Context12:27
Long context. So, inside of OG Assist, we have hit many hurdles, like with—especially with legacy models now, like with token limits or just way too much—just completely overloaded with context, especially as conversations get longer. So, we found that having some sort of rolling summarization was more effective than, you know, always stuffing in the latest and most recent messages.
Rather just, you know, give, like, a running summary after n number of messages, and maybe you only want the, like, n minus 5 most recent messages or n minus 10 most recent messages,right? And it may be that you're only talking about a specific topic now, but you may want to refer to context earlier, like 100 messages above.
Then
that's where kind of the memory component comes in, because when you have this rolling summary of a really long conversation, then you could do recall over that summarization. And, you know, if you ask the agent, "Hey, remember that thing that we talked about?"
then the agent within the thread will be like, "Yeah, I do know what you were talking about. I have kind of this short little tidbit." And it can, you know, follow up and do more kind of with that summary, rolling summary in mind.
So, that's kind of how we handled long context and memory, and it's worked pretty well for us. So, just wanted to share a little bit about that and how we've solved the long context problem.
Generative UI14:11
UI on the fly. So, in this example, I said to the agent, "Hey, generate me a long essay, but give me some examples about what the essay could be about." So, what's really cool is the agent had this primitive registered of this form, and it was able to build out this form for me at runtime and give me some options of what I could choose from.
So, it feels very personal, and it feels very kind of in the moment that it's able to give me these options just at runtime. So, this is kind of just a short little thing I wanted to include here about generative UI and how we are able to render UIs on the fly.
You can't scale what you can't see. So,
Observability15:00
this kind of section is about tracing and observability. Really, what's cool about Effect is you kind of get tracing out of the box. You know, when you use these Effect functions, they all get kind of tagged automatically with, like, these spans, and kind of the span gets picked up and feeds into these traces so that you can kind of get these kind of drill-downs of these function calls.
So, here is an example of a trace from the Effect team. I have it linked. You could see that, like, "Hey, when you hit this API, it goes to this endpoint, to this handler," and, you know, et cetera.
And it takes—and what's really cool is you could profile all your traces,right? So, this takes a total of this many seconds, and you can see where the bottleneck is here. If there's a failure, you can cross-reference it across services.
So, really, really important, especially working in agentic systems where we're integrating with other teams and other APIs and other platform capabilities. So, what's cool with Effect is you get all this tracing out of the box, and it really makes building this agentic experience, debugging it, and maintaining it just the breeze.
Tools and skills. So, not only did we make a big bet on Effect, but we also made a big bet on tools and skills. So, we believe that tools and skills are really all you need. And in this case, you could see on the left, we have this tool called GetDadJoke, and this is kind of the Effect way and kind of the building blocks of how we do things here at OpenGov.
Tools & Skills16:20
But, you know, this is pulled from the Effect website, but you could see, "Hey, this is how you make a tool," and then you add it to a toolkit, which is a collection of tools. And then you can register this toolkit with the language model.
So, for example, if you had a prompt that said, "Hey, generate some dad jokes about pirates," well, guess what? The agent has a tool that can help get a dad joke. So, really, this is the building blocks of how we did tools and eventually skills, and really just has paid off wonderfully for our organization.
So, really recommend trying out this Effect AI package from Effect and just trying out building out your own tools and skills.
Dev Velocity17:38
Developer velocity. So, not only do we build agents for our customers, but we also use agents here internally in OpenGov. So, we use a lot of Claude and Cursor, and it's just really been a game changer for our team.
It's funny because we're building tools and skills for customer-facing agents, and that has been great, but we're also building them internally as well to help accelerate our development workflows. So, things like Claude, Cursor, Cloud Agents, they really help accelerate how we read, write, review code, and ship.
So, it's just been such an accelerant. So, definitely just wanted to mention that before we wrap up.
Outro18:23
That's it. Thanks so much for watching. You've made it to the end. Let's build agents that ship to production.





