AIAI EngineerJan 5, 2026· 1:52:25

Claude Agent SDK [Full Workshop] — Thariq Shihipar, Anthropic

Thariq Shihipar of Anthropic presents the Claude Agent SDK, arguing that Bash and file-system-based agents outperform traditional tool-only approaches for autonomous tasks. He defines agents as systems that build their own context and trajectory, contrasting with structured workflows. The SDK, built on Claude Code, emphasizes the Bash tool as the most powerful primitive for composability and code generation, enabling non-coding tasks like data analysis. He demonstrates live-coding a Pokémon team advisor that dynamically fetches API data via scripts, and explains security through a 'Swiss cheese defense' of model alignment, AST parsing, and sandboxing. Shihipar also covers skills for progressive context disclosure, sub-agents for parallel work, and hooks for deterministic verification, stressing that agent building is an art of reading transcripts and iterating on context engineering.

  1. 0:00Intro
  2. 8:08Strong Opinions
  3. 15:33Bash Tool
  4. 22:43Agent Loop
  5. 25:45Tool Selection
  6. 50:54Spreadsheet Design
  7. 1:23:46Prototyping
  8. 1:44:52Q&A

Powered by PodHood

Transcript

Intro0:00

Thariq Shihipar0:21

Okay, yeah, thanks for joining me. I, uh, I'm still on West Coast time, so it feels like I'm doing this at, like, 7:00 a.m. So, yeah, but, um, glad to talk to you about the Claude Agent SDK. So, um, yeah, I think, like, this is going to be like a rough agenda of what we're going to talk about.

We're going to talk about, like, what is the Claude Agent SDK, why use it, there's so many other agent frameworks, what is an agent, what is an agent framework, um, how do you design an agent, uh, using the Agent SDK, or just in general.

And then I'm going to do some, like, live coding, or Claude is going to do some live coding on prototyping an agent. And, uh, got some starter code, but, uh, yeah, I— the whole goal of this is, like, you know, we got 2 hours, we're going to be super collaborative, ask questions.

This is also going to be not like a super canned demo in the sense that, like, we're going to be, like, thinking through things live, you know, I'm not going to have all the answersright away. And I think that'll be a good way of, like, building an agent loop, I think, is, like, really mu- very much, like, kind of an art or intuition.

So, um, but yeah, before we get started, just curious, a show of hands, like, how many people have heard of the Claude Agent SDK, or— okay, great. Cool. How many have, like, used it or tried it out? Okay, awesome.

Okay, so pretty good show of hands. Um, yeah, so I'll just get started on, like, the, like, you know, overview on agents. I think that, like, this is— I think something that people have seen before, but I think it still is taking some time to, like, really sink in, uh, how AI features are evolving, you know?

So I think, like, when GPT-3 came out, it was really about, like, single LLM features,right? You were like, oh, like, hey, can you categorize this, like, return a response in one of these categories. And then we've got more, like, workflow-like things,right?

Hey, like, can you, like, take this email and label it, or like, hey, here's my codebase, like, index for your RAG, can you give me, like, the next completion or the next, um, the next file to edit,right? And so that's what we'd call, like, a workflow, where you're very, like, structured.

You're like, hey, like, give me this code, give me code back out,right? And now we're getting to agents,right? And, uh, like, the canonical agent to me is Claude Code,right? Claude Code is a tool where you don't really tell it— we don't restrict what it can do, really,right?

You're just talking to it in text, and it will take a really wide variety of actions,right? And so agents, uh, build their own context, like, decide their own trajectories, are working very, very autonomously,right? And so, uh, yeah, and I think, like, as the future goes on, like, agents will get more and more autonomous, um, and we— uh, yeah, I think it's like we're kind of at a break point where we can start to build these agents.

Um, they're not perfect, you know, but it's definitely, like, theright time to get started. So, um, yeah, Claude Code, I'm sure many of you have tried or use. Um, it is, yeah, I think the first true agent,right? Like, the first, uh, time where I saw an AI working for, like, 10, 20, 30 minutes,right?

So, um, yeah, it's a coding agent. And, uh, the Claude Agent SDK is actually built on top of Claude Code. And, uh, the reason we did that is because, um, basically we found that when we were building agents at Anthropic, we kept rebuilding the same parts over and over again.

And so to give you a sense of, like, what that looks like, of course, they're the models to start,right? Um, and then in the harness you've got tools,right? And that's, like, sort of the first obvious step, like, let's add some tools to this harness.

And later on we'll give an example of sort of, like, trying to build your own harness from scratch too, and what that looks like, and how challenging it can be. But tools are not just, like, your own custom tools, they might be tools to interact with your file system, like with Claude Code.

Um, did the volume just go up, or were they not holding it close enough? Okay, there's some echo now. Anyways, um, you've got tools, tools you run in a loop, and then you have the prompts,right? Like the core agent prompts, the, um, the prompts for the tools, things like that.

Uh, and then finally you have the file system,right? And, or not finally, but you have the file system. The file system is a way of context engineering that we'll talk more about later,right? And I think, like, I— one of the key insights we had through Claude Code was thinking a lot more through the, like, context, not just a prompt, it's also the tools, the files, the scripts that it can use.

Um, and then there are skills, which we've, like, rolled out recently, and, uh, we can talk more about skills, uh, if that's interesting to you guys as well. Um, and then, yeah, things like, uh, sub-agents, uh, web search, you know, like, um, like research, compacting, hooks, memory.

There are all these, like, other things around the harness as well. Um, and, uh, it ends up being quite a lot. So the Claude Agent SDK is all of these things packaged up for you to use,right? Um, and yeah, you have your application.

So I think, like, uh, to give you a sense of, uh, yeah, to give you a sense of, like, maybe why the Claude Agent SDK is, um, yeah, like, like, so, yeah, people are already building agents on the SDK.

A lot of software agents, uh, you know, software reliability, security, incident triaging, bug finding, um, site and dashboard builders, if you're— these are extremely popular. If you're using it, you should absolutely use the SDK. Um, MS Office agents, if you're doing any sort of office work, tons of examples there.

Um, got some, like, you know, legal, finance, healthcare ones. Um, so yeah, there are tons of people building on top of it. Um, I want to— oh, yeah, okay. So why the Claude Agent SDK,right? Like, why did we do it this way?

It's, why did we build it on top of Claude Code? And we realized, basically, that as soon as we put Claude Code out, yeah, the engineers started using it, but then the finance people started using it, and the data science people started using it, and the marketing people started using it.

And, yeah, I think it just, like, it— we just realized that people were using Claude Code for non-coding tasks. And we felt, and as we were building, you know, non-coding agents, we kept coming back to it,right? And so, um, it's a, like— and we'll go more into why that just works, why we could use Claude Code for non-coding tasks.

Uh, spoiler alert, it's like the Bash tool. Um, but, yeah, it's, uh, it was something that we saw as an emergent pattern that we want to use, and we built our agents on top of it,right? And, uh, these are lessons that we've learned from deploying Claude Code that we've sort of baked in.

So, uh, tool use errors, or compacting, or things like that. Stuff that is, like, very— can take a lot of scale to find, you know, like, what are the best practices we've sort of baked into the Claude Agent SDK.

Strong Opinions8:08

Thariq Shihipar8:08

Um, as a result, we have a lot of strong opinions on the best way to build agents. Uh, like, I think the Claude Agent SDK is quite opinionated. I'll talk over some of these opinions and why, like, uh, why we chose them,right?

Um, but yeah, one of the big opinions is the Bash tool is the most powerful agent tool. So, okay, um, what, what are, like, what I would describe as the Anthropic way to build agents,right? And I'm not saying that you can only build agents using the API this way,right?

But this is, like, um, if you're using our opinionated stack on the Agent SDK, what is it,right? So, roughly, Unix primitives, like, the Bash and file system, and, you know, we're going to go over, like, prototyping an agent using Claude Code.

And my goal is really to sort of show you what that looks like in real time,right? Like, why is Bash useful, why is the file system useful, why not just use tools. Um, yeah, agents, uh, I mean, you can also make workflows.

I'll talk about that a bit later. The agents build their own context. Um, thinking about code generation for non-coding, um, like, we use Code Gen to generate docs, query the web, like, do data analysis, take, uh, unstructured actions.

So, um, there's a lot of, like, uh, this can be pretty counterintuitive to some people. And again, in the, like, prototyping session, we'll go over how to use code generation for non-coding agents. Um, and yeah, every agent has a container or is hosted locally because this is Claude Code.

Uh, it needs a file system, it needs Bash, it needs to be able to operate on it. And so it's a very, very different architecture. I'm not planning to talk too much about the architecture today, but we can at the end if that's what people are interested in.

In, or sorry, by architecture, I mean hosting architecture. Like, how do you host an agent, and, like, what are best practices there? Have you talked about that at the end? Um, yeah. So, well, let me pause there because I feel like I covered a lot already.

Any questions so far on the Agent SDK, agents, um, yeah, like, what you get from it?

Guest10:15

Can you explain what code generation for non-coding means exactly?

Thariq Shihipar10:20

Yeah. Um, this is, um, like, basically, when you ask Claude Code to do a task,right? Like, let's say that you ask it to, uh, find the weather in San Francisco and, like, you know, tell me what I should wear or something,right?

Like, uh, what it might do is it might start writing a script, uh, to fetch a weather API,right? And then start, like, maybe it wants it to be reusable. Like, maybe you want to do this pretty often,right? So it might fetch the weather API and then get the, like, maybe even get your location dynamically,right, based on your IP address.

And then it will, like, um, you know, check the weather and then maybe, like, call out to, like, a sub-agent to give you recommendations. Maybe there's an API for your closet or wardrobe,right? It's like, so that's an example.

I think that, like, it's kind of, um, for any single example, we can talk over how you might use Code Gen. Uh, a lot of it is, like, composing APIs is, like, the high-level way to think about it.

Yeah. Uh, yeah, and.

Guest11:30

Yeah. Uh, workflow versus agent, uh, like for repetitive tasks or, you know, like a process, a business process that is always the same, do you will still prefer to build an agent versus a fully deterministic workflow?

Thariq Shihipar11:44

Yeah. So we do have— oh, sure, yeah, yeah. Um, so the question was about workflows versus agents, and would you still use the Claude Agent SDK for workflows? Is thatright? Um, yes. And so, uh, I mean, we— I just— we just sort of tell you what we do internally, basically.

And what we do internally is we've done a lot of, like, GitHub automations and Slack automations built on the Claude Agent SDK. So, uh, you know, we have a bot that triages issues when it comes in. That's a pretty workflow-like thing.

But we've still found that, you know, in order to triage issues, we want it to be able to clone the codebase and sometimes spin up a Docker container and test it and things like that. And so it still ends up being, like, a very— like, there's a lot of steps in the middle that need to be quite free-flowing.

Um, and then you, like, give structured output at the end. So, um, yes. Allright, we'll take one more question and then keep going. So, yeah, in the blue.

Guest12:41

Yeah. Uh, so could you talk about security and guardrails? Like, if, if, you know, you're using Claude Agent SDK and, you know, you lean towards using Bash as the, you know, all-powerful generic tool, then is the onus on, uh, building the agent builder to make sure that, you know, you're preventing against, like, common attack vectors?

Or is that something that the model is doing by itself?

Thariq Shihipar13:03

Yeah. So I think this is sort of, like, the Swiss cheap— oh, yeah, sorry. So the question was, uh, permissions on the Bash tool,right? Or, like, how do you think about permissions and guardrails? Like, in, like, when you're giving the agent this much power over, you know, your— it's environment of the computer, how do you make sure it's aligned,right?

And so the way we think about this is, uh, what we call, like, the Swiss cheese defense,right? So, like, there is, um, like, on every layer some defenses, and together we hope that it, like, blocks everything,right? So obviously on the model layer, uh, we do a lot of, um, alignment there.

We actually just put out a really good paper on reward hacking. Super recommend you check that out. Um, so, like, definitely, I think Claude models, like, we try and make them very, very aligned,right? And, uh, so, yeah, there's the model alignment behavior.

Then there is, like, the harness itself,right? And so we have a lot of, like, permissioning and prompting. Um, and, uh, like, we do, uh, AST pass parser on the Bash tool, for example, so we know, um, fairly reliably, like, what the Bash tool is actually doing.

And definitely not something you'd want to build yourself. Um, and then finally, the last layer is sandboxing,right? So, like, let's say that someone has maliciously taken over your agent. What can it actually do? Uh, we've included a sandbox in, like, where you can sandbox network requests, um, and sandbox, uh, file system operations outside of the file system.

And so, uh, yeah, ultimately that's what they call, like, the lethal trifactor,right? It's like, um, like, the ability to, like, execute code in an environment, change a file system, um, exfiltrate the code,right? I think I'm getting the lethal trifactor a little bit wrong there, but, like, the idea is basically, like, if they can exfiltrate your, like, information back out,right?

Um, that's, like, they still need to be able to extract information. And so if you sandbox the network, that's a good way of doing it. Um, if you're hosting on a sandbox container like Cloudflare, uh, Modal, or, you know, Edebeat, Eternal, like, all of these, like, sandbox providers, they've also done, like, some level of security there,right?

Like, you're not hosting it on your personal computer, um, or on a computer with, like, your prod secrets or something. So, uh, yeah, lots of different layers there. And, yeah, we can talk more about hosting in depth. Um, so, okay, so I'm going to, uh, talk a little bit about Bash is all you need, you know?

Bash Tool15:33

Thariq Shihipar15:33

Um, I think this is something that, oh, yeah, um, this is, like, my shtick, you know what I mean? I'm just going to, like, keep talking about this until everyone, like, uh, agrees with me. Um, or, like, I think this is something that we found at Anthropic.

I think it is sort of something I discovered once I got here. Um, Bash is what makes Claude Code so good,right? So I think, like, you guys have probably seen, like, code mode or programmatic tool use,right? Like the, um, different ways of, like, composing SOPs.

Uh, Cloudflare put out some blog posts on that. We've put out some blog posts. Uh, the way I think about code mode is, like, or Bash is that it was, like, the first code mode,right? So the Bash tool allows you to, you know, like, store the results of your tool calls to files, uh, store memory, dynamically generate scripts and call them, compose functionality, like tail and graph.

Uh, it lets you use existing software like FFmpeg or LibreOffice,right? So there's a lot of, like, interesting things and powerful things that the Bash tool can do. And, like, think about, like, again, what made Claude Code so good?

If you were designing an agent harness, maybe what you would do is you'd have a search tool and a lint tool and an execute tool,right? Like, you know, N tools,right? Like, every time you thought of, like, a new use case, you're like, "I need to have another tool now,"right?

Um, instead, now Claude just uses Graph,right? Or it knows your package manager, so it runs, like, npm run, like, test.ts or index.ts or whatever,right? Like, it can lint,right? And it can find out how you lint,right? And it can run npm run lint.

If you don't have a linter, it can be like, "What if I install ESLint for you?" Right? So, um, this is, like, yeah, like I said, the first programmatic tool calling, first code mode,right? Like, you can do a lot of different actions very, very generically,right?

Um, and so to talk about this a little bit in the context of non-coding agents,right? So let's say that we have, uh, an email agent, and the user is like, "Okay, how much did I spend on ride sharing this week?"

Um, you know, like, it's got one tool call, or generally it's got the ability to search your inbox,right? And so it can run a query like, "Hey, search Uber or Lyft,"right? And without Bash, it searches Uber or Lyft, it gets, like, 100 emails or something, and now it's just got to, like, think about it.

You know what I mean? And I think, like, a good, like, analogy is sort of, like, imagine if someone came to you with, like, like, a stack of papers and like, "Hey, how much did I spend on ride sharing this week?"

Can you, like, read through my emails? You know what I mean? Like, that would be really hard,right? Like, uh, you need very, very good precision and recall to do it. Um, or with Bash,right? Like, let's say there's a Gmail search script,right?

It takes in a query function. Um, and then you can start to save that query function to a file or type it. You can grep for prices, you know? You can, uh, then add them together. You can check your work too,right?

Like, you can say, "Okay, let me grep all my prices, store those as, like, in a file with line numbers, and then let me then be able to check afterwards, like, uh, was this actually a price? Like, what does each one correlate to?"

Right? So there's a lot more, like, dynamic information you can do to check your work with the Bash tool. So this is, like, um, just a simple example, but, like, hopefully showing you sort of the power of, like, the composability of Bash,right?

So, uh, I'll pause there. Any questions on Bash is all you need, the Bash tool, anything I can make a little bit clearer? Yeah.

Guest 219:17

Do you have stats on how many people use YOLO mode? I assume it's like everyone?

Thariq Shihipar19:21

Uh, stats on YOLO mode, we probably do. Um, I mean, internally we don't, uh, but that's just, I think we just have a higher security posture. Um, yeah, I'm not sure. Uh, I can probably pull that. Any other questions on Bash?

Okay, cool. Um, yeah, just to give you, like, some more examples, like, let's say that you had an email API and you wanted to, uh, you know, like, go through, like, fetch my email, like, tell me who emailed me this week,right?

So you've got two APIs. You've got an inbox API and a contact API. Um, this is, like, a way you can do it via Bash. You can also do it via Code Gen. This is kind of, like, enough Bash that it is Code Gen,right?

Like, um, Bash is essentially a Code Gen tool. Um, and then, yeah, like, let's say that you wanted to, you had a video meeting agent,right? You wanted to say, like, find all the moments where the speaker says quarterly results in this earnings call,right?

You can use FFmpeg to, like, slice up this video,right? Um, you can use JQ to, like, uh, start analyzing the information afterward. So, um, yeah, lots of, like, powerful ways to use, uh, to use Bash. So, allright, I'm going to talk a little bit about workflows and agents.

Yeah, you can do both. You can use, uh, build workflows and agents on the Agent SDK. Um, yeah, agents are, like, Claude Code. So if you are, like, building something where you want to talk to it in natural language and take action flexibly,right, then that's where you're building an agent,right?

Like, you want, you have an agent that talks to your, like, business data and you want to get insights or dashboards or answer questions or, uh, write code or something, like, that's an agent,right? And then a workflow is kind of like, you know, we do a lot of GitHub actions, for example,right?

So you define the inputs and outputs very closely,right? So you're like, "Okay, take in a PR and give me a code review." Um, and, yeah, both of these you can use the Agent SDK for. Um, when building workflows, you can use structured outputs.

We just released this. Um, you can, yeah, Google Agent SDK structured outputs. Um, but, yeah, so you can do both. I'm going to primarily be talking about agentsright now. A lot of the things that you can, like, learn from this are applicable to workflows as well.

So, um, yeah, we'll talk about this. Uh, wait, show of hands. How many people have, like, designed an agent loop before? Okay, cool. Okay, great, great. Um, so, yeah, I mean, I think the number one thing, the meta-learning for designing an agent loop to me is just to read the transcripts over and over again.

Like, every time you see the agent run, just read it and figure out, like, "Hey, what is it doing? Why is it doing this? Can I, uh, help it out somehow?" Right? Um, and, uh, we'll do some of that later,right?

So we'll, uh, we'll build an agent loop. Um, but here is the, uh, the three parts to an agent loop,right? So, uh, first, it's gather context,right? Second is taking action, and the third is verifying your work,right? And, uh, this is, like, not the only way to build an agent, but I think a pretty good way to think about it.

Um, gathering context is, uh, like, you know, for Claude Code, it's greping and finding the files needed,right? Um, you know, for an email agent, it's, like, finding the relevant emails,right? Um, and so these are all, like, pretty, um, yeah, like, I think thinking about how it finds this context is very important.

Agent Loop22:43

Thariq Shihipar23:03

And I think a lot of people sort of, uh, skip this step or, like, underthink it. This can be, like, very, very important. Uh, and then taking action, um, how does it, like, do its work? Uh, does it have theright tools to do it?

Like, code generation, uh, Bash, these are more flexible ways of taking action,right? And then verification is another really important step. And so the basically what I'd sayright now is, like, if you're thinking about building an agent, think about, like, can you verify its work,right?

And if you can verify its work, it's, like, a great, like, candidate for an agent. If you can't verify its work, like, it's, like, you know, coding, you can verify by linting,right? And you can at least make sure it compiles.

So that's great. Uh, if you're doing, let's say, deep research, for example, it's actually a lot harder to verify your work. One way you can do it is by citing sources,right? So that's, like, a step in verification. But obviously research is less verifiable than code in some ways,right?

Because, like, code has a compile step,right? You can also, like, execute it and see what it does,right? So, um, I think, like, thinking on, you know, like, as we build agents, the ones that are closest to being very general are the ones with a verification step that is very strong,right?

So, uh, I think there was a question here? Yeah.

Guest 224:19

So when where do you generate the plan of the work you need to run?

Thariq Shihipar24:25

Hmm. Yeah, I mean, you might.

Guest 224:29

What's the question?

Thariq Shihipar24:30

Oh, yeah, sorry. The question was, when do you generate a plan, um, before you run through it? So, um, like, in Claude Code, you don't always generate a plan. Uh, but if you want to, you'd insert it between the gathering context and taking action step,right?

And so, um, plans sort of help the agent think through step by step, but they add some latency,right? And so there is, like, some trade-off there. Um, but, yeah, the Agent SDK helps you, like, do some planning as well.

So, yeah. Yeah.

Guest 225:00

Can you, like, make the agent create that to-do list for, like, 100% sure that it will create that to-do list and run by it?

Thariq Shihipar25:13

Uh, yeah, so the question was, will the agent create the to-do list? Uh, yes. Um, if you're using the Agent SDK, we have, like, some to-do tools that come with it. And so it will, like, maintain and check off to-dos, and you can display them as you go.

So, yeah. Um, any other questions about thisright now? Okay, cool. Okay, so I'm going to quickly talk about, like, like, how do you do this stuff? You, like, what are your tools for doing it,right? And, uh, there are three things you can do.

Tool Selection25:45

Thariq Shihipar25:45

You have tools, Bash, and code generation,right? And I think traditionally, I think a lot of people are only thinking about tools. And, uh, yeah, basically one of the call to actions is just figuring out, like, thinking about it more broadly,right?

So tools are extremely structured and very, very reliable,right? Like, if you want to sort of have as fast an output as possible with minimal errors, uh, minimal retries, uh, tools are great. Uh, cons, there are high context usage.

If anyone's built an agent with, like, 50 or 100 tools,right? Like, they take up a lot of context and the model kind of gets a little bit confused,right? Um, there's no, like, sort of discoverability of the tools. Um, and they're not composable,right?

And I say tools in the sense of, like, if you're using, you know, a messages or completion APIright now, um, that's how the tools work. Of course, like, you know, there's, like, code mode and programmatic tool calling, so you can sort of blend some of these.

Um, then there's Bash. So Bash is very composable,right? Like, uh, static scripts, low context usage. Uh, it can take a little bit more discovery time. Like, let's say that you have, whatever, you have, like, the Playwright MCP or something like that.

Um, or sorry, the Playwright CLI, the Playwright, like, Bash tool. Um, you can do Playwright --help to figure out all the things you can do, but the agent needs to do that every time,right? So it needs to, like, discover what it can do.

Um, which is kind of powerful that it helps take away some of the high context usage, but add some latency. Um, there might be slightly lower call rates, you know, just because, like, it has a little bit more time to, um, or it needs to, like, find the tools and what it can do.

Um, but this will definitely, like, improve as it goes. And then finally, Code Gen, highly composable, dynamic scripts. Um, they take the longest to execute,right? So they'd need linting, possibly compilation. API design becomes, like, a very, very interesting step here,right?

And I'll talk more about, like, uh, best, like, how to think about API design in an agent. Um, but, yeah, I think this is, like, how you, like, the three tools you have. And so, yeah, using tools, I think you still want some tools, but you want to think about them as atomic actions.

Your agent usually needs to execute in sequence, and you need a lot of control over,right? So, for example, in Claude Code, we don't use Bash to write a file. We have a write file tool,right? Because we want the user to be able to sort of see the output and approve it.

And, um, we're not really composing write file with other things,right? It's, like, a very atomic action. Um, sending an email is another example. Like, any sort of, like, non-destructive, like, destructible or sort of, like, you know, uh, unreversible change is definitely, like, a tool is a good place for that.

Um, then we've got Bash. Uh, so, for example, there are, like, uh, composable actions, like searching a folder, using GitHub, linting code and checking for errors or memory. Um, and so, yeah, you can write files to memory, and that can be your Bash, like, Bash can be your memory system, for example,right?

So, um, and then finally you've got code generation,right? So if you're trying to do this, like, highly dynamic, very flexible logic, composing APIs, uh, like, you're doing data analysis or deep research or, like, reusing patterns. And so, um, yeah, we'll talk more about, uh, code generation in a bit.

Um, any questions so far about, like, the SDK loop or tools versus Bash versus Code Gen? Yeah?

Guest 229:18

Yeah, uh, I was going to ask, um, do you have, are you going to have any ready-made tools for, like, offloading tool called results?

Thariq Shihipar29:27

Offloading tools called results, like, into the file system or?

Guest 229:30

Like, let's say it goes to Bash and then the context explodes.

Thariq Shihipar29:33

Hmm.

Guest 229:33

Is it, like, tied to a command that, like, threw everything up?

Thariq Shihipar29:37

Okay.

Guest 229:37

Or, or otherwise just, like, long outputs in your history?

Thariq Shihipar29:41

Sure, yeah, yeah, yeah.

Guest 229:42

How much of, like, all the time just offloading them to files?

Thariq Shihipar29:45

Yeah, yeah. I, I think that's a good common practice. I think, um, we, I, I remember seeing some PRs about this very recently on Claude Code about handling very long outputs. And I, I, I don't know exactly. Like, I, I think, I think we are moving towards a place where more and more things are being, like, just stored in the file system.

And this is, like, a good example. Yeah, like, storing, like, long outputs, uh, over time. Um, I think, like, generally, prompting the agent to do this is a good, uh, way to think about it. Or even if you have, I think, like, something I just do always now is, like, whenever I have a tool call, I, um, I save it, like, the results of the tool call to the file system so that you can, like, search process it and then have the tool call return the path of the result.

Um, just because, like, that helps it, like, sort of recheck its work. So, um, yes.

Guest 230:45

Um, do you find that you need to use, like, the skills, um, kind of structure to help Claude along to use the Bash better or out of the box, you know, that's not necessary?

Thariq Shihipar30:59

Yeah, so the question was about skills and, like, do we need skills to use Bash better? Um, yeah, for context skills, maybe I can learn skills. Okay, yeah, skills are basically a way of, like, uh, you know, allowing our agent to take longer complex tasks and, like, sort of load in things via context,right?

So some, like, for example, we have, uh, a bunch of DockX skills. And these DockX skills tell it how to do code generation to generate these files,right? And so, um, yeah, I think overall skills are, yeah, basically just a collection of files.

They're also sort of, like, an example of being very, like, file system or Bash tool build,right? Um, because they're really just folders that your agent can, like, CD into and, like, read,right? Um, and so, yeah, they give, like, what we found the skills are really good for is pretty, like, repeatable instructions that need a lot of expertise in them.

Uh, like, for example, we released a front-end design skill recently that I really, really like. And, um, it's really just sort of a very detailed and good prompt on how to do front-end design. Uh, but it comes from, like, our best, you know, like, uh, AI front-end engineer, you know what I mean?

And he, like, really put a lot of top thought and iteration to it. So that's one way of using skills. Um, yeah.

Guest 232:27

Quick question.

Thariq Shihipar32:28

Yes?

Guest 232:28

Why use that front-end skill?

Thariq Shihipar32:30

Sure.

Guest 232:31

And it's pretty cool. Thanks for, uh, publishing it. Uh, I want to understand, uh, there are multiple Claude MD plays, like Claude MD is also there, and it is also at the user level and at the project level.

And then there are skilled Dock MD plays. Like, is there, like, a priority order? Should some stuff be relegated to Claude Dock MD and some other stuff should only come to skilled Dock MD?

Thariq Shihipar32:54

Hmm. So the question was about skilled Dock MD versus Claude Dock MD and how to think about, uh, that,right? And, uh, I think, like, I will say all of these concepts are so new, you know what I mean?

Like, even Claude Code is, like, released at, like, eight or nine months ago,right? Like, um, and so skills were released like two weeks ago. Like, I, like, I won't pretend to know all of the best practices for, for everything,right?

Um, I think generally skills are a form of progressive context disclosure. And that's sort of a pattern that we've talked about a bunch,right? Like, with, like, uh, Bash and, you know, like, preferring that over, like, you know, purely, like, normal tool calls is, like, it's a way of, like, the agent being like, "Okay, I need to do this.

Let me find out how to do this and then let me read in the skilled Dock MD,"right? So you ask it to make a DockX file and then it, like, CDs into the directory, reads how to do it, writes some scripts, and keeps going.

So, um, yeah, I, I think, like, there's still some intuition to build around, like, what, what exactly you, like, define as a skill and how you split it out. Um, but, uh, yeah, I think, uh, yeah, lots of best practices to learn there still.

Um, yeah.

Guest 234:08

Uh, so yesterday, uh, we talked about the future of skills and how they evolve over time. Do you see these as ultimately becoming part of the model and you need less of the skills? Is this just a way to bridge the gap for now?

Thariq Shihipar34:23

Yeah, so the question was, are skills ultimately part of the model? Um, are they a way to bridge the gap? I missed Barry's talk at Barry and Madge's talk yesterday, but, uh, yeah, I think roughly the idea is that the model will get better and better at doing a wide variety of tasks and skills are the best way to give it out of distribution tasks,right?

Um, but I, I would broadly say that, like, it's really, really hard, especially, like, you know, if you're, like, uh, not at a lab to, like, tell where the models are going exactly. Um, my general rule of thumb is, like, I try and, like, rethink or rewrite my, like, agent code, like, every six months, uh, just because I'm like, uh, things have probably changed enough that I've, like, baked in some assumptions here.

And so, like, I think that, like, our agent SDK is built to as much as possible sort of advance with capabilities,right? Like, the Bash tool will get better and better. Uh, we're building it on top of Claude Code, so as Claude Code evolves, you'll get those wins out of the gate.

Um, but at the same time, like, you know, things are so different now, like, than they were a year ago in, in terms of, like, AI engineering,right? And I think, like, a general best practice to me is sort of like, "Hey, we can write code 10 times faster.

You should throw out code 10 times faster as well." Um, and I think thinking about, like, not so, like, hedging your bets on, like, where is the futureright now, but, like, what can we do today that really works,right?

And, like, like, let's get market share today and not be afraid to throw out code later. Um, if you're a startup, this is arguably your largest advantage that you have over competitors. They're, like, you know, larger companies have, like, six-month incubation cycles.

And so they're always, like, stuck in the past of, like, the agent capabilities,right? And so your advantage is that you can, like, be like, "Hey, the agent, the capabilities are hereright now. Let me build something that uses thisright now,"right?

So, um, yeah, uh, any, any other questions on for we're talking about skills and Bash. Okay, it seems like there are a lot of skill questions. So, um, yeah, uh, I, I think at the back someone you might have to shout.

Guest 236:40

Yeah, so why would you use a skill versus an API? They look very similar to, like, you could, like, Python program there could be a package,right?

Thariq Shihipar36:49

Yeah, so the question was why use a skill versus an API? Um, good question. I, I think that, like, um, when you, like, these are all forms of progressive disclosure, basically, to the agent to figure out what it needs to do.

Um, and I'll go over, like, uh, examples of, like, you just have an API,right, in, in our, like, in our prototyping session. Um, it's totally, like, use case dependent,right? Like, just, I, I think, like, I don't have a, like, I don't think there's a general rule.

I think it's like, read the transcript and see what your agent wants. If your agent always wants, like, thinks about the API better as, like, an API.ts file or something or API.py file, do that. You know, that's great.

Like, I think skills are, like, a, like, sort of an introduction into, like, thinking about the file system as a way of storing context,right? And they're a great abstraction. Um, but there are many ways to use the file system.

Um, and I, I should say that, like, something about skills is that, like, you need the Bash tool, you need a virtual file system, things like that. So the agent SDK is, like, basically the only way to really use skills to, like, their full extentright now.

So, um, yeah, yeah, back there.

Guest 237:59

Can we expect a marketplace for skills?

Thariq Shihipar38:03

Yeah, the question was can we expect a marketplace for skills? So, um, yeah, Claude Code has a plug-in marketplace that you can also use with the agent SDK. Um, we're evolving that over time. You know, like, it was, like, a very much a V0.

Um, and by marketplace, I'm not sure if people will be charging for this exactly. It's more just like a discovery system, I think. Um, but yeah, that existsright now. You can do slash plugins in Claude Code. Um, and, and you can find some.

So, yeah. Yeah.

Guest 238:31

What's your current thinking about when you're going to reach for, like, the SDK, you know, to solve a problem?

Thariq Shihipar38:37

When, yeah, so the question is when do I use the SDK to solve a problem? Uh, if I'm building an agent, basically, I, I think that, like, um, my overall belief is that, like, for any agent, the Bash tool gives you so much power and flexibility and using the file system gives you so much power and flexibility that you can always eke out performance gains over it,right?

And so, uh, yeah, in the prototyping part of this talk, we're going to, like, look at an example with only tools and an example without, with, you know, Bash and the file system and compare those two. Um, and yeah, that's what I mean by being Bash tool build.

I'm like, I, I just, like, start from the agent SDK, you know? And I think a lot of people at Anthropic have started, like, doing that as well. So, um, of course, I, I do want to say that there are lots of times where the agent SDK is kind of annoying because you've got, like, this network sandbox container and you're like, "I hate like, I don't want to do this."

You know what I mean? Like, I want to run on my browser locally,right? Um, I totally get that. I think it's, there is, like, a real performance trade-off. Um, the way I think about it is sort of like React versus, like, jQuery.

You know, like, I, like, I, when I was coming up, I was, like, very into web dev and, like, you know, I was using jQuery and Backbone. And then React came out and it was by Facebook. And they're like, "You have to, here's JSX.

Like, we just made this up." And, and now there's a bundler,right? I'm like, "Huh, it's so annoying." Um, but, like, they generally makes the model or it makes, it made web apps more powerful,right? And I think we're sort of like the agent SDKs are like the React of agent frameworks to me because it's like we build our own stuff on top of it so you know it's real.

And all the annoying parts of it are just like things where we're annoyed about it too, but we're like, it just works. Like, you have, like, you got to do this, you know? Um, so, yeah. Uh, yeah, okay.

More, more skill questions, I guess? Yeah,right here.

Guest 240:34

Uh, I want to go back to Bash question.

Thariq Shihipar40:35

Oh, sure. Bash question. Great. I love Bash.

Guest 240:37

You have custom internal, like, Bash tools.

Thariq Shihipar40:39

Yeah.

Guest 240:39

How do you let the agent discover that or do those have to become tool as tools?

Thariq Shihipar40:44

Okay, the question is if you have custom agent Bash tools, how do you let the agent discover that? By custom Bash tools, do you mean, like, Bash scripts or?

Guest 240:51

We have, we have Bash scripts, yeah.

Thariq Shihipar40:53

Yeah. Um, yeah, so I, I think, uh, where is it? You just put it in the file system and you tell it, like, "Hey, like, here is a script." Uh, you can call it, you know, I, I'm generally thinking in the context of the Claude agent SDK where it has the file system and the Bash tools are tied together.

This is kind of an anti-pattern I see sometimes where people are like, "Oh, like, we're going to host the Bash tool in this, like, virtualized place and it's not going to interact with other parts of, like, the agent loop," you know?

And that sort of, you know, makes it hard because if you got a tool result that's saving a file, then your Bash tool can't, like, uh, read it, you know what I mean? Unless it's all in one, one container.

So does that answer your question? Like.

Guest 241:37

Yeah, kind of. I mean, like, so you're saying you just put it in, like, system prompt or something?

Thariq Shihipar41:41

Yeah, just put it in system prompt and be like, "Hey, you have access to this." I, I would, like, sort of design all my CLI scripts to have, like, a dash dash help or something so that the model can call that and then it can, like, progressively disclose, like, every, like, sub-command inside of the script.

Yeah. Uh, yeah, back there.

Guest 241:58

Yeah. So, um, like, my question is now when to reach for the agent SDK. So have you designed or rather would you recommend someone use the agent SDK to build, like, a generic chat agent as compared to, like, "Oh, you know, I'm building an agent where you have some input and the agent goes and does some stuff and finally I care about the output," as compared to, let's say, someone, like, are you using or do you foresee using the agent to build, like, the agent SDK to build, like, Claude the, the app rather than Claude Code?

Thariq Shihipar42:29

Uh, yeah, so the question is when do we reach for the agent SDK? Uh, does, um, like, uh, like, would we use the agent SDK to build Claude.ai, which is a more traditional chatbot, uh, than Claude Code? Um, I, one, I think Claude Code is, like, a very, like, the interface is not a traditional chatbot interface, but, like, the inputs and outputs are, are,right?

Like, you input code in, you, you get, like, or you input text in, you get text out and you, it takes you actions along the way. Um, you might have seen that, like, when we rolled out doc creation for Claude.ai, um, now it has the ability to spin up a file system and, like, create spreadsheets and PowerPoint files and things like that by generating code.

And so that is, like, you know, we're in the midst of sort of, like, um, like, merging our agent loops and stuff like that. But, but broadly, like, uh, like, yeah, Claude.ai will, like, is getting more and more, like, you see it with skills and the memory tool and stuff, more and more file system build,right?

So, uh, we do think this is, like, a broad thing that you can use just, just generally and happy to talk through examples. And, um, yeah, one more question and then we'll keep going. Yeah.

Guest 243:45

Uh, still trying to understand the rule of thumb on when to build a tool or use a tool, when to graph something with a script or just let the agent go wild on the Bash. Because I'll give you an example.

Let's say I need to access a database from time to time. I can use an MCP. I can wrap it in a script and I can just let the agent call an endpoint from that directly from Bash,right?

Thariq Shihipar44:13

Yeah, great question. Great question. So it still trying to grok, like, when to use tools versus Bash versus Cogen. And he gave an example, like, "Okay, I have a database. Um, I want the agent to be able to access it in some way.

What should I do? Should I create a tool that queries the database in some way? Um, should I use the Bash? Should I use Cogen,right?" These are all, these are three ways of doing it. Um, I think that they are, like, you could use any of them.

And I, I think, like, part of it is, like, I, I think, unfortunately, there's no, like, single best practice,right? This is, like, kind of a system design problem. But let's say that you want to access your Bash, your database via a tool.

You would do that if your database was very, very structured and you have to be very careful about, like, I don't know, you're accessing, like, user-sensitive information or something like that and you're like, "Hey, I, I can only take in this input and I need to, like, give this output and I have to mask everything else about the database from the agent,"right?

Obviously, that, like, sort of limits what the agent can do,right? Like, it can't write a very dynamic query,right? Um, if you're writing a full-on SQL query, I would definitely use Bash or Cogen, uh, just because when the model is writing a SQL query, it can make mistakes.

And the way it fixes it is, is its mistakes is by, like, linting or, like, running the file, looking at the output, seeing if there are errors, and then iterating on it,right? Um, and so I generally, like, if I'm building an agent today, I'm giving it as much access to my database as possible and then I'm, like, putting in guardrails,right?

Like, I'm probably limiting its, like, write access in different ways. But what I probably what I would do is, like, I would give it write access and put in specific rules and then give it feedback if it tries to do something it can't do.

You know what I mean? And, and so I know this is, like, kind of a hard problem, but I think this is the, like, set of problems for us to solve,right? Like, we built a Bash tool parser. Um, and that's a super annoying problem.

Uh, but we need to solve that in order to, like, let the agent work generally,right? And same thing with, like, database. Like, like, yes, it's quite hard to understand what is a query doing, but if you can solve that, you can let your agent work more generally over time.

So, um, yeah, I, I think thinking about it, uh, like, flexibly as much as possible and keeping tools to be, like, very, very, like, sort of atomic actions,right, that you need a lot of guarantees around. Um, yeah, one more question.

Guest 246:48

On the same thing,right?

Thariq Shihipar46:49

Yeah.

Guest 246:49

Uh, how do you ensure the role-based access controls are taken care of in this kind of case?

Thariq Shihipar46:55

How do you, uh, so the question is how do you ensure that the role-based access, uh, access controls are taken care of? Usually that's in, like, how you provision your API key or your backend service or something like that,right?

Like, um, I think that, like, probably what I do is, like, I create, like, temporary API keys. Sometimes people create proxies in between to insert the API keys. Um, if you're concerned about exfiltration of that. Um, but yeah, I would create, like, API keys for your agents that are scoped in certain ways.

And so then on the backend, you can sort of check it's, like, you know, what it's trying to do and, like, uh, if it's an agent, you can, like, give it different feedback. So, yeah. Allright. I have one more question.

Guest 247:35

Um, anything you could tell us, uh, more about the, the memory tool, the internal memory tool?

Thariq Shihipar47:41

Um, I have, I, I'm not trying to, like, keep a secret. I, I don't know exactly. Like, I haven't read the code, but I, I think it generally works on, on the file system. And so, uh.

Guest 247:52

Have you exposed it to, uh, to the, uh, agent SDK or is it already been?

Thariq Shihipar47:58

Um, I would say that, like, we've had this question a bunch. I would just use the file system in the Claude agent SDK. I would just create, like, a memories folder or something and tell it to write memories there.

Um, it's like, I, I don't know the exact implementation of the memory tool, but it does use the file system in, in, in that way. So, yeah. Um, allright. Yeah, last question on this. Yeah.

Guest 248:20

How you are managing for the Bash and the code, how you are managing the, like, reusability? Suppose the same agent is rolled out to hundreds of users and, uh, same code every time it is generating and every time it is executing.

So how can we use the reusability?

Thariq Shihipar48:37

Yeah, that's a really good question. So, uh, yeah, let's say you have two agents interacting with two different people. The question is, like, how do you think about reusability between agents or how do agents communicate,right? Um, I think, uh, this is a thing to be discovered.

I think, like, I think there's a lot of best practices and system design to be done on, like, um, because traditionally with web apps, you're serving one app to, like, a million people,right? And with agents, like with Claude Code, we serve, like, you know, a one-to-one, like, container.

When you use Claude Code on the web, it's like, it's your container,right? And so there's not a lot of, like, communication between containers. It's a very, very different paradigm. I'm not going to say that, like, I know exactly the best system design to do that,right?

And, like, I think there's a lot of best practices on, like, okay, these agents are reusing work. Um, how can we give them, like, like, cut, like, general scripts that combine together the work that they've done? How can we make them share it?

Um, I would generally think this is sort of like a tangent, but on, like, agent communication frameworks, I would say that, like, we probably don't need, like, a whole we don't I, I think this is more of a personal opinion.

I think, like, we probably don't need to reinvent, uh, like, a new communication system. There are, like, the agents are good at using the things that we have, like HTTP requests and Bash tools and API keys and, uh, named pipes and all of these things.

And so, like, probably, like, the agents are just making HTTP requests back and forth from each other, you know, using HTTP server. Um, there's a bunch of interesting work there. I've seen people make, like, a virtual forum for their agents to communicate and they, like, post topics and, like, reply and stuff like that.

Um, kind of cool. I think there's a lot of things to explore and discover there. Yeah. Okay. Um, going to keep going a little bit. How are we doing for time? Okay. It's got an hour left, I think.

Okay. Um, cool. So an example of designing an agent, uh, this is, like, yeah, let's this is not the prototyping session, but I think this is, like, will be a good sort of, like, like, we way into it.

Spreadsheet Design50:54

Thariq Shihipar50:54

Let's say we're making a spreadsheet agent. Uh, what is the best way to search a spreadsheet? What's the best way to execute code in? Like, what's the best way to take action in a spreadsheet? What is the best way to lint a spreadsheet,right?

These are all, like, really interesting things to do. Uh, I'm going to do, like, a Figma and we can go over it. Um, if someone could grab a water as well, that would be great. I, like, could really use water.

I'm, uh, yeah, yeah. Okay. Um, thanks. Uh, okay. So we're going to, um, yeah, let's, let's talk through it. Uh, or why don't you spend, like, a couple minutes yourselves thinking about this question. You have a spreadsheet agent.

You want it to be able to search. You want it to be able to, like, gather context, take action, verify its work. How would you think about it,right? So, like, just spend some time thinking through that. Take some notes or something.

Guest 251:47

That's great.

Thariq Shihipar52:57

Okay. Has everyone had a little bit of time to think about this? Does anyone want more time or want to just dive into it? Okay. Uh, what's the best way for an agent to search a spreadsheet? Realizing I have to type with one hand now.

Um,

I should figure this out because I'm going to need a typewriter. Okay. Um, the okay. Searching a spreadsheet. Uh, any, any ideas? How do you search a spreadsheet? Like, what would you do?

Guest 253:25

CSV.

Thariq Shihipar53:27

Okay. You've got a CSV. Okay. Now, like, your agent wants to, like, search the CSV. What, what does it do?

Guest 253:34

It greps it.

Thariq Shihipar53:35

It greps it. Okay. Uh, what does the grep look like?

Guest 353:38

Looks at all the headers.

Thariq Shihipar53:40

Looks at the headers. Okay.

Guest 353:41

Headers of all sheets.

Thariq Shihipar53:43

Okay. Great. Yeah. And let's say I'm looking for the revenue in 2024 or something. Um, now I've got my headers, like, uh, I'm just going to pull up a spreadsheet,right? Um, let's say that the revenue is in there's a revenue column and then there's, like, a, uh, so yeah, let's see.

Okay. So yeah, let's say it's something like this,right? Like, um, how do I get revenue in 2026,right? So this is sort of like a tabular problem,right? Like, there is revenue here and there's also 2026 here,right? So it's like a multi-dimensional step,right?

We could look at the headers. That will then give us, uh, like, if you just pull this, you'll get 100, 200, 300,right? So we need a little bit more. And, uh, any other ideas?

Guest 254:52

Okay.

Thariq Shihipar54:52

Yeah.

Guest 254:53

There's a Bash tool for it. The Awk, A-W-K, I think.

Thariq Shihipar54:57

Awk? Okay. Yeah, yeah, yeah. And what would it awk for?

Guest 255:01

Well, it depends on what you're looking for.

Thariq Shihipar55:04

Yeah. Yeah, yeah, yeah. That's a question,right? Like, what is the user looking for,right? They're probably looking for something like this, like revenue in 2026,right? Um.

Guest 255:12

Maybe use the APIs to use the Google tools to add all the numbers together or VLOOKUP or something like this,right?

Thariq Shihipar55:20

Yeah. So idea is, like, use the APIs, like use the Google APIs to, like, look it up. Um, that's great. Uh, but yeah, let's say we're working locally. We need to sort of design these APIs. Yeah.

Guest 255:30

SQLiter, .db compare the CSV directly. It works pretty well.

Thariq Shihipar55:34

Oh, interesting. Okay. Yeah, I didn't know that. That's great. So yeah, you, you use SQLite to query a CSV. Um, that's a great, like, sort of creative way of thinking about API interfaces,right? Like, um, if you can translate something into a interface that the agent knows very well, that's great,right?

And so, like, if you have a data source that you can convert it into a SQL query, then your agent really knows how to search SQL,right? So thinking about this transformation step is really, really interesting. It's a great way of, like, designing, like, an agentic search interface.

So, um, yeah, over there.

Guest 256:08

Sorry, real quick. What were the other tools? Because you used CSV for some of the stuff as well.

Thariq Shihipar56:12

Yeah.

Guest 256:12

Um, is there any kind of ranking within the tool? Is Claude smart enough to start ranking theright tool for theright job? Because that's kind of what we're talking about here isright tool for theright job.

Thariq Shihipar56:21

Yeah. Is Claude smart enough to rank theright tool for theright job? Uh, yeah, if you prompt it, you know, like, or like, I, I think this is one of those things where, like, I don't know. Let's find out.

Like, let's read the transcript. Uh, if it's not, like, how can you help it?

Guest 256:33

Or is that everything you need at the moment?

Thariq Shihipar56:35

Yeah. Just sort of, like, I, I think all of these things are, like, an intuition, you know? It's like, like, kind of like riding a horse. Not that I've ever rode a horse, but I know. Just like, I imagine it's like riding a horse.

Um, yeah, like, you, you, like, you know, you're sort of giving these signals to the horse or calming it down. You're trying to understand what it how, how do you push it faster? You know what I mean? And sort of, like, it's a very organic, like, thing,right?

Um, like, I think we like to say that models are grown and not designed,right? And so we're, like, sort of understanding their capabilities. Yeah. Uh, yeah. And where it is. Yeah.

Guest 257:13

Quick question. So is there a way to, like, metadata to the spreadsheet? Can you give descriptions? That's in a different document.

Thariq Shihipar57:19

Mm. Yeah. That's.

Guest 257:19

For example, KPIs. I'm trying to get an idea if I can get to, like, build intelligence for ask questions about a spreadsheet.

Thariq Shihipar57:26

Yeah. So that's another great pattern is, like, okay, can you add metadata to a spreadsheet? So these are some questions that you might want to think about before, like, when you're thinking about search is, like, what pre-processing can you do to make the search better,right?

And so one example is that you translate it into, like, a SQL format or something where you use something that can query it,right? That's like a translation step. Another step is, like, maybe you have a tool or, um, like a, a pre-processing step where another agent annotates the, the spreadsheet and, and, like, adds information so that the agent can then, like, search across that information better,right?

So, um, yeah, one more.

Guest 258:05

Um, I was just curious.

Thariq Shihipar58:06

Oh, yeah.

Guest 258:07

Like, what I mean, all those tools sound great, but.

Thariq Shihipar58:10

Yeah.

Guest 258:10

Why can't the agent just, you know, do what was suggested, read the header, and then just get the data from rep? Like, I feel like that should be pretty trivial for, um, for, for re-tasking.

Thariq Shihipar58:21

Yeah. Probably I should have, like, prepared this in code. Um, but yeah, I, I built a ton of spreadsheet agents before. Basically, it's.

Guest 258:29

Does that work?

Thariq Shihipar58:29

It's kind of hard to do. Yeah, yeah. So, um, basically what, what I would think about is, like, so we've got, like, okay, I Sean, do you have suggestions on how it can how I can code at the same time?

Go ahead.

Guest 258:43

Install voice to text on your computer.

Thariq Shihipar58:45

Oh, I see. Yeah, yeah, yeah. Do you work at Whisperflow or something or?

Guest 258:50

Stick the mic in your shirt.

Thariq Shihipar58:52

There's a microphone button on the back. There's a microphone button on the back.

Guest 258:57

Stick the mic in your shirt.

Thariq Shihipar58:59

Oh, I, I just don't trust that stuff, man. Okay. Um, maybe I have to maybe I shouldn't be working in an AI lab. Um, okay. So, uh, let's see.

Guest 259:13

Make someone hold it for you.

Thariq Shihipar59:14

Make someone hold it for you. Hold on. Hold on. Okay. Um, like, search. So

one way to do it is, like, you, you see in spreadsheets,right? Like, you can say here, you can design formulas,right? So, like, B32, um,right? So this is a syntax, for example, that the agent's pretty familiar with, like B3 to B5,right?

And so you can design an agentic search interface, which is like this,right? Like B3, B5 or something,right? So, like, your agentic search interface can take in a range,right? Can take in a range string,right? And these are things that, like, the agent knows pretty well,right?

Like, you can, um, do SQL queries,right? The agent knows SQL queries pretty well,right? Um, and, uh, like, these you can also, uh, do XML,right? Sorry, the font is so small. Um,

okay. Uh, yeah, you can also do XML. I, I, I'm not sure if you guys know, but, like, uh, XLX files are XML in the backend,right? And XML is very structured. Uh, you can do, like, an XML search query.

Uh, and there are different libraries that can do that. So that's one example,right? It's like, how do you search and gather context? And I hope this sort of, like, illustrates to you that, like, gathering context is really, really creative,right?

Like, and, and, like, there's so many iterations. And if you just if you've only tried one iteration, it's probably not enough,right? Like, think about, like, as many different ways as you can. Like, try these out,right? Like, try SQL, try, try the search, try, try the grep and awk and, like, all of these things.

And, um, have a few tests that you're trying across different things and, and see what the agent likes and what it, what it doesn't like. Um, it's going to be different for each case.

Guest 21:01:09

Sorry.

Thariq Shihipar1:01:09

Yeah.

Guest 21:01:11

When you say agent, you're referring to Claude, the, the model or because we're building an agent here.

Thariq Shihipar1:01:18

Yeah.

Guest 21:01:19

And you're relying on already pre-existing knowledge of how you handle XML. Who's, who's doing that to model?

Thariq Shihipar1:01:26

Yeah. Because the question is, like, who, uh, where does the knowledge come from? Is it the model? Is it, like, what do what do I mean by the agent? Yeah. Generally, I think what you're looking for is, like, you have a problem.

You want to make it as in-distribution as possible for the agent,right? And so the agent knows a lot about a lot of different things. It knows a lot about, for example, finance,right? So if you ask it to make a DCF model, it knows what DCF is,right?

And you can if, if you want to give it more information, you can make a skill,right? But so it knows what DCF is. It knows what SQL is. Can it combine those things together,right? And so, like, uh, ideally, you want to, like, your, your problem is going to be out of distribution in some way,right?

Like, like, there's some, like, information that's not on the internet or something that you have, uh, or something somewhat unique to you, and you want to try and, like, massage it to be as in-distribution as possible. Um, and, uh, yeah, it's, it's very, very creative, I think.

Like, uh, you know, it's not like a it's not a science to me. It's very much like an art. So, um, yeah. Okay. So we've tried gathering context, then taking action. Um, we can probably do a lot of the same stuff here that we've done before,right?

Like, we can do, like, insert 2D array,right? Um, if we've got, like, a SQL interface,right, we can, um, we can do a SQL query. We can edit XML. Um, these are, like, often very similar,right? Like, taking action and gathering context.

You probably want a similar API back and forth. And then the last thing is verifying work,right? Like, how do you think about how do you think about that? Um, check for null pointers,right, is one of the ways to do it.

Um, any other ideas on, on verification or yeah.

Guest 31:03:23

Sorry. I'm, I'm a bit confused.

Thariq Shihipar1:03:25

Oh, yeah.

Guest 31:03:26

What did you say?

Thariq Shihipar1:03:27

Yeah, yeah.

Guest 31:03:28

Like, when, when you're using other SDKs to build the agent.

Thariq Shihipar1:03:32

Yeah.

Guest 31:03:32

I don't need to tell, like, how it should gather the context.

Thariq Shihipar1:03:35

Sure.

Guest 31:03:35

I just give it the context and explain.

Thariq Shihipar1:03:37

Yeah.

Guest 31:03:38

This is what like, basically I explain in plain English.

Thariq Shihipar1:03:40

Yeah.

Guest 31:03:41

What it's meant to do.

Thariq Shihipar1:03:42

Yeah.

Guest 31:03:43

And what I tend to do, and you tell me if I'm wrong, I actually end up creating a separate agent for QA.

Thariq Shihipar1:03:50

Oh, interesting.

Guest 31:03:51

To, to verify because I don't trust the agent to verify itself.

Thariq Shihipar1:03:55

Mm.

Guest 31:03:56

But I'm just I'm, I'm just a bit I'm confused about the level of detail I need to provide the agent in that example.

Thariq Shihipar1:04:04

Yeah. Okay. So the question is about, um, giving context to the agent versus having it gather its own context. Uh, you mentioned that you sometimes use a Q&A agent. Uh, can I ask, like, what, like, domain you, you're building your agent in or?

Guest 31:04:19

In, uh, cybersecurity.

Thariq Shihipar1:04:21

Okay. Sure. Yeah, yeah, yeah. Um, I think that I, I think I need to, like, look into more specifics, but the Claude Agent SDK is great for cybersecurity. And, like, I would generally push people on, like, let the agent gather context as much as possible, you know, like, let it find its own work as much as possible.

Um, you're trying to give it the tools to find its own work. The way I think about this is kind of like, let's say that someone locked you in a room and they were, they were, like, giving you tasks, you know, like, so that's what your job was, like a Mr.

Beast sort of, like, scenario,right? Like, you get $500,000 if you stay in this room for six months. Um, then, like, like, someone's giving you a message. What tools would you want to be able to do it,right? Like, would you just want, like, a list of papers or, like, would you want a calculator or, like, a computer,right?

Probably I would want a computer,right? I'd want Google. I'd want, like, all of these things,right? And so, like, I wouldn't want the person to send me, like, a stack of papers being like, "Hey, this is probably all the information you need."

I'd rather just be like, "Hey, just give me a computer. Give me the problem. Let me search it and figure it out,"right? And so that's how I think about agents as well. Like, they need, like, like, you know, they're stuck in a room.

Guest 31:05:38

I need to get them tools. So if you can go back to the slides you had, to the graph you had.

Thariq Shihipar1:05:44

To the graph? Like, like this you mean or?

Guest 31:05:47

Yeah, this topic. So basically.

Thariq Shihipar1:05:48

Yeah.

Guest 31:05:48

That gathering context is basically these are the tools I'm offering it.

Thariq Shihipar1:05:53

Yeah, exactly. Yeah. You, you're I'm giving it, like, maybe an API for code generation. Maybe I'm giving it a SQL tool. Maybe I'm giving it Bash. These are all, like, examples,right? So yeah. Do you have one more question?

Guest 21:06:05

Question. So, uh, for all the agents that you're, uh, having the certainty.

Thariq Shihipar1:06:09

Yeah.

Guest 21:06:09

Next to the, the you share the same context window and it says size?

Thariq Shihipar1:06:14

Interesting. Yeah. So do agents share the context window? I think, I think this is, like, an interesting question just overall about how you manage context. Um, I think and I haven't talked about this too much yet, but sub-agents are, like, a very, very important way of managing context.

Um, I think that this is, like, we're using more and more sub-agents inside of Claude Code. And I would think about, like, doing sub-agents very generally. So, like, what we might do for the spreadsheet agent is maybe we have a search sub-agent,right?

So, like, sub-agents are great for when you need to do a lot of work and return an answer to the main agent. So for search, let's say the question is, like, how do I find my revenue in 2026?

Maybe you need to do a bunch of results. Maybe you need to, like, uh, search the internet. Maybe you need to search the spreadsheet, things like that. And there's a bunch of things that don't need to go into the context of the main agent.

The main agent just needs to see the follow result,right? And so that's a great sub-agent task. Um, I don't have a dedicated sub-agent slide here, but, like, yeah, they're very, very useful and I, I think a great way to think about things.

Um, yeah.

Guest 21:07:21

And just to, just to build on that question, actually.

Thariq Shihipar1:07:23

Yeah.

Guest 21:07:24

For verification, for example, you could imagine doing that through a skill or a sub-agent. You might even want to have an adversarial, like, the security example is a great one. You want to have it really go to town on it and not really have any sympathetic relationship with the work already done.

Uh, it's a very I get it's a spectrum, but do you, like are you saying yes, you'd use a sub-agent here? You'd use a skill? How would you think about this?

Thariq Shihipar1:07:46

Yeah, definitely. So question on, like, uh, do sub-agents or oh.

Guest 21:07:50

Yeah, sure. It'll work just to make sure.

Thariq Shihipar1:07:52

Oh, sure. Okay. Yeah, yeah, yeah. Thank you. Appreciate it. Um, okay. Yeah. Uh, can you use sub-agents for verification? Uh, yes. I, I think this is a pattern. I think, like, ideally, the, the best form of verification is rule-based,right?

You're like, is there, like, a null pointer or something? Uh, that's, like, easy verification. Does it lint or compile? Like, like, as many rules as you can, try and insert them. And again, be creative,right? Like, for example, uh, in Claude Code, if the agent tries to write to a file that we know it hasn't read yet, like, we haven't seen the we haven't seen it enter the read cache, we throw it an error.

We, we tell it, like, "Hey, uh, you haven't read this file yet. Try reading it first,"right? And that's an example of sort of, like, a deterministic tool that we insert into the verification step. And so as much as possible, like, anytime you are thinking about, you know, verification, first step is, like, what can you do deterministically?

What, like, what, like, you know, outputs can you do? And again, like, when you're choosing which a like, types of agents to make, the agents that have more deterministic rules are better, you know? Like, they just, like, like, it, it just makes a lot of sense,right?

So, um, of course, as the models get better and better at reasoning, then you can have these sub-agents that check the work of the main agent. The main thing there is to, like, avoid, uh, context pollution. So you probably wouldn't want to, like, fork the context.

You'd probably want to start a new context session and just be like, "Hey, yeah, adversarially check, um, the work of, like, this, this output was made by a junior analyst at McKinsey or something. They graduated from, uh, like, not a great school.

Like, their GPA, like, you know, like, like, just, like, feed it a bunch of stuff and then tell it to critique it,right? Like, that's, like, one of the tools of the sub-agent,right? And so, um, yeah, the more you, like, uh yeah, as the models get better and better, that sort of verification will become better as well.

Um, but doing it deterministically is, like, a great start. Yeah. Question?

Guest 21:09:57

Just a question about the verify work.

Thariq Shihipar1:10:00

Yeah.

Guest 21:10:00

So, um, so let's say we found null pointers. It's probably easy to just say, "Okay, fix it." But, like, you know, let's say we deployed to production and the client is using it. That's not us. And they somehow get into a spot where the whole spreadsheet's deleted.

And so, like, like, on what level do we need to bake in, like, the ability to, like, undo tools and stuff? Because, like, um, let's say the QA agent returns that their spreadsheet is empty.

Thariq Shihipar1:10:31

Yeah.

Guest 21:10:32

Not necessarily is able to undo work. So, like, you know, like, what was your advice there?

Thariq Shihipar1:10:37

Yeah. So the question is, like, how do you think about state and, like, undoing and redoing, being able to, um, fix errors basically,right? I think this is, like, uh, a really good question. And honestly, another sort of, like, um, like, when you think about, like, what are agents good at,right?

Like, or what problem domains are agents good at? How reversible is the work is, like, a really good intuition,right? So code is quite reversible. You can just, like, go back. You can undo the git history. We, we come with, like, you know, these atomic operationsright out of the gate,right?

Like, I use git constantly through Claude Code. I, I don't type k commands anymore,right? So, um, that's, like, a really good example. A really bad example is computer use, you know, because computer use has is not reversible in state,right?

Like, let's say you go to, like, doordash.com and you add, like, the user wants you to order a Coke and you add order a Pepsi. Now, like, you can't just go back and click on the Coke. You have to, like, go to the cart and you have to remove the Pepsi,right?

And so your mistake has, like, compounded this, like, you know, this state and the state machine has gotten more complex,right? And, and so, like, whenever you're dealing with, like, very, very complex state machines that you can't undo or redo of, it does become harder,right?

And I think one of the questions for you as an engineer is, like, can you turn this into a reversible state machine? Kind of like you said. Can you store state between checkpoints such that the user can be like, "Oh, my spreadsheet is messed upright now.

Just go back to the previous, uh, checkpoint,"right? Uh, potentially even can the model go back to previous checkpoints. Um, I, I think someone had this, like, time travel tool, um, that they were giving one of the coding agents, which was kind of cool, where you're like, it's like, you can time travel back to a point before this happened.

You know what I mean? Uh, it's kind of fun. I, I think, like, all of these tools, some of them don't work that well yet, but, you know, we'll, we'll get there. Um, yeah, thinking about state and verification is, is very useful,right?

So, um, yeah. Question at the back?

Guest 21:12:44

Yeah. Um.

Thariq Shihipar1:12:45

Yeah.

Guest 21:12:46

I'm kind of curious about scale. Um, so what if the spreadsheet is, like, millions of rows and million and, and thousand hundreds of thousands of columns,right? Uh, or just, like, any sort of database. Like, in that type of situation, how would you go about searching?

There's obviously a context limit you have to content for.

Thariq Shihipar1:13:06

Yeah, this is great. Um, I probably should have done the spreadsheet example as my coding example. For, for a preview, my coding, like, agent is a Pokémon agent. Um, probably spreadsheet would have been better. Okay. Uh, the question was, what if the spreadsheet is very big?

If you have a million rows, uh, how do you think about?

Guest 21:13:26

In 100 columns. I mean, like, 100,000.

Thariq Shihipar1:13:28

100,000, yeah, 100,000 columns or 100 columns or whatever. Like, how do you think about it,right? Like, your database is also very big. Like, how do you, how do you do that? Um, I think for all of these things, uh, one, of course, as the data becomes larger and larger, it's just a harder problem.

Like, you know, it, it just absolutely is. Your accuracy will go down,right? Like, Claude Code is worse in larger code bases than it is in smaller code bases,right? As, as the models get better, they will get better at all of that.

Um, for all of these, I would think about, like, how would I do this? If I had a spreadsheet that was, like, a million columns and a million rows, what would I do? I, I mean, I would need to start searching for it,right?

I would need to be like, like, if I'm searching for revenue, I'd be like searching Control F revenue, and then I'd go check each of these, like, results, and I'd be like, is thisright? And then, like, I'd see, like, hey, is there a number here?

And then I'd probably keep a scratch pad, like a new sheet where I'm like, hey, like, equals revenue equals this, you know, and, and, and store this reference and, and keep going. So I, I think that's a good way of thinking about it is, like, the model shouldn't you should never, like, read the entire spreadsheet into context because it would, it would take too much,right?

Like, um, you want to give it, like, the starting amount of context. And that's also how you work,right? Like, let's say that you open up the spreadsheet. What you see is rows is this,right? You see, like, the first 10 rows and the first, like, you know, 30 columns or something,right?

That's what you see. You don't load all of it into contextright away. You probably have an intuition for, like, hey, I should load more of this into context,right? And, and, like, oh, I should navigate to this other sheet,right?

And this other sheet has more data,right? Um, but you need to, like, sort of you gather context yourself,right? And so the agent can operate in the same way. It can, like, navigate to these sheets, read them, like, try and, like, keep a scratch pad, keep some notes, and keep going.

So that's how I would think about it. Uh, yeah, at the back?

Guest 21:15:24

Yeah. So my question is about managing context pollution. It actually, I guess, relates to the previous question. Um, do you have a rule of thumb for, you know, what fraction of the context window do you use before you start hitting diminishing returns or just it becomes less effective?

Thariq Shihipar1:15:40

Mm-hmm. Yeah. The question is, yeah, context management. Do you have a rule of thumb for, like, uh, how much of the context window to use before it becomes less effective? This is actually, I'd say, a pretty interesting problemright now.

Um, I think a lot of times when I talk to people who are using Claude Code, they're like, "I'm on my fifth compact." I'm like, "What?" Like, like, I've I, like, almost have never done a compact before. You know what I mean?

Like, I have to, like, test the UX myself by, like, like, forcing myself to get compacted. Um, just because, like, I, I tend to, like, clear the context window very often,right, when I'm using Claude Code myself, just because, like, um, at least in, in Code, the state is in the, the files of the code base,right?

So let's say that I've made some changes. Uh, Claude Code can just look at my git diff and be like, "Oh, hey, these are the changes you've made." It doesn't need to know, like, my entire chat history with it, you know, in order to continue a new task,right?

And so in Claude Code, I clear the context very, very often. And I'm like, "Hey, look at my outstanding git changes. I'm working on this. Can you help me extend it in this way?" Right? That's, like, a way of thinking about it.

And, um, when you're building your own agent, like, let's say we're building a spreadsheet agent, it gets a little bit more complex 'cause your users are less technical,right? And they don't know what a context window is,right? Um, that is, like, I'd say a hard problem.

I think there's, like, some UX design there of, like, can you reset the conversation state,right? Like, can you maybe every time the user asks a new question, can you do your own compact or something? And can you, like, uh, summarize the context?

Um, does it like, in a spreadsheet, a lot of the state is in the spreadsheet itself, so it probably doesn't need, you know, to know the entire context. Um, can you store user preferences, um, as it goes so that you remember some of this stuff?

You know, like, there's a lot of, like again, like, it's an art. There's, like, so many different angles and ways in which you can do this,right? Um, but yeah, you are trying to, like, sort of minimize context usage.

Um, you probably don't need Sonnet Million Context or something. You know what I mean? Like, you just need good context management, like UX design. Yeah. Um, yeah.

Guest 21:17:53

Uh, just, I just wanted to ask. The sub-agents were made to protect the conduct of the core agent,right?

Thariq Shihipar1:17:59

That'sright. Yeah. Sub-agents were made to protect the core.

Guest 21:18:01

Would the spreadsheet will we be able to use multiple sub-agents and try to make a process where we chunk up the spreadsheet in the case where it's super large, so then the agents can kind of run through each portion, like, in parallel of each other?

Thariq Shihipar1:18:11

Yeah. Yeah. I mean, um, yeah. So, like, one of the things I love about Claude Code is that we are, like, the best experience for using sub-agents. Like, especially sub-agents with Bash. It is very, very good. I didn't really quite realize, uh, all the pain.

Um, I think if anyone's going to KubeCon, I believe Adam Wolfe is giving a talk on KubeCon about how we did the Bash tool. Adam's a legend and the Bash tool, they did such a good job. Um, when you're running parallel sub-agents at the same time, Bash becomes, like, very complex and there are lots of, like, like, race conditions and stuff like that.

And, and so there's a lot of work that we solved there,right? So this is, like, one of the things I love about Claude Code is you can just be like, "Hey, like, spin up three sub-agents to do this task," and it will do that.

And in the Agent SDK as well, you, you can just ask it to do that. So number one, sub-agents are a great primitive in the Agent SDK, and I haven't seen anyone do it as well. So that's, like, a big reason to use it.

Um, yes, generally you want it you want these sub-agents to preserve context. Let's say you have if you have a spreadsheet, you could potentially have multiple read sub-agents going on at the same time,right? So maybe the main agent is like, "Hey, can this agent read and summarize sheet one?

Can this agent read and summarize sheet two? Can this agent summarize sheet three?" And then they return their results, and then the agent maybe spins up more sub-agents again,right? So this is, like, another knob you have. Um, and I, I think what I want to say is, like, there's like, we've talked so many about so much about, like, all these different creative ways that you can, like, do things.

This is, like, the level at which you should think about should have to think about your problem. You should not really, in my opinion, think about, like, uh, like, how, like, how do I spin off a process to make a sub-agent or, like, you know, like, the system engineering between, like, uh, behind, like, what is a compact or something,right?

So, like, we take care of all of this for you in the harness so that you can think about, like, hey, what sub-agents do I need to spin off,right? And, like, how do I create a, a, agentic search interface?

And how do I, like, verify it's work? These are the really core and hard problems that you have to solve. And anytime you spend not solving these problems and, and solving, like, lower level problems, you're probably not delivering value to your users, you know?

And, and so, um, yeah, I, I think sub-agents big fan of the Agent SDK sub-agents. Yeah. Uh, yeah. Question?

Guest 21:20:35

So, uh, like, we have this, uh, uh, checksum and the verification task.

Thariq Shihipar1:20:40

Yeah.

Guest 21:20:40

So where exactly do we need to put the verification? In this example, I let's say after generation of the SQL query.

Thariq Shihipar1:20:46

Yeah.

Guest 21:20:46

I can verify it is theright query is generated or not. That is the one path. Second path is, like, generation the query directly executing, and once I will get the output, then I will, uh, do the verification. So, uh, and how do how agent can choose dynamically, like, which one is theright path?

Thariq Shihipar1:21:04

Yeah. So the question is, like, where do you do verification? Uh, is it only at the end? Do you do it in the middle? Like, things like that. I would say, like, everywhere you can. Just, like, constantly verifi verification,right?

Like, uh, like I said, we do some verification in the read step of the of Claude Code,right? So that's, like, a great example. Um, you can do it at the end. You should absolutely do it at the end.

But at any other point, if you have rules or heuristics especially, uh, like, if, for example, you're like, "Hey, one of my rules is that you shouldn't do, like, the, the total number of columns you should searches should be under 10,000 or under 1,000 or something," that's, like, a, a nice way of doing it,right?

Like, similarly here, like, maybe you shouldn't be inserting, like, a huge, like, row, like, of, of values. Like, give feedback to the model. Be like, "Hey, chunk this up,"right? You throw an error and give it feedback. And the great thing about the model is, like, it listens to feedback.

It will read the error outputs,right? And then it'll just keep going. So yeah, verification is definitely, like, I, I know I have it in this, like, as a sort of a loop, but, um, it's definitely more like, verification can happen anywhere and, and should happen anywhere.

Like, like, put it in as many places as you can. So, um, allright. I do need to start doing some of the prototyping, but I'll, I'll take one more question. Soright here. Yeah.

Guest 21:22:21

How do we say how do we form the steps? I mean, like, how do we say the agent that go search first and then.

Thariq Shihipar1:22:26

Yeah.

Guest 21:22:27

Do this step and then do that step? How does the loop actually start from the start point to the end? How do we actually design it?

Thariq Shihipar1:22:33

You just tell it. So, like, uh.

Guest 21:22:35

Like, like, is it in a system prompt or?

Thariq Shihipar1:22:38

Yeah, in the system prompt. Yeah. So, like, with Claude Code, we just give it the Bash tool and we're like, "Hey, like, gather context, read your files, uh, do stuff, like, run your linting." You know what I mean?

Um, and so, yeah, again, with the agent, you don't need to enforce this,right? You don't need to tell it, "Hey, like, you need to do this," because, like, sometimes it might not be necessary,right? Like, let's say that someone is asking a read-only question for your spreadsheet.

You don't need to, like, verify that, uh, like, you're that there are no compile errors,right? Because there's you haven't done any write errors, write operations,right? So, um, let the agent be intelligent and, and, like, in the same way that you would like that same freedom when you're doing your work,right?

Uh, you're trapped in this box or whatever, like, same way,right? Uh, so, okay, cool. I, I, I do want to try and see if I can do some prototyping now that we have this, um, uh, the, the holder as well.

Um, okay. Yeah. Execute lint. We've done a bunch of Q&A. Okay. Prototyping. Okay. Let's say that you have an agent,right? Like, you want you want to build an agent. You come out of this talk and you're like, "Great.

I have a bunch of ideas. How, how do I do this?" Um, I think what I say overall is, like, building an agent should be simple. Your agent at the end should be simple, but simple is not the same as easy,right?

Prototyping1:23:46

Thariq Shihipar1:23:59

So, like, it should be very simple to get started. And it is. Just go to Claude Code, give Claude Code some scripts and libraries and, uh, custom Claude MDs and ask it to do it,right? And that's what we're going to do,right?

Um, that's, like, it should be so easy to be like, "Hey, this is my API. This is, like, an API key. Uh, can you, like, go search, like, you know, I don't know, like, my customer support tickets or something and organize them by priority or something like that,"right?

And then look at what Claude Code does and, and it and iterate on it,right? And this is, like, a great way of, like, just skipping to, like, the hard domain-specific problems that you have,right? So you have a lot of, like, domain problems, like, how do you organize your data, your agentic search, how do you, like, create guardrails on your database?

These are all questions that you can just start solvingright away with Claude Code,right? And so try and, like, build something that feels pretty good with Claude Code. And I think generally what I've seen is that you can do this and get really good results just out of the bat using Claude Code locally,right?

And, and you should have high conviction by the end of it,right? And so, um, yeah, I think, like I forgot this, uh, more info. Watch my AI engineer talk. Uh, this is, like, a deck for internal that we were using.

Um, okay. So, uh, yeah, I'm gonna be inserting this. So yeah, you're getting what we what we show customers,right? So, um, okay. Uh, yeah. So yeah, use, use Claude Code. Uh, again, simple, but simple is not easy,right? So, like, the amount of code in your agent should not be, like, super large.

It doesn't need to be huge. It doesn't need to be extremely complex, but it does need to be elegant. It needs to be, like, what the model wants. You want to have this interesting insight. Let's turn the, the model into a SQL query.

Oh, let's turn this spreadsheet into a SQL query and then go from there,right? So, um, think about it that way. And Claude Code is, like, a great way of doing that. So, okay. Uh, let's make a Pokemon agent,right?

This is what we're gonna do. Uh, Pokemon is a game with a lot of information. There are thousands of Pokemon, each with a ton of moves. Um, uh, we want to be pretty general. And so there is actually, like, a Poke API.

Um, and the reason I chose Pokemon is just 'cause, like, I know that you guys have your own APIs as well,right? And they're all, like, very unique,right? And, uh, so I wanted to choose something with a kind of complex API that I haven't tried before.

Um, so the Poke API has, like, you know, you can search up Pokemon, like Ditto. Uh, you can search up, like, items and things like that. Um, and so it's got this, like, yeah, this custom API. You've got, uh, everything in the games,right?

So, um, and yeah, like, one of the quest things your agent might want your user might want to do is make a Pokemon team,right? I love Pokemon. I know very little about making an interesting Pokemon team for competitive play.

Uh, could my agent help me with that? That'd be that'd be cool,right? So, um, my goal is to make an agent that can chat about Pokemon and then we will, like, you know, see what we can do,right? And, and, and how far we get.

So, um, I've done, like, some of this work already and I will, like, open up and show you. So, um, the first step and the prompt here is, like, the first step is I'm, I'm gonna do mostly code generation for this,right?

And so, um, let me.

Guest 31:27:32

Is that gonna be on GitHub somewhere?

Thariq Shihipar1:27:34

Uh, actually, it is. Uh, yeah. So on my personal GitHub. Oh, yeah. I was going to commit all of this as well.

Guest 31:27:43

Is that on your official GitHub?

Thariq Shihipar1:27:44

Yeah.

Guest 31:27:45

Nice.

Thariq Shihipar1:27:45

Um, yeah, yeah. So, uh, I think my personal GitHub is let's see. Sorry.

Guest 31:27:51

Is it secure GitHub or does it have malware in it?

Thariq Shihipar1:27:56

You, you, you guys are AI engineers. Yeah. Like, if you can get owned, that's, that's your fault. Um, yeah. So, um, yeah, you can you can clone this if you'd like. Um, I need to push the last changes.

So, okay. So, um, yeah. Can, can you guys see this? Should I put it in dark mode instead or is this fine? Like, um.

Guest 41:28:18

Dark mode.

Thariq Shihipar1:28:19

Dark mode. Okay.

Okay. Is this better?

Guest 41:28:31

Yeah.

Thariq Shihipar1:28:31

No? You want a different dark mode?

Dark hard. Okay. I think it's gonna be it's gonna be gonna get guys. Um, okay. Allright. Let's see. I how does this work? Can you guys still hear me or?

Guest 41:28:49

Yeah.

Thariq Shihipar1:28:50

Yeah. Okay. Um, okay. So here's an example of, like, I've taken the, the prompt I gave it was, "Hey, I go search Poke API for its API and create a TypeScript library,"right? And so this is all vibe coded.

Um, and so you can see here that it's created this, like, interface for Pokemon,right? And so it's created, like, this Pokemon API. I can get by name. I can list Pokemon. I can get all Pokemon. I can get species and abilities and stuff like that.

And so, like, this is just a prompt that I gave it,right? And it generated this, like, TypeScript API. It also did it for moves. Um, and then it's created this, um, like, uh, it's created this, like, API that I can use import Poke API,right, from the Poke API SDK.

And, uh, yeah, you can see, like, sort of how it's, like, set, set this up. And, uh, now in contrast,right? And, and so this is the Claude.md,right? This is the TypeScript SDK for the Poke API. Um, this is, like, the, the modules in the Poke API.

Here are some of the key features. Um, uh, I'm asking it to write scripts in the examples directory, and then it will execute those scripts to help me with my queries,right? Um, and I give it some example scripts.

It doesn't always need all this information,right? Like, uh, but yeah, fetching Pokemon, listing the resources, getting data, things like that. So this is, like, my agent really. It's like a prompt I gave it to generate a TypeScript library and then this Claude.md, and I, I can chat with it in Claude Code.

I'll also show you a version of it that is just tools,right? So here I'm using the messages completion API,right? And I've given it a bunch of tools from the API. So, like, get Pokemon, get Pokemon species, uh, get Pokemon ability, get Pokemon type, get move.

So you've defined all of these tools. And you can see that, like, you know, I also just gave it a prompt and told it to make the tools. Um, it doesn't want to make a hundred tools,right? Like, there's a ton of Smogon or sorry, um, Poke API data.

Um, but, like, it, it, you know, there's only so many parameters it can do. So it's got this, like, tool call. And now, um, and I, I made, like, a little chat interface with it,right? So let me now go here and say, like, uh, this is my tool calling.

Um.

Guest 41:31:23

Did you push to the latest?

Thariq Shihipar1:31:25

Did I? Yes.

Great. So yeah, here we've got this chat.ts,right? Um, I, I use bun when I'm prototyping stuff just 'cause, like, I don't wanna compile from TypeScript to JavaScript. Um, and, uh, again, bun is, like, linting built into it. Uh, it's a way of, like, simplifying for the agent so the agent doesn't need to remember to compile.

But TypeScript is better for generation 'cause it has types,right? So I'm gonna start this, like, bun chat and then I'm gonna try, like, okay, what are the generation two water Pokemon, um, and you'll see that it's, it's starting to, like, search.

And I'm logging all the tool calls here. This is very, very important,right? Because, like, it needs to, like, do the tool calls. And so you can see that what it's doing is, like, it's searching a bunch of Pokemon.

Um, and then it told me, okay, here are the water Pokemon for Gen Two,right? It's got Totodile, Croconoff, or Alligator. You can see sort of, like, how it's thought, like, in between each step, it's thinking through, um, the previous steps,right?

Now, like, let's say that I want to do with Claude Code. I think I might need to, uh, I might need to delete this example.

Guest 41:32:47

Sorry.

Thariq Shihipar1:32:48

Um, oh, yeah.

Guest 41:32:49

Small question. How do you log the, the tool calls? It's like a just an argument you can?

Thariq Shihipar1:32:56

Oh, yeah. This is, um, this is, like, in the normal API,right? So I just, like, uh, in the model, every time it logs it, I just call this. This is in the, like, normal Anthropic API. Um, in the SDK, I, I'll get back to get to the SDK.

Um, it's just, like, you just log every assistant message. So, um, just doing console logs, please. Does that make sense or?

Guest 41:33:22

Yeah.

Thariq Shihipar1:33:22

Okay. Yeah.

Guest 41:33:23

So, so the chat interface you're showing.

Thariq Shihipar1:33:25

Yeah.

Guest 41:33:25

Is that just using the regular API or?

Thariq Shihipar1:33:27

Yeah, that's using the regular APIs.

Guest 41:33:28

So not the agent SDK?

Thariq Shihipar1:33:29

Not the agent SDK. Yeah, yeah, yeah. And so what I'm gonna do here is, um, here, I'm gonna delete the script because I don't want it to cheat. Um, but okay. So here, you, you know that, um, I've, I'm just opening Claude Code.

I've created a bunch of files here. I'm gonna say, like, can you tell me all the generation two water Pokemon? Um, and then we'll see what it can do,right? So, um, I forget if I need to prompt it to write a script or something.

I think I'll be fine. We'll, we'll see what happens.

Guest 31:34:01

Do you mind going to the core SDK file and just showing you talked about getting context and then action and then verification? Can you show that in the code and how we're configuring the tool description?

Thariq Shihipar1:34:13

Yeah. So, uh, we haven't done the SDK part yet. So, so far, I've just put, put some APIs in Claude Code. Yeah, yeah, yeah.

Guest 31:34:24

I'm sorry. I thought I missed that.

Thariq Shihipar1:34:25

No, no, no. Yeah, yeah, yeah. Of course. Okay. Um, but yeah. So, okay. You can see here, um, it's, it's given me a lot more,right? And, um,

yeah, it's giving me a lot more. So it, it, it's, it's saying there's 20 water Pokemon,right? And I think this is roughlyright. I've, like, um, uh, what did it do?

Oh, I think it just knows. Okay.

No, that's funny. It lagged up on this. Um,

allright. Um, anyways, uh, yeah, the Pokemon is slightly in distribution, which is, which is, I, I guess, good. Um, but yeah. So, like, what, what it will do is, like, it will try and, like, write, like, a script.

And, uh, because you don't want it to think as much,right? So here it's like, okay, what I'm going to do is, um, let's see. Gen Two Water Type Pokemon. Yeah. Where is it?

Okay. So yeah, you can see here it, it knows, like, okay, the start of the generations. It fetches these, uh, per API. Um, I guess it's decided not to use, like, my prebuilt API here. Um, and then, uh, yeah.

And, and then runs it,right? So, um, I think I need to, like, improve the Claude.md for this. But anyways, you can see that, like, it's able to, like, check 200 plus Pokemon and then check for their type and, and, you know, get their, get their information,right?

So this is, like, uh, just a quick example on, like, how to do cogen and how to use Claude Code to do it,right? So, um, it will run this script and then, like, uh, um, like, keep going,right? So, uh, it will give me the output.

And, um, yeah, basically what I want to show, let's see, we have roughly 15 minutes left. Um, yeah.

Guest 41:36:33

It's time to play Pokemon.

Thariq Shihipar1:36:35

It's time to play Pokemon. Yeah, yeah. Actually, this is one of the demos I was thinking of doing. Um, Claude Code plays Pokemon. So, like, let's say you want to do, like, an agentic version of Claude plays Pokemon.

How would you do it? Um, what you would do, I think, is, like, you would give it access to the internal memory of the, uh, the ROM,right? And so let's say that it wanted to find its party. It could search that in memory.

And Pokemon Red is, like, a very well in distribution, uh, reverse engineered, uh, game,right? And so it could search in memory to be like, "Hey, these are the Pokemon. Um, these are, like, this is how I figure out where the map is.

This is how I navigate it,"right? So this is, like, maybe actually, I'd like to the reader if you wanna try it out. It's like, um, there is, like, a Node.js GBA emulator. Um, I think I have to legally say you have to go buy Pokemon Red and try it.

Um, but yeah, I, I think, like, uh, yeah, good example. Anyways, here. So it's, it's fetched all of them and it, it's listed all their types. And, um, yeah, you can see how it's, like, used code generation to do this,right?

So, um, a quick example of using Claude Code to prototype this. Um, now there can be, like, more interesting, like, data here. So, um, I do want to leave time for examples. So I, I think I'll just sort of, like, for questions.

So I'll just sort of go through, like, an example. Let's say you're making competitive Pokemon. Competitive Pokemon has a lot of different variables and data. So this is, like, a, a text file from this online, like, a library basically, which stores, like, all of the Pokemon and their, like, moves and who they work well with and don't work well with and, you know, like, who they're countered by and all of these things,right?

So there's a ton of data here,right? And it's all in text file, um, which is actually pretty good for Claude Code,right? Because I can say, like, okay, um, hey, I'm gonna give it a little bit more data. Normally I'd put this in the, um, check the data folder.

Tell me I, I want to make a team around Venusaur. Can you give me some suggestions based on the Smogon data? Um, and Smogon is, like, this online API. And so I'm, I'm not entirely sure what it'll do here yet.

I haven't done this career before. Uh, but we'll see. I think it'll be, it'll be fun. Um,

where am I? Let's, oh, I see it.

Okay. Um, yeah, but what I wanted to do is sort of grap through this, this data,right? And, and sort of figure out from itself, from first principles, not having seen this data before, how can I, like, answer my queries,right?

So, um, while it does, it does that, I'll, I'll take any questions. Yeah?

Guest 41:39:32

Um, first of all, great work, John. Uh, so this is, like, really on top of Claude Code. And so my question is, if we were to deploy this customer-facing.

Thariq Shihipar1:39:43

Yeah.

Guest 41:39:44

Are we supposed to have Claude Code running in, like, uh, like a swarm or are we somehow able to take Claude Code far out just to use Claude and the agent SDK?

Thariq Shihipar1:39:55

Hmm. Yeah. So let, let me show you, like, very quickly, like, what the, what, what it looks like to use the agent SDK here. Um, so I've already done this file system,right? And again, I want you to think about the file system as a way of doing context engineering,right?

Like, this is, like, a lot of the inputs into the agent. So my actual agent file is, like, 50 lines,right? Um, and it's mostly just, like, random, like, boilerplate,right? Like, I guess, yeah, it's decided to stop it from, uh, writing scripts outside of that custom scripts directory.

Again, fully backloaded. So, um, yeah, you can see, like, it just runs this query, takes in the working directory, um, and, uh, like, like, runs it in a loop,right? And so probably I'd want to, like, turn into, like, some allowed tools here and stuff, but it, it's very simple.

And, and so, um, if I were to, like, productionize this, the first step I'd do is, like, okay, I, I've tested it on Claude Code. It seems to do pretty well. I write this file, then I put it.

There are two ways to do it. So one is I do think that, like, local apps might be coming back with AI because I think that, like, there's such an overhead to running it. Like, for example, Claude Code is a front-end app,right?

Like, it works on your computer. So maybe the way I ship this as a Pokemon app is, like, hey, I have, like, an app that you install and it works locally on your computer and it's writing scripts. I think that's one way of doing it,right?

Um, the other way is, yeah, you have, you host it in a sandbox. Um, and again, there's a bunch of different sandbox providers that make it really easy. Like, Cloudflare has a good example, um, of using the agent SDK.

And it's just, like, sandbox.start, you know, and then, like, bun-agent.ts. And that's kind of all it takes,right? Like, it's, like, like, they've abstracted away a lot of it. Um, so you run, like, the sandbox, um, and then you communicate with it.

And, um, yeah, I think there is, like, some very interesting stuff that I'm not sure I had time to get to, but, um, like, I, I think some interesting questions are, like, um, yeah, like, how do you do this sort of, like, service?

Now we're just fitting up a submit, like, a sandbox per user. Um, there's a lot of, like, I'd say best practices to solve here. One thing I just wanna call out for you guys to think about, um, if you're making a, an agent with a UI, like, let's say that you have, uh, yeah, my Pokemon agent and I wanted to have a UI that is adaptable to the user,right?

Like, maybe some users are doing team building, some users are helping it with their game, some users just want pictures of Pokemon. How would, how would I have an agent that adapts in, in real time to my user,right?

Um, the way I would do it is in my sandbox, I would have a dev server,right? And the dev server would expose a port. Um, it would run on Bun or Node or something. It would, like, expose a port.

The agent could edit code and it would live refresh. And, and your user would be interacting with that website. This is how a lot of, like, site builders, like, Lovable and stuff work,right? They, they use sandboxes and they host essentially a dev server.

And so thinking about this for your users, if you want a customized interface, this is a great way to do it. Um, okay. Let's see. Let's see what it did. Um,

okay. Cool. Okay. So, um, it's, like, written this, like, script. It's generated, like, showed me some base stats and suggested a, like, um, uh, a move set and some teammates. And you can see sort of, like, let's see, what did it do?

Um, control E.

Um, yeah. Okay. So you can see here what it started doing is, like, it started searching for Venusaur,right? And it started finding, uh, those types, the, the, the, like, those Pokemon. And when it does that, it also gets other Pokemon that mention Venusaur.

So it gets, like, its teammates and its counters and stuff,right? And it's sort of over this time found interesting Pokemon,right, that, like, it might work with,right? So it's done a bunch of these searches and it's gone to these profiles.

It's found its most common teammates and, and written a script to, to analyze it,right? And so this is all based on a text file. Of course, I could have pre-processed the text file a little bit more. Um, but yeah, it's, like, done this sort of, like, interesting, um, an analysis for me,right?

And again, I'll, I'll push up more code to the GitHub repo and, um, I'll also tweet about this. I'm on Twitter. I'm, uh, TRQ212. Uh, I tweet a lot. So, uh, definitely, like, mostly about agent SDK stuff. Um, but yeah, we have about eight minutes left, so I wanna spend the rest of the time taking questions about kind of anything, you know?

Q&A1:44:52

Thariq Shihipar1:44:52

I'm, I'm sorry we didn't get to do more prototyping. Um, but, uh, yeah. Are you there?

Guest 41:44:58

Yeah. I was gonna say, with the Claude play, can you, uh, sort of plug this in with that just to see if the agent will, uh, be more selective with the teammates that, uh, tries to capture?

Thariq Shihipar1:45:07

Yeah. Put it in, in Claude plays Pokemon. Yeah, yeah, yeah. I do want to make Claude Code plays Pokemon. I think that would be fun. Yeah, yeah, yeah. I, I think Claude plays Pokemon. I think we try and keep it, like, a pure reasoning task as much as possible.

Yeah. Um, other questions? Yeah.

Guest 41:45:19

I was curious about how people are monetizing Claude Code SDK. I mean, they're monetizing it, you know, kind of like, what about, like, if it's not going to Claude? You kind of, like, lose the opportunity to get all the margins if you do inputting agents.

Thariq Shihipar1:45:30

Yeah.

Guest 41:45:31

I'm curious, like, how have you people been shipping your own Claude Code SDK to key so that they kind of take the usage space or just price point?

Thariq Shihipar1:45:38

Yeah. I, I do think overall, especiallyright now, agents are kind of pricey. You know what I mean? Because, like, um, the models are have just started to get agentic. We really focus on, like, having the most intelligent models, you know?

And, like, you generally, this is just, like, an overall, like, SaaS business software thing. You'd rather charge fewer people more money that really have, like, a hard problem, you know? And so I think this is still good. Like, you probably should find, um, you know, these hard use cases, but I would say, like, number one, make sure you're solving a problem that people want to pay for,right?

It's, it's, like, the number one step,right? And then number two, um, yeah, I think you could do subscription or token based. I, I, I think this kind of comes down to, like, how much you expect people to use your product, uh, versus, like, how much you expect them to, like, use it occasionally.

Like, Claude Code obviously people use a lot. And in order to, like, we do a mix of, like, if we give you some rate limits and if you exceed it, we do, uh, usage-based pricing. Um, I think that, like, yeah, it's very, like, dependent on your own user base and kind of, like, what they will do.

But I will say monetization is something you should think about upfront and design your, you know, agent around because it's hard to walk back these promises. So, um, yeah, back there.

Guest 41:47:00

Um, I haven't heard you talk at all about folks. I would be curious to hear your take on how folks like that in their career.

Thariq Shihipar1:47:06

Uh, yeah. There's so much to talk about. Um, hooks are great. We, we, we do ship with hooks. Um, hooks are a way of doing deterministic verification in particular or inserting context. So, um, you know, we fire these hooks as events and you can register them in the, in the agent SDK.

There's, like, a guide on how to do that. Um, examples of things you might use hooks for is, like, for example, um. Yeah, you can run it to verify the, like, a spreadsheet each time. Uh, you could also look, like, let's say I'm working with an agent and, uh, I'm, the agent is doing some spreadsheet operations and the user has also changed the spreadsheet.

This is an interesting, like, place to use a hook 'cause you can be like, "Hey, has after every tool call, insert changes that the user has made." Uh, and you, and so you're giving it kind of live context changes, um, in an interesting way.

So, um, yeah, I think, uh, uh, yeah, there, there's more stuff on, like, the docs about hooks. Um, I am happy to, like, talk about it afterwards as well. Yeah. More questions? Yeah.

Guest 41:48:09

So when common, like, workflow I, I do is.

Thariq Shihipar1:48:12

Yeah.

Guest 41:48:13

Let's say I say sample data. I go through the sample data in Claude Code.

Thariq Shihipar1:48:18

Yeah.

Guest 41:48:18

Then I realize, okay, it's working.

Thariq Shihipar1:48:20

Yeah.

Guest 41:48:20

I want to take the same conversation that I've already done 'cause I'm going through a few questions there.

Thariq Shihipar1:48:24

Yeah.

Guest 41:48:25

And convert that into an agent.

Thariq Shihipar1:48:26

Okay.

Guest 41:48:27

Uh, which is that I followed a few steps. Now it's actually working.

Thariq Shihipar1:48:30

Mm-hmm.

Guest 41:48:31

But I don't want to rewrite all of the code to write the SDK, like, it's like because it works.

Thariq Shihipar1:48:38

Yeah, sure. I, I, yeah. So, like, let's say you've done this prototyping, you found something that works. What I would do is, like, I'd summarize the Claude.md. Like, obviously, like, when I tried doing this one time, it, like, didn't use my API directly and it wrote JavaScript.

I should have been more specific in my Claude.md to be like, "Hey, you should use this." Um, I, yeah, I, I think, like, so that's one thing. Um, the second thing is, uh, yeah, just summarize into Claude.md, have the helper scripts that you need, and then, like, write something like this agent.ts.

Guest 41:49:11

Correct.

Thariq Shihipar1:49:11

To, like, to run the agent SDK. Uh, yeah. More questions? Yeah. In the grey?

Guest 41:49:16

Uh, yeah. I, uh, just tried to put a Pokemon agent and it makes lines, but I was using the helper scripts to answer. It tries a couple times. Like, my SDK isn't very good 'cause it's just a side coded it.

Thariq Shihipar1:49:26

Sure, sure.

Guest 41:49:26

But then it tries twice and then it's like, "Oh, here's your comparison table." But it's just 'cause it's in distribution. Do you have any advice for that kind of problem?

Thariq Shihipar1:49:33

Yeah. This is a good question. And, and, you know, like, I'm, I think there is some messiness,right? Like, I, I think one of the things. If an agent knows an answer, um, and you want to, like, sort of, like, fight it kind of to be like, "Okay, like, you know, it's generation nine now and, like, Venusaur stats have changed and there's this, like, this new, like, era."

Like, um, this is hard. I actually think, uh, one of the ways of doing that is hooks. So you can say, for example, like, "Hey, uh, don't if, if you've, like, returned a response without writing a script, you know, you can check that.

You can be like, give feedback to be like, "Please make sure you write a script. Please make sure you read this data." Right? And, and you can use hooks to, like, give that feedback in. In the same way that in Claude Code, um, we have these, like, rules, like, make sure you read a file before you write to it,right?

So add some determinism. Uh, it can definitely be, like I said, it's an art, you know, sometimes, you know. Yeah. Maybe, like, ri, like, riding a horse, I guess, probably. Um, yeah. In the grey.

Guest 41:50:33

How are you guys dealing with, like, large code bases? I'm working, like, a 50 million plus line code base, and so.

Thariq Shihipar1:50:38

Yeah.

Guest 41:50:39

The correct tool doesn't work really.

Thariq Shihipar1:50:40

Mm-hmm.

Guest 41:50:41

Um, so I'm having to build, like, my own, like, semantic indexing type thing to kind of help with that,right?

Thariq Shihipar1:50:46

Sure, sure.

Guest 41:50:47

Is there any kind of, like, added Anthropic maybe thinking about how that can be more native to the product? Like, you know, in a couple months, is the thing I'm writing just gonna go away or, like, how, how are you guys thinking about that?

Thariq Shihipar1:50:57

Okay. Your last question in a couple months, is it thinking to go away? Just generally, yes. Yeah. Like, anytime you ask about AI, yeah. Uh, I think that, um, semantic search, uh, this is a Claude Code question more than agent SDK question, but happy to answer it.

Like, um, we, you know, there are trade-offs. So semantic search, it's more brittle. Um, I think you have to, like, index and, and, and search. And we, it's not necessar the model is not trained on semantic search. And so I think that's sort of, like, a problem.

Like, you know, grep, it's trained on because it's, like, it's easy to do that. But, like, semantic search, you're implementing your bespoke query. Um, for, like, very large code bases, you know, we have lots of customers that work in large code bases.

I think what I've seen is sort of, like, they just do, like, good Claude.mds. You start in, you know, try and make sure you start in the directory you want. Have, like, good, like, verification steps and hooks and links and things like that.

And so, um, you know, that's what we do. We don't have, you know, a custom we, we dog through Claude Code,right? So, um, yeah. Okay. Yeah. Last question.

Guest 51:52:03

We have to close, unfortunately, actually.

Guest 41:52:05

Uh, let's give it up for Tariq, everyone.