Intro0:00
Alright, hello London!
And I hope everyone's been enjoying the AI Engineer Europe so far. There's so many amazing speakers; I've been, like, watching talks and talking to people for days now, and it's been immense. I'm Sam, I lead development of GitHub's MCP server, and yeah, I'm here to talk about mostly challenges we've faced building and scaling our remote server, how we've overcome them.
And, uh, before I start, I just, like, I like messing with people, so you know, here's a quick show of hands. Who's used an MCP server? Good, good. Who's used GitHubs? Who has a hot take? No, um, and yeah, has anyone built a server or a client?
Oh, nice, quite a few. And yeah, has anyone contributed to the specification?
Oh, yeah! I got one. That's actually the first one, I think. Other than the MCP Dev Summit, there were quite a lot of them. But, um, yeah, anyway, it's really awesome to see so many hands, so I'm glad I've actually come to theright place.
GitHub MCP Launch1:28
But yeah, for GitHub, you know, our MCP journey started, at least in public, in April last year. And, uh, we actually opened source to our local MCP in April last year. And we've just turned 1 years old, so I'm super stoked by that.
But, um, yeah, back then,right, there was a tremendous buzz. We were the most starred repo on GitHub of the particular week. And, uh, like, the exposure meant we got a high volume of public contributions, rapidly filling gaps in platform coverage that people kind of wanted to add tools and things.
Tool Overload2:05
And, you know, not everything was perfect,right? After a month or so of new features, agents, in some ways, were getting worse at using GitHub, and context windows were getting blown out quicker. And, you know, we picked, I think, over 100 tools, and certainly at the time, that was just too many.
LangChain had already produced research they published in February that year, you know, of the exact kind of problems we were seeing. More tools don't make better agents, you know, they get confused and forgetful. Well, I say more tools, like more context, and more tools shove directly into the context, to be precise.
But, uh, yeah, GitHub's a really expansive platform, and we provided tools, you know, for repos, issues, PRs, actions, projects, like, even more things. But the hard part of solving this was, like, we didn't want to prevent users from having the tools individually that they needed and they used.
And suffice to say, our user base is pretty diverse. And probably even, like, on GitHub platform at the moment, there might be, like, one or two claws as well. And for the record, there's a team of us who work on it.
It's not just me. And my team is awesome. But, uh, yeah, so to try and fix some of this, you know, I quickly added this thing, tool sets, which was, you know, a kind of grouping concept of related product tools, and users could just pick which ones they wanted and configure it.
I also, like, added a dynamic tool selection thing where agents could discover sets of tools and then turn on in chunks. And we never released it, but I made a kind of RAG version of the same, you know, for kind of semantic tool search and discovery.
But, uh, like, what do you think happened, even in spite of all this stuff?
Context block.
Everyone used the default settings. It was really annoying because, like, in a way, we had all these elegant solutions. All they did was require users to actually, you know, configure the JSON a little bit. Most users just don't.
Maybe it's even partially a spec problem because, you know, for, like, every proposal so far for grouping to the MCP specification, for various reasons, has been rejected. And there have been several attempts. And, like, in a sense, every mode or configuration we add, you know, one could argue is papering over potential gaps.
Like, or gaps in client implementations. So, like, as an example, we have a read-only mode, and roughly 17% of our users use it, but it maps one-to-one to the read-only, sorry, yeah, the read-only hint annotation. But, like, no client exposes that as a method of filtering servers.
I think some gateways now do, but anyway, it's an interesting, easy win for more enterprise use cases where people often only want that. But, yeah, we needed to find better solutions to context reduction. And you don't need to worry too much about the specifics.
This is dated now. But, like, we started trying to optimize, and we looked at the usage patterns on our remote server. And initially, you know, we cut the amount of context used by focusing the tools more specifically to the general case and based on usage to, like, about 49% reduction of the initial load.
And then we subsequently also grouped CRUD tools and brought that down even more. And I think, like, I think you get about 40 tools if you use the default configuration, and then you can kind of expand or contract that based on your own preference.
But, yeah, like, it's easy to customize. And we've also, like, recently had a massive push to, you know, reduce output tokens of a lot of tools as well. And in this example, you know, just by tailoring exactly what comes of the list pull requests, it's, like, actually lost more than 75% of the tokens used in the output.
Optimizing Context5:50
So, you know, in terms of how token-hungry GitHub server is, like, it's a moving target. We're constantly changing things that improve it. And if you haven't used it in a while, like, it's likely very different from a few months ago even.
And, um, yeah, anyway, like, and we haven't ruled out more advanced approaches like code mode, and we're always experimenting internally. But, uh, on the heels of this, we also dug into our data, and we found some more opportunities.
Tool Reliability6:43
So, yeah, like, we made a big push to reduce tool failures as well. And the success rate is roughly, I think, over 95% at this point. But, uh, like, not all failure is preventable because agents don't necessarily know which repos they haveright permission on.
They still hallucinate. But, uh, we've been able to identify significant numbers of errors that could be overcome, mostly by encoding a sort of agent intent into our tool surface. And, you know, you might have to make five API calls to make it more robust, but, you know, in that case, we do that in the server side to reduce round trips because that, you know, saves context, saves time, and usually makes a massively better experience, you know, makes the agents more successful.
And, yeah, we also started to run evals last year. I'm not going to go into detail. That link takes you to a blog article that my colleague Xenia wrote about doing it. But, like, one of the gists is, instead of micro-optimizing individual tool descriptions, you know, you try to test them against each other to try and make sure that they're called at theright times and not called at the wrong times so that in the pool of each other, they don't fight for, like, you know, like, the perfect tool description that makes the agent call it all the time is terrible, as is the reverse of that.
So you need to try and get that as tight as possible. But yeah, this could be a whole other talk. Security, on the other hand, is something that's, like, a kind of constant menace in all of this. I've seen lots of people talking about this.
Security Setup8:07
And it's a real problem in some ways for us because, you know, we have a lot of people using plain text access tokens for MCP in the wild. And usually they're stored somewhere the agent can access. They're frequently long-lived, they're often over-privileged, and they're kind of sat there just waiting to be abused.
End users, like, I don't think they're choosing this, you know? Like, it's actually hard to make configuration easy and secure at the same time. And clients have to make use of system keyrings or encrypted storage. And, like, VS Code does.
But, you know, the MCP spec also provided a better way with remote HTTP, which, you know, is all the way back to April last year as well. And we embraced this, of course. And we wanted to make secure connection the path of least resistance.
We didn't want users to have to download a local runtime. And, you know, our remote server supports OAuth 2.1, and my team even helped add the proof key for code exchange support, which is commonly known as Pixie, to GitHub's authorization server to improve the security posture for client apps.
But as I said, we hoped OAuth would be the path of least resistance. And again, perhaps some of you might know what happened.
Everyone expected us to support the dynamic client registration. And for us, like, it created more problems than it solves because, like, if you implement it kind of properly, it's hard not to have unbounded growth of app databases and challenges of how you would bucket them for rate limits, and there isn't a reliable app identity.
So we disconsidered it and rejected it. And, like, we feel like it was a well-intentioned mistake. And we're, you know, we're not the only authorization server to not support this. And, um,
even, like, MCP itself,right, it decided that client ID metadata is probably the way to go. And I can't promise that we're going to support it, but I promise that I am trying to get us to support it. And that should make logging in, like, massively easier.
But, yeah, more on that in the future. And also, speaking of security, some of you may have seen this. This was a fun day. But, like, you know, Invariant Labs published this. And, you know, like, it's a correct, sort of, correctly done prompt injection exfil attack for getting private data out of GitHub.
Prompt Injection10:39
And the thing is, you know, they called specifically GitHub's MCP server out. And I think that, you know, we do provide the tools that can enable that if you just kind of enable them all. But, uh, it applies to almost every agent setup, whether they use MCP or not, or whether they use GitHub MCP.
You know, like, the lethal trifecta stuff, which I'm not going to rehash now because I think many of you have probably seen it, or you can look it up, like Simon Willison's blog post on that's excellent. But, you know, the utility of agents is in direct conflict with kind of protecting this stuff.
And it's, like, it's an active space trying to work out how to prevent these problems. But it's not solved, and it's very much not unique to GitHub. And we have users with wildly different risk profiles. You know, like,
you know, we even have people that have, like, air-gapped GitHub enterprise server instances in, like, much more secure. And then, you know, obviously, the cloud bros, et cetera, are also just running straight to GitHub with, like, you know, probably full token access to the agent and everything.
And that's kind of also interesting,right? And, like, I'm not naysaying any of this. It's just it's cool to kind of see what people do and see if we can actually support the different use cases and security postures while everyone experiments with this stuff.
And we also kind of use, like, lean on auth to manage tools as well. And this is something I'm pretty happy with. If you log into GitHub MCP with a path token, we just immediately filter the tools down by the scopes that the token has.
Auth Scopes12:27
You don't have to do anything other than give it the token. On OAuth, we support step-up auth. So, you know, you can get a, we could return a scope challenge, and then it will interactively ask the user if they want to allow the scope.
And if you do, then you can, like, continue the tool call. It doesn't fail, which I think is also nice. And VS Code, for example, supports that. And I initially worked on this with them just because they already have a token to use GitHub.
And what they wanted was that if their baked-in token doesn't have permissions to use everything, that instead of just failing, there was a mechanism for users having a clean install and then an upscoping later if they need it.
And, yeah, lastly, server tokens as well. Like, they didn't have a, like, on actions and things. They didn't have a user. So user-specific tools are kind of out there. And then by removing those, we're just removing kind of constant sources of failure and wasted context at the same time.
Stateless Architecture13:47
We run a completely sort of stateless server setup. And we have been using Redis for session storage. You know, it's standard observability in debug kind of stack. Like, this is not a weird picture, but I guess one of the weird things for some people is a lot of people are running a stateful MCP server process in the singular and have kind of struggled with how you get it into this shape.
But
for us, like, we did a few things because it's very dynamic. But, like, one of the fun things we did is
we actually make a brand new, in the SDK sense, a brand new server instance on every single request. And we add the tools to it at the start. So whatever your configuration is, it just builds this, and then you get what you've asked for or what you're allowed to use because some things have policies that impact whether you've got tools or not.
And, yeah, like, we've been able to scale to, at this point, we serve around 7 million tool calls a week. And, you know, we don't have session affinity. Even the sessions, we generally only use them to identify. It's the only way to identify the self-reported client identity that comes through MCP.
So it's useful for us to understand, like, what clients people are using the server with. So, yeah, like, we use sessions for that. But, um,
yeah, we also have, like, wanted to bring experiments to all of you and everyone. And we have this thing that's an insider's mode. And all it does is it turns on certain feature flags and things for experiments that we're happy to just ship to anyone who wants to use them.
Experiments15:14
And this just takes you to the documentation. But, like, an example of something that we haven't released generally yet, but is on insiders, is our MCP apps. And, like, just, you know, I set up the example before I came in.
But, like, it's quite nice when you're talking to the agent to have the opportunity to kind of edit the AI-generated issue, especially if you're, you know, you're working heavily in professional open source stuff and you want to make sure that it's you posting and it's not going to get closed as a sort of bot-generated thing.
Like, this is a nice human-in-the-loop thing that MCP enables. And I much, you know, I wasn't sure how much I would like it at first, but then I've come to love it because I kind of care about how my issues and things are received by people.
And this is just a really great way to make sure that I can check that.
Future Outlook16:30
So, yeah, like, in terms of where I think it's going, like, something along these lines, I think a near future, you know, server discovery will hopefully be automatic. And tool use will probably become more compositional, like Bash, or piping tools into other tools, streaming data through them, or like, you know, Cloudflare's code mode approach, or Anthropic's tool search tool API, which just landed in Claude Code a couple of weeks ago.
And OpenAI recently added a similar API as well. Sorry, OpenAI added a similar API too. And, you know, I fully expect that, like, thousands of tools will be normal very soon. We're trying to iron out all the problems that prevented it in the first place.
And that'll probably reverse many of the fewer tools decisions. And users hopefully won't even have to know what MCP is. They'll just convey what it is they want to do and the auth setup and, like, you know, the tool selection.
Things will become truly autonomous. And I don't think we're that far away from this, but we're kind of in this experimental phase where we're not really there yet. But I think harnesses like Pi are also interesting because you can build a weird client that maybe optimizes this in a really good way yourself.
So I would encourage people to experiment with crazy clients. I feel like you never know you could be, like, the next,
Closing & Q&A17:53
like, well, if you're super lucky, you could be, like, the next Claw,right? You could publish something that goes so viral it totally changes the agentic game. I wanted to end on a high and look at some numbers. So, like, GitHub itself, it's actually got over 11 million Docker downloads of our standard IO server, which is by far not the most used version of it either.
We've got 126 contributors now and over 2,300 issues and PRs, which it's been over seven a day, like, every single day for over a year now, which I do look at almost every single thing eventually. So it's been, like, quite a year.
I mean, other. Some repos have it even worse, but, like, I also love it, so I please keep doing it. And, yeah, we've got almost 4,000 forks, which blows my mind. I kind of want to know, like, the weirder things that people have done that they haven't contributed back.
Yeah, nearly 30,000 stars, and we're fast approaching 8 million tool calls a week. And GitHub itself is also facing a new challenge.
This is really intense,right? And it shows no sign of slowing down. I still wanted you to keep opening issues and PRs for us. Like, we will cope, but, you know, this is new territory. And, you know, everything's, like, mildly on fire for everyone, I think, these days.
And it's just exciting and fun. But, yeah, thank you so much for having me.
I think I got, like, 30 seconds. I don't know if anyone has anything they want to ask, but.
What's your take on piping tool calls?
You know what? I think, like, things like trying out MCP CLIs and things like that is a fun avenue. I don't think it's entirely ironed out, but, like, one thing you can do, take the read-only tools from some MCP, wrap it in a CLI, and just give it a proper help and just see how the agent does.
Like, stuff like that is surprisingly effective. And, you know, like I say, I want people to mess with this stuff, so I would encourage you to just try it if you're interested. Allright, I'm zero seconds. I will answer you, but in person, if that's okay.





