AIAI EngineerFeb 5, 2025· 22:39

Scaling AI in Education: A Khanmigo case study: Shawn Jansepar

Shawn Jansepar, Khan Academy's Director of Engineering, details building Khanmigo, an AI tutor and teacher assistant on GPT-4 with OpenAI, arguing generative AI can democratize one-on-one tutoring. He describes a rapid prototyping culture that launched Khanmigo in three months via a company-wide hackathon, replacing traditional agile with a prototype-to-beta-to-launch framework. Technical challenges include math accuracy through a math agent and chain-of-thought prompting, refactoring prompts into a component architecture for testing, and managing scale using multiple models and dedicated Azure compute. Jansepar highlights ethical design: a Socratic tutor that avoids answers, teacher moderation, and a writing coach with revision history. Khanmigo now has 200,000 paid users, half in school districts, with teacher tools sponsored by Microsoft free to US teachers. Plans include releasing a math tutoring benchmark and evaluating smaller models like Phi-3.

  1. 0:00Intro
  2. 0:53Mission & Vision
  3. 2:27Khanmigo Demos
  4. 5:07Writing Coach
  5. 8:46Classroom Impact
  6. 10:20Building Khanmigo
  7. 15:54Technical Hurdles
  8. 20:12Scaling & Cost
  9. 22:06Conclusion

Powered by PodHood

Transcript

Intro0:00

Shawn Jansepar0:12

Allright, thanks so much for coming to my talk. Um, yeah, I'm going to talk to you today about how we've been scaling AI in the classroom, so kind of a case study of Khanmigo. My name's Shawn Jansepar, I'm the Director of AI and Learning at Khan Academy, and I've also been doubling up as a product leader, so thankfully when we're trying to decide on features and roadmap, I don't have to negotiate with anyone but myself.

There are other product managers, but it makes that process a lot easier. So today I'm going to talk about problems we're solving at Khan Academy, how we're solving those problems, walking through a few demos, how we transformed Khan Academy into an AI-first organization, and some of the technical challenges solved, and what we're focused on next.

So, for those of you who don't know, Khan Academy is a nonprofit with a mission to provide a free world-class education for anyone, anywhere. We have 140 million registered users, 7 billion learning minutes just in 2021 to 2022, 1.5 million very, uh, very active learners who use Khan Academy for 2 hours a month, and we—most recently we have about 200,000 paid Khanmigo learners and teachers, half of which are in school districts who are getting access in the classroom, half of which who are just paying on their own.

Mission & Vision0:53

Shawn Jansepar1:23

And we're specifically committed to serving historically under-resourced learners in our education system, so we, through our partnerships with districts, tend to try to partner with schools who have a higher percentage of students who are in the free and reduced lunch category.

So we're really trying to, you know, reach the students who, um, normally don't get access to this type of technology in education. And since the pandemic, outcomes have been worse than ever, so quickly, this is just a kind of graphic showing you that over time, in the third grade, students are largely grade-level proficient, but by the time they get to the eighth grade, the average student is 1.1 grade levels behind.

So we believe that AI can be used to successfully fill those unfinished learning gaps at scale by democratizing one-on-one tutoring for anyone, anywhere. So that's where Khanmigo comes in. It's our tutor for learners, and it's our assistant for teachers.

And just even taking a step back for a second, emulating tutors has really actually been our north star the entire time, but thus far we've been limited by the technology. We've been doing, you know, self-paced video, self-paced practice.

Sal famously said in his TED Talk that his cousins prefer the automated version of their cousin to their real cousin, and LLMs for us are just kind of a tool that help us accelerate towards that vision. We're really not just trying to use AI for AI's sake.

Khanmigo Demos2:27

Shawn Jansepar2:40

So how are we using AI to solve tutoring and teaching assistance at scale? There was a really great quote from Notion's CEO where they mentioned that this is as much of an interface revolution as it is a technology revolution, and they are really good at interfaces.

We think it's the same for us, but we're really good about building educational technology at scale for millions of learners around the world. So I'm not going to go through this whole list, but just generally things we're doing: building a Socratic tutor that doesn't give away answers, optimizes for the accuracy of math, is context-aware of your work, and is deeply integrated into our content.

Multilingual, text-to-speech, etc. So I'm going to go into some screenshots. So one example here is just comparing Khanmigo to ChatGPT, and you'll notice in this screenshot that here there's just kind of this diving into, allright, you asked me a question, I'm going to give you the answer, which can be great if you're an older learner and you know how to learn, but as learners in the classroom, we don't want students just getting access to the answer.

We want them to think about it deeply, we want more of a Socratic method where you're asking leading questions to help guide you to the answer. So you'll notice here, Khanmigo didn't give away the answer, it's asking leading questions.

And then on the math accuracy front, you'll notice that ChatGPT said, "You correctly distributed negative 4." In fact, in this example, you did not correctly distribute negative 4, it was incorrect. But Khanmigo's noticing that, uh, it notices that there was a small mistake in the distribution, and then it kind of back- it backs up and tries to help you from there.

So we're also been doing a lot to kind of embed Khanmigo directly into our platform and into our content, so you can immediately just say, "Hey, tutor me, I got this wrong," and then Khanmigo's passed the context in of, it has the full understanding of what this question is, what is the step-by-step solution, what is the, uh, thing that you entered, and then it'll analyze your solution and help you with it, maybe give you helpful tips as to why you might have landed there.

Um, so that's kind of after you got the question wrong. Before you get the ques- before you submitted the question, we wouldn't just dive into this. Um, things, you know, general chatbots are great, but there's a lot of UX affordances that are always—that's very important for education.

Things like having a great math input widget for doing complex equations, um, specialized graph rendering, etc. And then also UX for, uh, use cases like essay writing, which I'm going to dive into. So we've been thinking about how do we go beyond just that traditional chat interface.

Um, I'm sure you all have kind of had the experience when you were writing essays in grade school. You wrote that essay, your teacher marked it, there's some red ink on it, you never probably reviewed that essay again, you moved on.

Writing Coach5:07

Shawn Jansepar5:18

What I think that education—or what I think LLMs can do in education that is going to revolutionize the way we think about it—is that it's going to help us shorten feedback loops. There's never going to be a situation maybe where in the future you won't be able to submit an essay unless it's at least a B, because the teacher will know that you've been iteratively working on that essay with an AI.

And so I know there's a lot of concerns about cheating, so I'll talk about that in a second, but this is just kind of like an early preview into our writing coach, so it kind of breaks down the essay writing process into these different categories.

So in this example here, you're outlining an essay with an AI directly, and you might have questions about how to write that thesis, but it won't let you move on until you've written a pretty strong thesis, and it'll give you that feedback along the way.

And then there's drafting the actual meat of the essay. From there, you can kind of ask helpful for helpful tips, but if you ask it to write the next paragraph for you, it'll refuse. And then there's that last step around revision, so it'll give you feedback on axes of evidence, structure, style and tone, and conclusion.

And then this is where things get really interesting with the teacher. So the teacher now has full access to the history of what you built with the AI to kind of validate that you were the one who wrote that with the AI, and it wasn't just something that you copied and pasted from ChatGPT.

So you can kind of get a sense of that by looking at the conversation history. Um, we're building kind of these replay functionalities, and then it'll also say like, "Hey, these, you know, these 54 or 56 words just showed up out of nowhere."

So we really think that essay writing isn't actually going to be hurt by generative AI, but we are really going to revolutionize the way that works. Similarly for code, we also have a coding platform, and we're planning on doing the same for code writing at some point.

We're also building a lot of fun engagement things. Uh, students love this kind of stuff. You know, you can reclaim energy points to customize Khanmigo. Um, but then there's integration into the classroom. So, you know, one of the reasons why Khan Academy started focusing on classrooms instead of just focusing on independent learners who find us on Google is because kids who come to us on Google, it's great, but you have to care enough about learning to come to Khan Academy in the first place.

We really wanted to reach those kids who thought they were bad at math, who hated learning, and so it's important for Khan Academy to be integrated into the classroom and have that tight feedback loop between what the student is doing and so that the, uh, with the teacher so that the teacher can get insights into that.

So, for instance, we have a moderation feature where if any conversation is kind of going off the rails, is inappropriate, a teacher will be flagged and notified. So we kind of tried to do things when we were launching Khanmigo, uh, to turn some of the fears into actual features.

Um, this is an example of a classroom snapshot where a teacher can say, "Hey, how is my class doing?" And because it has access to all of that rich data about the usage of Khan Academy in the classroom, it'll give you that overview, and it'll also, um, give you the ability to kind of dive into skills reports and dive deeper into our graphs.

And then from there, as a teacher, you might be able to do things like, "Hey, I want to write a progress report to Reese's family." That's going to save you a huge amount of time because all this data is directly available.

We built a ton of other teacher tools, things like lesson planning, uh, exit tickets, etc. I would love to demo them, but there's not enough time. Uh, huge shout-out to Microsoft, they recently sponsored us bringing these teacher tools to all US teachers in the States and soon around the world for free.

Um, so feel free to check it out. Anyone can sign up.

Classroom Impact8:46

Shawn Jansepar8:46

And, oh, actually, I need to plug in the audio.

Instead of kind of talking about the results in the classroom, I thought I would just do a quick fun video showing how that's going.

Guest8:59

Talking to their class, how much they love using it. Let's talk to them.

Guest 29:03

Yes, let's do this. Let's go.

Guest9:08

This regular class was so much fun. Students in this class were opening up their imaginations, writing stories, learning about history, and they learned about America's first president.

Guest 29:18

What did they tell you about George Washington?

Guest9:20

Another lesson they learned is Khanmigo's here to help, but it's not going to give you the answer.

Guest 29:25

No, it just, it just helped me out. It didn't give me the answer, but it helped me out definitely.

Guest 39:31

It helped me write, like, stories, and it's so amazing.

Guest9:35

Now we're in Ms. Barreiro's fifth grade class, and everybody here is saying hi to Khanmigo.

Guest 29:40

Hi, Khanmigo.

Guest9:42

In Ms. Barreiro's fifth grade class and Mr. Rodriguez's sixth grade class, one thing was for sure: Khanmigo was here to help with personalized learning and adaptive feedback.

Guest 39:54

I feel like it's like a tutor. It's like a real person sitting next to you, like, telling you how to do a problem. It helps you figure out certain things that you wouldn't understand, but then he is like a teacherright behind you.

Guest 210:08

I absolutely love Khan Academy. For moving into this space, AI is our future, and I can't wait to see how Khan Academy soars in the AI space in education.

Shawn Jansepar10:20

Allright. So I want to dive a little bit into how we built Khanmigo and how did we transform Khan Academy into an AI-first organization. Um, you know, it didn't happen overnight. So, you know, first we kind of wrote up some requirements about, like, what would the perfect AI tutor and assistant, what would that be.

Building Khanmigo10:20

Shawn Jansepar10:38

We estimated the work, then we wrote up this nice Gantt chart of the timeline. Just kidding. We didn't do any of that. Uh, the keys to our success were really rapid prototyping and iteration. Um, actually, there's a great book that I highly recommend.

It's called Creative Selection. Um, it was about how the, uh, iPhone was created. Um, really kind of driven through this process of iterative development, um, driving it through demos and, you know, feeding the demos that do really well and, you know, putting an end to the demos that don't do really well.

And I think what we really learned is that magic really comes from when you pair incredible builders with domain experts who have some unique insight. And ideally, really, these should be members of your team who deeply understand the customer you're building for, as opposed to maybe just sending things off to an R&D team.

At least that's been in our experience. Uh, the history of our OpenAI partnership. So it actually all started when OpenAI needed some AP Bio questions to test their new model. Uh, we got a first sneak peek in September.

And then in December, actually between September and December, we didn't have much headway. It was a lot of like, well, what are you doing, what are you doing. And then we said, you know what, let's just do some rapid prototyping together and red team.

Um, and then we narrowed down on a few use cases. And then once we decided that we really liked the way that this tutor felt in our platform, we said, hey, let's try to launch something alongside a GPT-4, but let's do a company-wide hackathon to build a lot of new features, see what the team can come up with.

The best ideas from day one, we tried to ship those with GPT-4 in our launch in March. Um, and then we also participated in reinforcement learning where a few members of our content team were embedded within, um, with OpenAI to kind of train it on 100, uh, math tutoring examples.

So I would say some keys to our success. One is that we resisted perfectionism, which I would call an issue of Khan Academy of the past. We didn't let perfect be the enemy of the good. Uh, we made quick decisions for what we would call two-way door decisions, things you can reverse.

But for one-way door decisions, things like trust and safety, we thought deeply about those problems. Um, we trusted our intuition, we leveraged expertise if learning scientists, um, and educators, uh, kind of wanted to kind of work through that with us instead of us trying to A/B test every single decision.

And then we also iterated rapidly. I mentioned that with demo-driven development. And then we also kind of had some fixed timelines to avoid Parkinson's law where work can fill all available space. We also ended up driving forward what we called at Khan Academy, we have this process called architecture decision records, and then more recently transitioned that to organization decision records.

So whenever we have like a big strategic shift at Khan Academy or a big process change at Khan Academy, we document what is the problem we're trying to solve, who is the driver, who is the approver, who's contributing.

And then once we've kind of written that all out, we share that with the team so that they can digest it at their own time and pace. But what we did was we changed some of our Khan Academy core values.

So take a stand, we added trusting our intuition, um, we added resisting perfectionism to deliver well, so that that could kind of trickle down in the way we do performance reviews, etc. with the team and how we hire.

And then we also did a big change to our product development process. So I'm sure this is probably a tale that many of you have experienced, but we had an old process that we would consider agile. You'll notice that all these circles on the left, you know, they feel great.

There's a lot of arrows moving in certain directions. They look like you can go backwards, but, and then being a very heavyweight process, a lot of like heavy documentation you had to write for every single feature, etc. Which makes sense when you're building something, I think, that you know you're just going to ship to all users and it has to be high quality.

But in our case, we were doing a lot of prototyping, um, and so we kind of created this new process around there's a different level of quality of your code that you should write depending on your confidence in the feature.

So if you're not very confident and you're not sure, that's just a prototype. Go build it, be scrappy. Uh, if it's a kind of you're somewhat confident, then maybe build it as an A/B test or to a beta, test it with a few subset of users.

And then if you feel really confident in it and you're going to share it with all users, that's when the quality of the code needs to be higher. We have some checklists for that that people have to kind of go through.

There's a higher level of accessibility required when you get to that stage, etc. But then as long as you're following this framework, different teams are empowered to kind of build software in the way that makes sense for them, as opposed to having a one-size-fits-all for every team.

We also added an AI platform team as kind of a glue between our infrastructure team and product team. So there was obviously just a lot of AI platform team chat. So I think a lot of this is familiar, but it's really this team that owns the AI developer and prompt engineer experience of Khanmigo.

Um, so for example, we recently built this component interface in Go that when you as a developer conform to that component interface, it automatically means that things like tracing and LangFuse are available as long as you conform. Um, they own things like the AI router so that you as a product team don't have to worry about that.

It just kind of this transparently taken care of for you. Uh, we have a chat component, um, that has to have APIs and be extensible for product teams. And it's also a team that will consult with other product teams when they're building features if they need some, some help and expertise, as well as what we're kind of tentatively calling Khanmigo as a service where we're trying to work with third-party companies to embed Khanmigo into their products instead of them having to rebuild a brand new tutor from scratch.

So what are some of the technical hurdles? Uh, what were some of our technical hurdles that we overcame, and what are the challenges ahead? So on the tutor accuracy side, things we've done introduced a math agent. So we kind of give the AI the autonomy to defer to a math agent to work through problems, you know, get access to a basic calculator or, you know, Python might do some more advanced things if it's a more advanced math problem.

Technical Hurdles15:54

Shawn Jansepar16:19

Um, we're doing something called chain of thought prompting. So it's not just going directly to the model and saying, you know, help the student. We're actually doing things like we have this process that will say, write out how you think the student might have arrived at that answer.

Write out how you think it should have arrived at that answer. And then it kind of feeds that package over to another call to the AI that supplies that and says, using this context, how would you tutor that student?

We noticed that that helped, um, improve the performance dramatically. Um, we've been providing it with additional context like step-by-step solutions, additional context via RAG. And we've also been building this tutoring accuracy dashboard monitor, uh, to monitor quality over time.

So we can kind of get a sense of as we're shipping new changes to our prompts, as we're upgrading models, as we're, you know, making any change, we can make sure that, A, we don't have any regressions to the math quality because that's a really key part of this experience for learners.

And we also can, you know, ultimately just make sure we're trending in theright direction.

In terms of next steps, one thing that we're about to launch very, very soon, um, is a math accuracy benchmark data set. Soright now, when you look at data sets that different model providers are grading against, it's often, you know, is it good at doing math?

Is it good at reasoning? But there's nothing that says, how do I grade whether or not this is an accurate tutor? So we're going to be releasing that with a paper shortly. Um, we're also looking into integration, uh, with tools like Wolfram for some of the more complex math and then fine-tuning as well.

So prompt engineering, it's been a journey. So where we started was a lot of hacking together prompts really, really quickly. Lots of spaghetti code. I'm sure that's the same at a lot of your organizations. Um, we had a kind of varied test coverage depending on the activity.

So again, we had a lot of really good test coverage for the math portion of it. How good does it do at math tutoring? But we had a lot of other activities that we were just like testing ideas and throwing things at the wall and see if they stuck.

So things like, you know, talk to a historical figure. We didn't really have any test cases for that. When we wanted to iterate on models, it was really just, you know, somebody on our team who knew it well enough that would just go in and test it.

And a lot of it was also just looking in logs, uh, for completion calls when we had to debug a particular scenario that was a failure case that we wanted to dive deeper into.

Where we're at now is we've refactored that code using the component architecture that I mentioned earlier. So you can kind of have clearly defined inputs and outputs, and it makes it a lot easier to test and know, like, what are sub-pieces of the code that are worth testing.

Um, actively developing a kind of minimal viable, uh, test coverage suite across all of our activities. So we have a broad enough set of coverage and we have tools, uh, to run these manual test cases really quickly, but still with humans in the loop.

Um, and then we're also trying to move in more towards in the direction of model-graded evals to quickly test new models. But of course, some of that's not possible without humans on the loop. So it really depends on the scenario.

And then also, I think I mentioned this already, all stack traces are now available in LangFuse. If you want to go and see what happened, you can kind of go in, look through that, you know, stack trace and, um, debug from there.

In terms of next steps, this is not something we necessarily have an answer around, but we were looking for continued ways to get more reliable model-graded evals. Um, roughly we find that when we label things ourselves, compare it to the, you know, what the model will grade if we're doing things like, was there any, you know, math mistakes in this?

Uh, the AI is catching it roughly 70% of the time. And we also want to add a lot more languages, but, you know, how do we do that without bloom, uh, ballooning the test and iteration complexity? It's, there's probably not going to be a way to do that, but we'd like to reduce that complexity.

Scaling & Cost20:12

Shawn Jansepar20:12

And then in terms of, this is the last one, uh, in terms of scale, performance, and cost, we need Khanmigo to be reliable enough for teachers and students to use it daily, cheap enough for lower-income districts to afford it, and fast enough such that it doesn't lead to user frustration.

So things we've done so far is leveraging multiple models in production for different workflows, uh, fall back to shared capacity when we're reaching limits on our dedicated capacity, and continuing to do things like perceived performance improvements. Like if you notice in this, uh, bottom screenshot, we added this UI when we actually added a few changes to the experience that made it slower.

And I said to the team, we can't ship this unless we add a perceived performance improvements change so that users know that something's happening here. And anecdotally, somebody on our team said, "Hey, after we shipped that, my, my kid was using it and I said, do you think it feels slow?"

And that person's, that person's kid said, "No, I don't think so. Khanmigo's doing math." So that was a win. And then another one is just we have gotten a sizable donation from Microsoft, um, of Azure LLM Compute to bring Khanmigo, uh, free to all teachers.

So that's obviously helping as well.

In terms of next steps, uh, we're planning on doing things like load balancing between OpenAI and Azure for redundancy. Um, we want to evaluate more efficient models. We're in the middle of evaluating models like GPT-4.0 as well as small language models like Phi-3 and actively kind of working with Microsoft on Phi-3.

They're building what they're calling Phi-3.14, which is a small to medium-sized, um, language model that's optimized for math tutoring. And we also want to get to a point where we can dynamically scale up and down depending on our needs.

Uh, when we launched Khanmigo, we needed more throughput than OpenAI could offer on the shared capacity, and we wanted faster performance. So we ended up moving towards a dedicated instance. But, you know, that dedicated instance is great. We need to provision it for the peaks, but it's not doing a whole lot at night.

Conclusion22:06

Shawn Jansepar22:06

So ideally, we're, we're making more efficient use of our compute. That's pretty much it. I just challenge us all to leverage artificial intelligence to help us enhance human potential and human intelligence. Thank you.