Loading use case index…
Loading use case index…
AI use case
Khan Academy's Khanmigo generative-AI tutor produced a 6.1 percent improvement in next-item correctness across 15 million tutoring threads over six months of product tests, vali…
Core facts from this catalog record. Primary narrative lives in the hero above; full raw fields follow in the next section.
Every column from the source row, in stable order. URLs open in a new tab.
Title
Khan Academy Khanmigo AI Tutor: 6.1% Improvement Across 15M Threads
Content
Khan Academy ran roughly 20 substantive product tests on its generative-AI tutor Khanmigo over six months, covering more than 15 million tutoring threads, and reported a 6.1 percent improvement in next-item correctness after giving Khanmigo structured access to each student’s Khan Academy learning record. The latest round of work, summarized in the Khan Academy engineering blog, narrows in on what actually moves outcomes and what does not. "Across roughly 20 substantive product tests in this space over six months, we now have a much clearer picture of what makes a difference," the team wrote in the post. The wins came from feeding Khanmigo real signal about a student’s recent performance patterns and skill gaps — Khan Academy learning history — rather than from generic prompt tweaks or model upgrades. Khanmigo was first launched three years earlier as a Socratic AI tutor for students and a teaching assistant for educators, built on OpenAI’s GPT-4. The new round of testing focuses on the math tutor and addresses two common failure modes of large-language-model tutors: surface-level reasoning and over-helpfulness. Each test compared a new version against the existing product experience, with results evaluated for probability of moving key metrics before broad rollout. The biggest measured gain — a 2.7 percent bump in next-item correctness when Khanmigo was given structured Khan Academy learning history, plus another 3.4 percent in a follow-up refinement, totaling 6.1 percent — came from giving the tutor better data about the student rather than from making it faster or more verbose. A separate test that fed Khanmigo hard-to-parse student data produced a neutral result with no measurable effect. Khan Academy positions Khanmigo as both a student-facing tutor and a teacher-facing assistant. The new evaluation framework is built on real classroom usage patterns from more than 15 million tutoring threads, run continuously rather than as one-off launches. A full paper describing the metrics, infrastructure, and experiment results is scheduled for the 27th International Conference on AI for Education. For operational scale, Khan Academy reports Khanmigo is now used alongside its full learner base on Khan Academy, with each new feature ship gated on statistical significance of metric movement. The team frames the work as incremental, evidence-driven optimization rather than a single breakthrough — meaning the path to better tutoring runs through hundreds of small A/B tests on real student interactions, not through bigger models.
Continue exploring AI deployments in the catalog.
Back to use casesCity
Mountain View
Company/Organization
Khan Academy
Continent
North America
Country
United States
Category
Diversified Consumer Services
Type
Deployment
Id
e414e9b2-5830-40d1-82ee-f68715e9979f
Created At
2026-06-30T10:11:25.022168+00:00