Sherpa Teacher-RL Lifts Student LLM Accuracy by 20.5 Points
TL;DR
- Teachers trained with Sherpa raise instructed students' performance across all archetypes by an average of 20.5 percentage points.
- On MathTutorBench, Sherpa raises the Qwen3-8B teacher's overall pedagogy score from 52.5% to 79.2% and lifts student accuracy from 48.6% to 69.0%.
- In a 960-judgment pairwise study, 96 high school teachers preferred Sherpa's responses over the base model in 79.6% of comparisons.
A multi-turn reinforcement learning framework called Sherpa, from researchers at Stanford, Georgia Tech, and Carnegie Mellon, lifts a Qwen3-8B teacher model's pedagogy score on MathTutorBench from 52.5% to 79.2%, and raises instructed students' math accuracy by an average of 20.5 percentage points across student archetypes.
The paper, posted to Hugging Face, takes an unusual training signal. "Rather than predefining what constitutes good teaching, Sherpa separates pedagogical specification from teacher optimization: diverse instructional needs are instantiated on the student side, while the teacher is optimized based on student improvement," the authors write. The teacher is rewarded directly for the delta in a simulated student's accuracy after each tutoring exchange, with the student instantiated as one of several LLM archetypes conditioned on different learning preferences.
The authors ran a human preference study using 96 high school teachers over 320 dialogue contexts, totaling 960 pairwise judgments. Sherpa's responses were preferred over the base Qwen3-8B teacher in 79.6% of comparisons, with the win rate holding across all four student preferences, including two the model had not seen in training.
Overall student accuracy climbs from 48.6% at baseline to 69.0% after Sherpa training. The gains transfer across student backbones: a Llama-3.1-8B student's accuracy rises from 37.5% to 53.5%, though a smaller Gemma-3-1B student gains only 7.4 points. Code and model weights are posted at SALT-NLP/Sherpa. It lands in a thin lane for us, with just 21 education-focused stories tracked over the last 90 days, most of them about how children use chatbots rather than how models are tuned to tutor.
Originally reported by huggingface.co
Read the original article →Original headline: Sherpa Paper Trains LLMs to Teach Adaptively via Multi-Turn RL, Lifts Student Accuracy 20.5 Points