Simulated students train AI tutors, self-tested over GPT-5.4
Microsoft and the University of Illinois' simulated students can train AI tutors cheaply, self-tested better than GPT-5.4 role-play across 60 students, pending independent verification.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
Training AI tutors on simulated students can mimic real students more accurately than prompting GPT-5.4 to play them, and produced a chess tutor that scored highest among three versions in expert review.
Previously, training AI tutors required feedback from real students, which was costly and slow.
Microsoft and the University of Illinois' StudentSim preprint is built on Qwen3-4B: it first learns common mistakes and revision behavior from pooled subject data, then adapts to individual students from a few records. In the authors' own tests across 60 students in chess, English, and math, it mimicked students more accurately than GPT-5.4 prompted to play them. A chess tutor trained with it scored highest among three versions in expert review.
The authors stress this is only a proof of concept: chess has an engine for judging moves, while essays and open-ended math lack reliable scoring; the benchmark is first-party and not independently reproduced.