Eduardo Model Cuts AI Tutoring Compute Costs
New research introduces Eduardo training method, enabling a 27B model to match frontier AI tutors with significantly lower compute.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
The Eduardo-27B model matches the performance of Gemini-3.1-Pro and Claude Opus 4.8 on two tutoring benchmarks while using only 1/2.4 to 1/6.2 of the "thinking tokens" required by those frontier models.
Traditional RL-trained AI tutors often suffer from reward hacking, where they simply provide answers to maximize scores, fostering student dependency. This study introduces a "masked near-transfer post-test," forcing the model to improve rewards by guiding students to solve problems independently, thereby distinguishing true teaching from mere telling.
The team has open-sourced an 8,671-problem dataset, the training environment, and model weights. Current results are based on author-reported benchmarks and await independent community reproduction to verify robustness in broader scenarios.