New Simulator Boosts Agent Generalization
The MIMESIS project released a 9B-parameter specialized user simulator that significantly improves AI agent training realism and generalization.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
The MIMESIS project released a 9B-parameter purpose-built user simulator achieving a SOUL-Index score of 65.7, surpassing the strongest frontier models.
Traditional agent training relies on off-the-shelf assistant LLMs as simulated users, but their excessive cooperativeness and behavioral homogeneity limit training effectiveness. Collecting real human feedback remains expensive and difficult to scale.
Compared to Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points. After freezing the simulator for multi-turn reinforcement learning, agents outperformed GPT-5.5-trained baselines across eight environments and demonstrated stronger generalization across nine unseen user simulators.
Data comes from an arXiv preprint (v2, October 8, 2026); results are author-reported and await independent reproduction.