Agents writing game rules: best combo solves only 52.78%
GameLogicBench, the first per-tick game-rule benchmark, shows the best single run solving 52.78% of tasks across 20 model-scaffold combos.
ImportanciaLocalEvidenciaE2 no replicadoAnálisisRápido
The best single run by a coding agent implementing game logic solved only 52.78% of tasks, according to the first benchmark that checks game rules at every simulation tick: 72 Godot tasks, with 403 hand-designed scenarios expanded into 1,451 test cases.
Previously no benchmark checked game rules tick by tick, so incorrect rule implementations went undetected. Across 20 combinations of language models and scaffolds, the best score was that 52.78%; under the Claude Code scaffold, all twelve models scored lower as tasks grew from isolated mechanics to repository-scale features. Most failed submissions were runnable but implemented the rules incorrectly.
The authors' own ablation found that without mutant-based validation, incorrect agent submissions passed the evaluator; with open network access, agents copied code from public repositories. These are first-party preprint results; the success rates and evaluator validity await independent reproduction.