EfficientTDMPC claims new sample-efficiency best on HumanoidBench and DMControl
Thomas Evers and four co-authors report that EfficientTDMPC, with three changes, reaches a new sample-efficiency best on HumanoidBench and DeepMind Control Suite.
重要度局所的証拠E2 未複製
EfficientTDMPC makes three changes to TD-MPC-family model-based reinforcement learning: aggregating multi-horizon planning objectives across different rollout depths, adding a state-action value ensemble for MuZero-style methods, and penalizing uncertain return estimates with pessimistic reanalyze when generating policy targets. The authors self-report a new sample-efficiency best on HumanoidBench and DeepMind Control Suite.
Before these three changes, the TD-MPC family gave readers no way to judge whether such modifications could yield sample-efficiency gains.
The best is a self-reported first-party result with no third-party reproduction yet; it comes from the arXiv preprint by Thomas Evers and four co-authors (first submitted May 15, 2026, updated to v4 on September 17, 2026, arXiv:2605.16692).
Source: arXiv abstract page: EfficientTDMPC (v4, 2026-09-17) ↗