CSWAM self-reported to lift RoboTwin 2.0 Clean-to-Randomized success from 10.16% to 45.18%
Authors self-report: CSWAM adds a V-JEPA 2.1-based causal semantic expert to FastWAM-style world action models, lifting RoboTwin 2.0 Clean-to-Randomized success from 10.16% to 45.18%.
ImportanceLocalPreuvesE2 non réplicable
Readers can now know: CSWAM, a method that self-reports lifting RoboTwin 2.0 Clean-to-Randomized transfer success from 10.16% to 45.18%, while retaining efficient action-only inference. Previously, FastWAM-style world action models lacked causal semantic conditioning on sparse observation histories, leaving cross-randomization transfer success at roughly 10%.
The measurements were self-reported by eight authors: adding a V-JEPA 2.1-based causal semantic expert to FastWAM-style world action models, conditioning action denoising on sparse observation histories, and after embodied pretraining, RoboTwin 2.0 Clean-to-Randomized success rose from 10.16% to 45.18%; on two real-robot tasks across three OOD difficulty levels, average success rose from 27.5% to 70.0%. These are author self-reported simulation and self-test results.
Boundary: results are not independently reproduced; the work was submitted to arXiv on September 16 (v2 updated September 17, arXiv:2609.18462).