Visuotactile representation leads baseline by 11 points on sim benchmark
VT-MUSE reports leading the strongest baseline by 11 points across all simulation benchmark tasks, with clear gains on real robots.
重要度局所的証拠E2 未複製
With the VT-MUSE visuotactile manipulation representation framework, all simulation benchmark tasks lead the strongest baseline it evaluated by 11 percentage points, with clear gains also in real-robot experiments, as self-reported by the authors.
Prior methods mostly encoded vision and touch independently before fusing them, ignoring contact timing. This framework first jointly adapts encoders via cross-modal temporal alignment and masked-view consistency, then combines a conditional variational latent model with tactile history, feeding a lightweight Transformer policy through gated cross-attention.
These results are self-reported by the authors; both the simulation benchmark and real-robot experiments were self-tested by the authors, with no third-party reproduction mentioned. arXiv preprint, submitted August 21, updated to v2 on September 17 (arXiv:2608.21290).