M2Tok's multi-head multi-codebook action tokenization reportedly cuts reconstruction error and lifts VLA success
M2Tok splits latent action features across heads with independent codebooks; its authors report reconstruction loss significantly below prior discrete action tokenizers and higher VLA success rates.
ImportanciaLocalEvidenciaE2 no replicado
M2Tok, an action tokenizer that splits latent action features across heads with an independent codebook per head, reportedly reduces reconstruction error and improves VLA task success rates, with code released open source.
Earlier discrete action tokenizers were limited by a single codebook's quantization expressiveness, incurring higher reconstruction loss and constraining downstream VLA performance.
Chunpu Xu, Yao Mu and nine authors in total self-report that M2Tok's reconstruction loss is significantly lower than existing discrete action tokenization methods, and that VLAs built on it achieve higher success rates on RoboTwin, Simpler-Env and 3 zero-shot real-world tasks.
These are self-reported results with no third-party replication yet; the work was submitted to arXiv on September 16 (id 2609.18259, labeled ECCV 2026).