ACT encoder ablation's 35%→2% drop not reproduced in reruns of original code; latent zeroed at inference
Bo Kang, sole author, self-reports: rerunning the CVAE encoder ablation of Action Chunking Transformers on two simulation tasks in the original code, the paper's 35% to 2% success-rate drop did not reproduce, though small changes remain uncertain.
ImportanceLocalEvidenceE2 unreplicated
Rerunning the CVAE encoder ablation of Action Chunking Transformers on two simulation tasks in the original code, the paper's 35% to 2% success-rate drop did not reproduce, so readers now know that ablation result is not robust; small gains or losses remain uncertain, training duration and checkpoint selection can reverse which policy wins, and the cause of the drop is unidentified.
Previously one could only judge the encoder's role from the original paper's ablation numbers, while at inference ACT already zeroes the latent and does not use it; the author's timing shows skipping the encoder improves training throughput, and the code and evaluation tools are open-sourced.
The above is a self-reported result by Bo Kang as sole author, arXiv:2609.16745 (submitted September 15, updated to v2 on September 16), with no third-party reproduction yet, and the cause of the drop remains unidentified.
Source: arXiv:2609.16745 abstract page ↗; full-text HTML v2 ↗