DiffAdapterVLA injects trajectory tokens into late VLM layers, self-reported low-latency closed-loop planning on NAVSIM
DiffAdapterVLA injects explicit trajectory tokens into the late layers of a driving VLM so trajectory states co-evolve with driving conditions, with self-reported low-latency closed-loop planning on NAVSIM.
ImportanceLocalEvidenceE2 unreplicated
DiffAdapterVLA injects explicit trajectory tokens into the late layers of a driving vision-language model, letting trajectory states co-evolve with driving conditions depth-by-depth inside the backbone, refining trajectories recursively with lightweight layer-wise adapters and dropping the separate planner. Previously, trajectory generation in driving VLMs was decoupled from driving-condition evolution and often relied on a separate planner, adding latency and complexity.
The authors self-report high-quality, low-latency closed-loop planning with a small number of trainable parameters on the NAVSIM benchmark.
Boundary: results are self-reported with no third-party replication; the work is by a team of eight including Changxin Lu, submitted to arXiv as a preprint on September 14 (v2 updated on the 16th, arXiv:2609.15322).