DiffAdapterVLA embeds trajectory tokens into driving VLM late layers, low latency and few parameters on NAVSIM
DiffAdapterVLA injects explicit trajectory tokens into the late layers of a driving vision-language model, with authors self-reporting high closed-loop planning quality, low latency and few trainable parameters on NAVSIM.
ImportanceLocalEvidenceE2 unreplicated
DiffAdapterVLA injects explicit trajectory tokens into the late layers of a driving vision-language model, letting trajectory states co-evolve with driving conditions during the backbone's forward computation while training only lightweight adapter modules; the authors self-report high closed-loop planning quality, low latency and few trainable parameters on the NAVSIM benchmark.
Previous approaches did not let trajectory states co-evolve with driving conditions in the backbone's forward computation, making it hard to balance planning quality, latency and parameter count.
The measurement was done first-party by Changxin Lu and seven co-authors: on the NAVSIM benchmark they self-report high closed-loop planning quality, low latency and few trainable parameters.
Boundary: results are author self-reported with no third-party reproduction yet; the preprint was submitted to arXiv on September 14 (v2 updated on the 16th, arXiv:2609.15322).