Single-pass latent steering halves VLM driving inference latency
Inference latency for vision-language model end-to-end driving can be cut by about 50% versus standard two-pass CFG guidance, while instruction following and driving scores improve rather than degrade — a result self-reported by Meibo Hu and three co-authors.
Previously such models followed instructions weakly and typically relied on standard two-pass CFG guidance to compensate, at the cost of running inference twice. The authors propose Latent-Centroid Steering (LCS), which replaces this with single-pass latent-space steering that projects toward a precomputed instruction centroid.
By the authors' own measurement, on the Bench2Drive closed-loop and nuScenes open-loop benchmarks, instruction following and driving scores both beat the two-pass CFG baseline, with inference latency reduced by about 50%. The code repository is public.
The result comes from an arXiv preprint (submitted July 31, updated to v3 on September 16, page marked IROS 2026) and has not yet been reproduced by others.
Sources:arxiv.org