Visual cues boost video-planned robot navigation, CueNav reports 70% narrow-passage success
CueNav from TUM/MIRMI and MIT CSAIL self-reports that bird's-eye and body-view cues nearly double maze navigation success over cue-free planning.
ImportanceLocalEvidenceE2 unreplicated
Giving a video model visual cues — a bird's-eye view and partial body view — can substantially raise robot navigation success: the authors self-report that maze navigation success nearly doubles versus a cue-free planner, narrow-passage success reaches 70%, and they demonstrate zero-shot semantic-conditioned navigation and deployment of the same planner across robot platforms. Previously, video models planned navigation without cues before an inverse dynamics model converted predicted video into actions, with notably lower success.
The method, called CueNav, comes from a TUM/MIRMI and MIT CSAIL team; all figures above are author self-reported, the abstract does not state whether tests were simulation or real hardware, results are independently unverified, and the project page shows code not yet released; the preprint was submitted to arXiv on September 15 and revised September 16 (arXiv:2609.16737).