SyncVoice claims zero-shot video dubbing SOTA in one model covering Chinese and English
Xiamen University and Xiaomi MiLM Plus teams self-report that SyncVoice adds a text-visual fusion module to pretrained TTS for lip-synced zero-shot Chinese-English video dubbing.
ImportanceLocalEvidenceE2 unreplicated
A model called SyncVoice can time synthesized speech to visual cues such as mouth movements in video, and its authors claim state-of-the-art zero-shot dubbing on the LRS3 dataset, with a unified Chinese-English model trained on about 600 hours of Chinese and 1,190 hours of English audiovisual data.
Previously, pretrained TTS generated speech from text alone without sensing visual cues, making it hard to align dubbing with lip movements, and Chinese and English dubbing typically required separate handling.
These are self-reported results by authors from Xiamen University, Xiaomi's MiLM Plus team and collaborators, with no independent verification yet.
Boundary: the results have not been reproduced by third parties; the work is an arXiv preprint (id 2512.05126), first submitted on November 23, 2025 and revised to v2 on September 15, 2026.