Preprint: Conditional Residual Prediction Helps Causal Video Model Reach 82.78 on VBench Without Teacher Distillation
New training method enables a 2B-parameter causal video model to score 82.78 on VBench, approaching bidirectional performance.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Key Finding: A new training method called Conditional Residual Prediction (CRP) enabled the 2-billion-parameter causal video diffusion model Optica to achieve a score of 82.78 on the VBench benchmark, using only approximately 15 million training videos.
Context: Traditional causal video models, suitable for streaming and long-form generation due to their autoregressive nature, typically yield lower quality than bidirectional models of the same size. Existing solutions often rely on knowledge distillation from large bidirectional teacher models, which is complex and computationally expensive.
Result: CRP reduces the model's over-reliance on history by predicting the target first and adding historical conditions as residuals. Experiments show this approach nearly closes the 6.14-point gap with bidirectional models trained under the same setup, without requiring any bidirectional video model during training.
Limitations: Results are self-reported in a preprint and have not yet been independently reproduced. The validation focuses on 480p, 5-second video generation tasks.