Small Model Cuts Video Agent Costs
New training framework lets a 4B model match Gemini on video benchmarks while slashing API token usage.
BedeutungLokalBeweisE2 nicht repliziertAufbereitungSchnell
The VideoResearchAgent framework enables Qwen3.5-4B to achieve 40.48% accuracy on the Video-BrowseComp benchmark, comparable to Gemini-3-Flash-Preview.
This result comes from a preprint submitted on October 4. Traditional video agents rely on slow, unreliable live web interactions, while fixed local simulations often induce retrieval-specific shortcuts that fail to generalize to the open web.
The team introduced the RDR-GRPO algorithm, which randomizes candidate rankings, distractors, and metadata during training to reduce overfitting. Compared to the untrained model, cumulative API token consumption dropped by 74.9%.
These are author-reported results based on a specific simulator environment; independent reproduction on real open-web video has not yet been verified.