Codec metadata speeds VLM inference up to 3.3x
CodecSight uses video codec metadata to guide VLM inference, reaching 3.3x concurrent streams and up to 93% fewer FLOPs, but self-reported.
ImportanceLocalEvidenceE2 unreplicated
Video codec metadata can serve directly as a runtime signal for streaming VLM inference: pruning image patches before visual encoding and selectively refreshing the KV cache across windows by frame type, with no model-specific training or offline profiling, lets concurrent streams reach up to 3.3x the baseline and cuts executed FLOPs by up to 93%.
Previously, speeding up streaming VLM inference usually required model-specific training or offline profiling, which was costly and hard to transfer to new models and workloads.
Yulin Zou and eight other authors self-report: across three VLMs and four video workloads, their vLLM-based implementation reaches up to 3.3x the baseline in concurrent streams, up to 5.3x faster average time-to-first-token, up to 93% fewer executed FLOPs, and at most a 4.64 percentage point drop in task quality. These are the authors' self-reported benchmark results.
The results do not cover other models or workloads, and no third party has reproduced them; the preprint was first submitted on April 7 and updated to v4 on September 15, arXiv:2604.06036.