Intermediate-layer representations may beat final layers: LAYERSCOPE self-tested on seven video models
A label-free layerwise evaluation reports intermediate-layer representations beating final layers, but results are author-run on seven models only.
ImportânciaLocalEvidênciaE2 não replicadaTratamentoRápido
Intermediate-layer representations of video and multimodal models can outperform final-layer and model-default outputs, according to the authors' own runs — LAYERSCOPE makes this comparable layer by layer without task labels, across seven architecturally diverse models.
Layerwise evaluation previously depended on task labels, which cost annotation effort and made systematic comparison of layers hard.
The team used local, global, distributional and correspondence-based geometric metrics across seven architecturally diverse models on video and multimodal classification, clustering and text-to-video retrieval tasks from MVEB/MVEB+, and report that no single geometric metric consistently predicts downstream performance; the results are the authors' own runs from a preprint, and whether they generalize to more models awaits third-party testing.