Fewer parameters, faster training, but don't treat it as a universal solution when switching tasks
Leitura de longo prazo · 《ALBERT: A Lite BERT for Self-supervised Learning of Language Representations》(2019)
ALBERT is an approach to saving parameters in text models: it reduces the size of word representations and shares parameters across layers, enabling faster learning of sentence meanings with fewer parameters. When you see it achieving high scores on text tasks with fewer parameters, first ask where those saved parameters come from.
The mechanism it provides is this: parameters are saved by reducing the size of word representations and sharing them across layers, while using sentence-order prediction to retain multi-sentence task capabilities. Therefore, when reading about low-parameter text models, ask whether these two aspects still fit your specific task.
If dealing with non-English languages, extremely small datasets, or if the width per layer is too large, do not use it to judge general optimality. In some single-sentence tasks, removing the sentence-order objective has limited impact.
"ALBERT" (2019) | Next review date: 2027-09-20
Fontes:arxiv.org