Nash Decoding Beats Scale
Nash decoding allows masked language models to outperform autoregressive models up to 18 times larger on question-answering benchmarks.
The preprint models text revision as a multi-player game: each token position is a player, vocabulary items are actions, and the goal is a Nash equilibrium that maximizes joint probability. On CLAPNQ, PubMedQA and CoQA, the authors self-report that masked models (which predict all positions in parallel) using Nash decoding achieve higher F1 and ROUGE scores than autoregressive models (which generate token by token) with up to 18x more parameters, without fine-tuning. The cost is significantly increased test-time computation, not quantified in the abstract.
Results are from the authors' own evaluation; no independent reproduction has been reported.
Fuentes:arxiv.org