MoRE shares expert pools across layers: self-reported lower perplexity than standard MoE at 114M–1.15B
MoRE lets adjacent layer groups share an expert pool while each layer keeps its own router plus a learnable depth embedding; authors self-report lower perplexity than standard MoE and weight sharing at 114M–1.15B under equal budgets.
重要度局所的証拠E2 未複製
MoRE lets adjacent layer groups share a single expert pool while each layer keeps its own router, distinguished by a lightweight learnable depth embedding; the authors self-report that at three scales from 114M to 1.15B parameters, under equal compute and parameter budgets, perplexity is lower than standard MoE and weight-sharing architectures, requiring only minor changes to existing MoE implementations.
Previously, standard MoE kept each layer's experts separate, while weight-sharing architectures struggled to distinguish layers; MoRE aims to combine parameter efficiency with layer distinction via a shared pool plus depth embeddings.
The perplexity comparisons are self-reported by Eric S. Qiu, Kilian Q. Weinberger and seven authors in total, with no third-party benchmark named, and the scales are small.
The results have not been independently reproduced; the preprint was submitted to arXiv on September 16, updated to v2 on the 17th, with the page noting acceptance to COLM 2026 (arXiv:2609.18176).