Deeper models find superposition geometries beating uniform polytopes
A first-party toy-model study reports asymmetric geometries with lower error than uniform polytopes, challenging Anthropic's conjecture; unverified on real LLMs.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
Asymmetric geometries for storing neural-network features can achieve lower error than the uniform polytopes previously thought optimal — a first-party toy-model result by Bartosz Rzepkowski and colleagues.
Superposition is the phenomenon where a model packs more features than it has dimensions into a limited space. Anthropic's 2022 "Toy Models of Superposition" conjectured that for independent features of equal importance and sparsity, the best storage geometry may be uniform polytopes; previously, shallow models only found symmetric solutions.
The study redoes the same task with a deeper architecture of a linear encoder plus MLPs, reports that asymmetric geometries yield lower error, and derives two theoretical benchmark decoders explaining why shallow models only find symmetric solutions. Code is open-sourced. This is a first-party result, with experiments covering only 4 features compressed to 2 dimensions; whether it generalizes to real LLMs remains unverified, and the study was published on LessWrong.