Stronger play pushes computation deeper: monotonic shift from Elo 700 to 2500
The author's own ablations show Maia-3's computation migrating to deeper layers as skill conditioning rises, opposite to a common prediction; generalization to language models is untested.
ImportanciaLocalEvidenciaE2 no replicadoAnálisisRápido
The stronger a chess model plays, the more its computation concentrates in deeper layers: author David Litman's own ablations found that as Maia-3's Elo input rises from 700 to 2500, the causal center of mass migrates deeper, monotonically, for every move type measured, with knight forks showing the largest shift.
The plausible prior prediction was that skilled features should be computed earlier for reuse, meaning computation should move forward as skill rises; the measured direction is the opposite.
The subject is Maia-3, an 8-layer chess transformer that takes player Elo as an input, so skill can be varied without changing weights; the measurement was done head-by-head via ablations, using the author's own chessformer-lens tool. The sample is this single chess model, and whether it generalizes to language models is unknown.