Statistical mechanics explains double descent: more parameters act like stronger weight regularization, unverified
In a single-author preprint, Congzhou M Sha proposes that training trajectories are finite-time diffusing particles in the training-loss energy landscape, making added parameters equivalent to stronger weight regularization; not peer-reviewed.
BedeutungLokalBeweisE2 nicht repliziert
Under Congzhou M Sha's statistical mechanics account, double descent works like this: the training trajectory is a particle diffusing for finite time in the training-loss energy landscape, which produces an effective weight decay; by equipartition of energy, adding parameters at fixed training loss lowers the temperature and moves toward stationary paths, whose L2 norm can only decrease as parameters increase — which the author says is equivalent to stronger weight regularization.
Previously the phenomenon lacked this kind of unified theoretical explanation; this is a single-author, self-reported claim.
The claim has not been peer-reviewed or independently verified; the preprint was submitted to arXiv on September 16 and updated to v2 on September 17 (arXiv:2609.19076).
Source: arXiv:2609.19076 abstract page ↗