New Optimizer Enables Single-GPU Training of 13B Models
The Clean optimizer reduces second-order method memory to linear cost, achieving 26% faster convergence than AdamW and enabling single-GPU pre-training of 13B models.
ImportanciaMaterialEvidenciaE2 no replicadoAnálisisRápido
The Clean optimizer reduces the memory complexity of second-order methods from quadratic to linear, enabling the pre-training of a 13B-parameter model on a single 80GB GPU.
Traditional second-order optimizers like SOAP accelerate convergence but incur prohibitive memory costs. Clean uses randomized Nyström approximation to estimate preconditioners, preserving curvature information while significantly compressing state usage.
Author-reported benchmarks show Clean reaches AdamW's final performance 26% faster in wall-clock time. Its low-precision variant, Q-Clean, reduces optimizer memory consumption by over 50% compared to Muon. These results are currently from a preprint and have not yet been independently reproduced.