ReadingResearchRadarInvestment framework
Sign in / Sign up中文
Sign in / Sign up中文
ReadingResearchRadarInvestment framework

Reading

2026-10-041 posts

When Training a Model and Unsure How Big the Step Size Should Be, Adam Sets It Automatically for Each Parameter

Material

Long-term reading · 《Adam: A Method for Stochastic Optimization》(2015)

When training an AI model, you repeatedly fine-tune thousands of parameters along the gradient; how far each step goes is determined by the learning rate — too large and it oscillates, too small and it's too slow. Adam (proposed in 2015) is a method that sets the step size automatically: it records the magnitude and variability of each parameter's gradient, giving stable parameters large steps and jittery parameters small steps, and it is insensitive to overall gradient scaling.

Today, open any deep learning framework or tutorial and the default optimizer is most likely this one; new methods still use it as the comparison baseline when publishing papers. The rule it established — "hyperparameters barely need tuning" — has not been replaced to this day.

If you want theoretical convergence guarantees beyond convex problems, or want to use it to judge models and data the paper didn't test, don't use it to draw conclusions; evidence on long-term performance in non-convex deep learning settings is limited.

Adam: A Method for Stochastic Optimization (2015) | Next review 2027-09-20

Sources:arxiv.org

Research

You’re all caught up in this view