Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

AdamW Optimizer

A variant of the Adam optimizer that decouples weight decay from gradient updates to improve regularization.

Last reviewed: July 25, 2026

AdamW is the optimizer used to train the overwhelming majority of modern transformer-based language models, including most open-weight LLMs. It’s a modification of the Adam optimizer — itself a combination of momentum (tracking a moving average of past gradients) and adaptive per-parameter learning rates (scaling each parameter’s update by its recent gradient magnitude) — that fixes a specific flaw in how the original Adam implementation handled weight decay.

The Problem AdamW Solves

Weight decay is a regularization technique that nudges parameters toward zero each step, which helps prevent overfitting by discouraging any single weight from growing too large. In the original Adam optimizer, weight decay was implemented as if it were part of the loss function’s gradient — meaning it got divided by the same adaptive per-parameter scaling factor as the actual gradient. This coupling meant that parameters with historically large gradients received proportionally less weight decay, undermining the regularization effect in a way that wasn’t obviously visible in training loss curves.

AdamW decouples the two: it applies weight decay directly to the parameters as a separate step from the adaptive gradient update, rather than folding it into the gradient calculation. This seemingly small change, introduced in a 2017 paper by Loshchilov and Hutter, produced measurably better generalization in practice and has since become the default choice.

Why It Matters

Nearly every major transformer training run — from BERT to GPT-style models to Llama — uses AdamW as its base optimizer, typically paired with a learning rate scheduler (like cosine decay with warmup) and gradient clipping for training stability. Understanding AdamW’s decoupled weight decay is relevant when tuning hyperparameters for fine-tuning: the appropriate weight decay value is independent of the learning rate in a way it wasn’t under plain Adam.

Common Hyperparameter Defaults

AdamW is typically configured with a handful of hyperparameters beyond the learning rate itself: beta1 and beta2 (controlling the decay rates of its two moving-average estimates, commonly 0.9 and 0.999 respectively, though large-scale LLM training often lowers beta2 for better stability), epsilon (a small constant preventing division by zero), and the weight decay coefficient itself. While these defaults work reasonably well across many settings, large-scale pretraining runs typically tune them carefully alongside the learning rate schedule and batch size, since at the scale of a multi-million-dollar training run, even small hyperparameter misconfigurations can measurably affect final model quality or, in the worst case, cause outright training instability partway through a long run.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.