Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

Learning Rate Scheduler

A training framework that adjusts optimizer step sizes dynamically over training epochs.

Last reviewed: July 25, 2026

A learning rate scheduler adjusts the optimizer’s step size — how large a weight update to make in response to a given gradient — over the course of training, rather than using one fixed value for the entire run. Nearly every LLM training run, whether pretraining from scratch or fine-tuning, uses some form of learning rate schedule rather than a constant rate.

Why the Learning Rate Needs to Change

Early in training, a model’s weights are far from any good solution, and a relatively higher learning rate helps it make fast initial progress. As training proceeds and the model approaches a good solution, a high learning rate risks overshooting past good minima or causing the loss to oscillate rather than settle. Gradually reducing the learning rate as training progresses generally produces better final results than holding it constant throughout.

Common Schedule Shapes

Warmup is a short initial phase (often the first 1-5% of training steps) where the learning rate starts near zero and ramps up to its target peak value, which helps avoid instability from large, poorly calibrated updates when the model’s weights are still close to their random initialization.

Cosine decay smoothly reduces the learning rate from its peak value down toward zero (or a small minimum value) following a cosine curve shape over the remainder of training, and is currently the most common schedule shape for LLM pretraining, generally outperforming simpler linear or step-based decay schedules.

Practical Relevance

Getting the schedule wrong — too-short warmup, too-aggressive decay, or a peak learning rate that’s poorly matched to the model size and batch size — is one of the more common causes of unstable or underperforming training runs, which is why published training recipes for major open-weight models typically specify their exact schedule in detail.

Why Warmup Specifically Matters for Transformers

Transformers are particularly sensitive to learning rate choices early in training because of how layer normalization interacts with randomly initialized weights — without a warmup period, an aggressive learning rate applied to poorly calibrated initial gradients can cause the loss to spike or even diverge entirely within the first few hundred training steps, sometimes permanently corrupting the model before it has a chance to find a reasonable region of the loss landscape. This sensitivity is specific enough to transformer architectures that learning rate warmup, while used elsewhere in deep learning, is considered close to mandatory for transformer training in a way it isn’t always for other architectures, and skipping it is one of the more common causes of failed or unstable large-scale training runs reported anecdotally by practitioners.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.