Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

Stochastic Gradient Descent (SGD)

An optimization method that estimates gradients using mini-batches to update parameters iteratively.

Last reviewed: July 25, 2026

Stochastic gradient descent (SGD) is the foundational optimization algorithm behind nearly all neural network training, including the training of large language models — though in practice, most modern LLMs use adaptive variants like AdamW that build on SGD’s core idea rather than plain SGD itself.

How It Works

Training a neural network means iteratively adjusting its weights to reduce a loss function that measures how wrong its predictions are. Computing the true gradient of that loss with respect to every weight, using the entire training dataset, is prohibitively expensive for datasets with billions of examples. SGD’s key insight is to estimate the gradient using only a small, randomly sampled subset of the data — a mini-batch — at each step, rather than the full dataset. This produces a noisier, less precise gradient estimate, but one that’s vastly cheaper to compute, and in practice this noise itself has a beneficial regularizing effect, helping training escape shallow local minima that a more precise gradient descent might get stuck in.

SGD vs. Adaptive Optimizers

Plain SGD applies the same learning rate to every parameter uniformly, often combined with momentum (a moving average of past gradients that smooths out the update direction). Adaptive optimizers like AdamW extend this idea further, tracking a separate effective learning rate for each individual parameter based on the history of its gradients, which tends to converge faster and more reliably for the large, complex loss landscapes of deep transformer models. This is why AdamW, not plain SGD, is the default choice for training modern LLMs, even though SGD remains foundational to how the field describes and reasons about the optimization process, and is still commonly used for training smaller computer vision models.

The Batch Size Tradeoff

The size of the mini-batch used at each SGD step is itself a meaningful hyperparameter: very small batches produce noisier gradient estimates (which can help escape poor local minima but slow convergence and underutilize parallel hardware), while very large batches produce smoother, more accurate gradient estimates but require proportionally more memory and can sometimes converge to solutions that generalize slightly worse to new data — an effect studied extensively in the deep learning literature. Techniques like gradient accumulation let practitioners decouple the effective batch size used for the gradient calculation from the memory footprint required at any given moment, letting them approximate the training dynamics of a much larger batch without needing the GPU memory to hold it all at once.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.