Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

Gradient Clipping

A safety training strategy that caps maximum gradient values to prevent numerical instability.

Last reviewed: July 25, 2026

Gradient clipping is a training stability technique that caps the magnitude of gradients before they’re used to update model weights, preventing a single unusually large gradient from destabilizing training. Without it, deep neural networks — especially transformers trained on noisy, diverse data — are prone to occasional gradient spikes that can cause a sudden, catastrophic jump in the loss, sometimes irrecoverably corrupting the model’s weights partway through a long, expensive training run.

How It Works

The most common approach, gradient norm clipping, computes the overall norm (magnitude) of the gradient vector across all parameters, and if that norm exceeds a chosen threshold, scales the entire gradient down proportionally so its norm equals the threshold. This preserves the direction of the gradient update while limiting its size, which is generally preferred over simple value clipping (capping each individual gradient component independently), since value clipping can distort the direction of the update rather than just its magnitude.

Why It Matters for LLM Training

Training runs for large language models can cost millions of dollars and run for weeks on thousands of GPUs; a training run that diverges or corrupts partway through wastes that investment unless caught and rolled back from a checkpoint. Gradient clipping is one of several standard safeguards (alongside careful learning rate scheduling and weight initialization) used to keep these long training runs numerically stable, and it’s applied by default in essentially every major LLM pretraining and fine-tuning setup, typically with a clipping threshold around 1.0.

Norm Clipping vs. Value Clipping in Practice

Most production training frameworks default to global norm clipping, computing a single norm across all parameters’ gradients combined and scaling everything down proportionally if that combined norm exceeds the threshold — this preserves the relative direction between different parameters’ updates, which value-based clipping (capping each individual gradient value independently) can distort. A typical clipping threshold for transformer training is around 1.0, though the ideal value depends on model size, batch size, and learning rate, and is usually tuned alongside those other hyperparameters rather than set independently. Gradient clipping is applied after the full backward pass computes gradients but before the optimizer actually updates the weights, making it a lightweight, low-cost safeguard that’s included by default in virtually every deep learning framework’s training loop, and removing it from a large training run without a specific reason is generally considered a red flag rather than a reasonable optimization.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.