Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

FP8 Quantization

Compressing model weights to lower-precision numbers (8-bit, 4-bit) so models need less memory and run faster with minimal quality loss.

Last reviewed: July 25, 2026

What is quantization?

Quantization stores model weights in fewer bits — FP16 → INT8 → INT4 — shrinking memory roughly proportionally. A 70B-parameter model needs ~140 GB at FP16 (multiple GPUs) but ~40 GB at 4-bit (one big GPU). Since LLM inference is memory-bandwidth-bound, smaller weights also mean faster token generation.

Why it matters

Quantization is the difference between “runs on a GPU cluster” and “runs on one GPU” — or a laptop. It’s the enabling technology for local/edge LLMs (llama.cpp, Ollama, GGUF files) and a primary cost lever for self-hosted serving. FP8 is now standard on serving hardware; 4-bit (GPTQ, AWQ, GGUF Q4) is the common self-hosting sweet spot.

Quality trade-off, honestly

8-bit is essentially lossless. Good 4-bit methods lose a little — often unnoticeable in chat, but measurable on reasoning, math, and code, and worse for small models (quantizing a 7B hurts more than a 70B). Below 4-bit degrades sharply. The rule: a bigger model quantized to 4-bit usually beats a smaller model at full precision in the same memory footprint.

What people get wrong

  • Trusting perplexity alone. Run your own task evals on the quantized model; aggregate metrics hide task-specific damage.
  • Confusing weight-only and activation quantization. Weight-only (GGUF, GPTQ) saves memory; FP8/INT8 activation quantization is what unlocks faster compute on modern GPUs.
  • Forgetting the KV cache. At long contexts the attention cache, not weights, dominates memory — it can be quantized too.

Why FP8 Is a Recent Development

Unlike INT8, which has been used for inference quantization for years, FP8 as a training and inference format is newer, enabled specifically by hardware support introduced in NVIDIA’s Hopper architecture (H100 GPUs) with dedicated FP8 tensor cores. FP8 comes in two common variants: E4M3 (4 exponent bits, 3 mantissa bits), which favors precision over range and is typically used for weights and activations, and E5M2 (5 exponent bits, 2 mantissa bits), which favors range over precision and is often used for gradients, which can have a wider dynamic range during training. Using FP8 for both training and inference — rather than just inference, as with INT8 — is what distinguishes it as a genuinely new capability rather than an incremental improvement on existing quantization techniques.

Practical Impact

Because FP8 tensor core throughput on Hopper-class GPUs is roughly double that of BF16, models trained or served with FP8 can see substantial speedups, and this is a major part of why FlashAttention-3 and other recent inference optimizations specifically target FP8 support — it’s currently one of the most impactful levers available for reducing both training cost and inference latency on modern GPU hardware, provided a workload’s accuracy can tolerate the reduced precision.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.