INT8 Precision
An 8-bit integer data format used to reduce model VRAM footprints and accelerate compute stages.
Last reviewed: July 25, 2026
INT8 is an 8-bit integer numeric format used to represent model weights (and sometimes activations) with roughly a quarter of the memory footprint of standard 32-bit floating point, making it one of the most aggressive commonly used precision levels for deploying large models on memory-constrained hardware. Where BF16 and FP16 remain floating-point formats used mainly during training, INT8 is primarily an inference-time optimization applied through post-training quantization.
How Weights Get Converted to INT8
Since neural network weights are naturally floating-point values, converting them to 8-bit integers requires a quantization scheme that maps the floating-point range to the 256 discrete integer values INT8 can represent (typically -128 to 127). This mapping uses a scale factor (and sometimes a zero-point offset) computed per-tensor or per-channel, and the model’s outputs are dequantized back to floating point for the parts of a computation where precision matters most.
The Accuracy Tradeoff
INT8’s coarse resolution — 256 possible values compared to the millions representable in FP16 — introduces quantization error, but modern quantization techniques (including calibration against representative data, and per-channel rather than per-tensor scaling) keep this error small enough that INT8 quantized LLMs often show only a small, sometimes negligible, drop in benchmark accuracy compared to their FP16 originals, while cutting memory usage roughly in half again versus FP16 and running noticeably faster on hardware with dedicated INT8 acceleration.
Where It’s Used
INT8 quantization is common for deploying LLMs to edge devices, mobile hardware, and cost-sensitive high-throughput serving, where the memory and speed gains outweigh a small accuracy cost — libraries like bitsandbytes and frameworks like TensorRT-LLM provide production-ready INT8 quantization support.
GPTQ and AWQ: Popular INT8/4-bit Methods
Beyond generic INT8 quantization, specialized methods like GPTQ (Generative Pre-trained Transformer Quantization) and AWQ (Activation-aware Weight Quantization) have become popular specifically for quantizing LLMs, since they account for the particular sensitivity patterns of transformer weights rather than applying uniform quantization blindly. AWQ in particular identifies which weights are most sensitive to quantization error based on the magnitude of activations that pass through them, and preserves higher precision selectively for those weights while quantizing the rest more aggressively — a targeted approach that tends to outperform naive uniform quantization at the same average bit-width, and is widely supported in serving frameworks like vLLM and TensorRT-LLM for production deployment of quantized open-weight models.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.