Post-Training Quantization (PTQ)
An offline compression technique that converts weights to lower precision after model training finishes.
Last reviewed: July 25, 2026
Post-Training Quantization (PTQ) is the process of converting an already-trained model’s weights (and sometimes activations) to a lower-precision numeric format after training has finished, without any further training or fine-tuning required. It’s the fastest and most widely used way to shrink a model’s memory footprint and increase inference speed.
How It Works
PTQ typically runs a small “calibration” step, passing a modest sample of representative data through the model to observe the actual range of values its weights and activations take on, then uses that information to choose scale factors that map the original floating-point range onto the target lower-precision format (commonly INT8 or 4-bit formats) as accurately as possible. More sophisticated PTQ methods calibrate these scale factors per-channel (a separate scale for different slices of a weight matrix) rather than a single scale for an entire tensor, which noticeably improves accuracy for a similar compression ratio.
PTQ vs. Quantization-Aware Training
Because PTQ is applied after the fact, it’s dramatically cheaper than quantization-aware training (QAT), which requires a training or fine-tuning run to let the model adapt to quantization noise during learning itself. PTQ can typically be completed in minutes to a few hours using a small calibration dataset, versus the much larger compute and time investment of a training run. The tradeoff is that PTQ generally produces a somewhat larger accuracy drop than QAT at the same target precision, particularly at aggressive compression levels like 4-bit — a gap that’s driven the development of increasingly sophisticated PTQ calibration methods aiming to close that difference without the cost of retraining.
Where It’s Used
PTQ is the default approach used by popular quantization tools like GPTQ, AWQ, and bitsandbytes’ 8-bit and 4-bit modes, and it’s the standard first choice for teams deploying an existing pretrained or fine-tuned model to memory-constrained or latency-sensitive production environments.
Per-Channel vs. Per-Tensor Calibration
A meaningful technical detail in PTQ implementation is whether scale factors are calibrated per-tensor (a single scale factor for an entire weight matrix) or per-channel (separate scale factors for different slices, typically corresponding to output channels or attention heads). Per-channel calibration generally preserves more accuracy, since different channels within the same weight matrix often have meaningfully different value distributions, and forcing them to share a single scale factor wastes precision on whichever channels don’t match that shared scale well — this is why most modern PTQ implementations default to per-channel calibration despite its slightly higher implementation complexity and marginally larger metadata overhead compared to simpler per-tensor schemes.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.