QLoRA
A fine-tuning method that backpropagates gradients through frozen, 4-bit quantized base models into LoRA adapters.
Last reviewed: July 25, 2026
QLoRA (Quantized Low-Rank Adaptation) is a fine-tuning technique that combines two separate memory-saving methods — 4-bit quantization of a model’s frozen base weights, and LoRA’s low-rank trainable adapters — to make it possible to fine-tune very large language models on a single consumer or prosumer GPU, rather than requiring a multi-GPU server.
How It Combines the Two Techniques
Standard LoRA freezes a model’s pretrained weights and only trains small low-rank adapter matrices injected alongside them, which dramatically reduces the number of trainable parameters and the corresponding optimizer memory needed — but the frozen base weights still have to be loaded into GPU memory at their original precision, which for a 70-billion-parameter model can require well over 100GB even before accounting for the training process itself. QLoRA’s contribution is quantizing those frozen base weights down to 4-bit precision, cutting their memory footprint roughly in half again versus the already-reduced BF16/FP16 formats commonly used, while the small LoRA adapters being actively trained remain in full precision to avoid degrading training quality.
The NF4 Data Type
QLoRA introduced NormalFloat 4 (NF4), a 4-bit data type specifically designed around the observation that trained neural network weights tend to follow a roughly normal (bell curve) distribution — rather than spacing its representable values evenly (as a standard integer format would), NF4 places its 16 representable values at positions that are information-theoretically optimal for a normal distribution, capturing more useful precision where weight values are actually concentrated.
Practical Impact
QLoRA’s 2023 paper demonstrated that a 65-billion-parameter model could be fine-tuned on a single 48GB GPU with results comparable to full 16-bit fine-tuning, a result that significantly broadened who could practically fine-tune frontier-scale open-weight models, and QLoRA (implemented via the bitsandbytes library) remains one of the most widely used fine-tuning approaches in the open-source LLM ecosystem today.
QLoRA’s Paged Optimizers
Beyond 4-bit quantization and LoRA adapters, the original QLoRA implementation introduced “paged optimizers,” which use NVIDIA’s unified memory feature to automatically page optimizer states between GPU and CPU memory when GPU memory would otherwise be exceeded during a sudden spike in memory usage (like an unusually long sequence in a training batch). This additional mechanism prevents out-of-memory crashes during otherwise-successful training runs caused by occasional memory usage spikes, functioning as a safety valve that lets QLoRA training push closer to the actual limits of available GPU memory without the training run failing entirely the first time it encounters a slightly larger-than-typical batch.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.