Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

Product Quantization (PQ)

A lossy vector compression technique that reduces database RAM footprints.

Last reviewed: July 25, 2026

Product Quantization (PQ) is a vector compression technique used in large-scale vector search systems to dramatically reduce the memory footprint of stored embeddings, at the cost of some retrieval accuracy — it’s the compression component behind the widely used IVF-PQ indexing method.

How It Works

Instead of storing a vector’s full sequence of floating-point values, product quantization splits each vector into several smaller sub-vectors, and for each sub-vector position, learns a small “codebook” of representative values (typically 256 of them, via a clustering algorithm like k-means) from the actual data. Each sub-vector in the dataset is then replaced by the index of its nearest codebook entry — a single byte, since 256 fits in 8 bits — rather than storing the original floating-point sub-vector directly.

Why This Saves So Much Memory

A typical 768-dimensional embedding stored as 32-bit floats requires about 3KB per vector. Splitting that vector into, say, 96 sub-vectors of 8 dimensions each, and replacing each sub-vector with a single codebook index byte, reduces the stored representation to just 96 bytes — roughly a 30x compression ratio. This compounds enormously at scale: a billion-vector index that would require terabytes of memory in raw float form can potentially fit in a fraction of that with product quantization applied.

The Accuracy Tradeoff

Because each sub-vector is approximated by its nearest codebook entry rather than stored exactly, distance calculations between quantized vectors are themselves approximate, meaning searches over a PQ-compressed index have somewhat lower recall (a higher chance of missing the true nearest neighbor) than searching uncompressed vectors. This tradeoff is generally accepted for use cases where the alternative — needing enough memory to store billions of raw vectors — simply isn’t practical or cost-effective.

Combining Product Quantization With Reranking

Because product quantization introduces approximation error into distance calculations, production systems often combine it with a reranking step: an initial PQ-based search retrieves a larger candidate set quickly using compressed vectors, and then a final pass re-scores just those candidates using their original, uncompressed representations (stored separately, often on cheaper storage than the compressed index needs to be held in memory) for a more accurate final ranking. This two-stage pattern — fast approximate search followed by precise reranking on a small candidate set — is a recurring architecture across large-scale vector search systems, appearing in similar form whether the underlying compression technique is product quantization, scalar quantization, or another approximation method entirely.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.