IVF-PQ Index
A vector search index combining Inverted File indexing and Product Quantization for memory efficiency.
Last reviewed: July 25, 2026
IVF-PQ (Inverted File with Product Quantization) is a vector search indexing method designed to handle extremely large embedding collections — hundreds of millions to billions of vectors — where the memory overhead of graph-based indexes like HNSW becomes impractical, trading some search accuracy for a substantially smaller memory footprint.
How It Works
The “IVF” component partitions the entire vector space into a number of clusters (using a clustering algorithm like k-means), and each vector is assigned to its nearest cluster. At search time, rather than comparing the query against every vector in the index, the search first identifies the handful of clusters closest to the query vector, then only searches within those clusters — dramatically reducing the number of comparisons needed, at the cost of occasionally missing a true nearest neighbor that happens to sit in a cluster the search didn’t examine.
The “PQ” component, product quantization, further compresses each vector by splitting it into smaller sub-vectors and replacing each sub-vector with the index of its nearest entry in a small, pre-trained codebook, rather than storing the full-precision sub-vector. This can shrink a vector’s storage footprint by a factor of 10x or more compared to storing raw floating-point values.
The Tradeoff
Combining both techniques gives IVF-PQ a dramatically smaller memory footprint than HNSW at similar vector counts, which is what makes it practical for billion-scale collections that wouldn’t fit in memory as an HNSW graph. The cost is generally lower recall (a higher chance of missing true nearest neighbors) compared to HNSW at equivalent search speed, since both the cluster-based search pruning and the quantization-based compression each introduce their own approximation error. It’s a common choice in systems like Milvus and FAISS specifically for use cases where scale outweighs the need for maximum search precision.
Combining IVF-PQ With Reranking
Because IVF-PQ trades some accuracy for its large memory savings, production systems sometimes add an additional reranking step on top: after IVF-PQ retrieves a set of approximate candidates using the compressed representation, the system can re-score just those candidates using their original, uncompressed vectors (which may be stored separately on cheaper, slower storage) for a final, more accurate ranking. This two-stage approach — fast approximate search over compressed vectors, followed by precise reranking on a small uncompressed candidate set — recovers much of the accuracy that pure IVF-PQ search would sacrifice, while still gaining the bulk of its memory benefits for the large-scale first-pass search.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.