Vector Embeddings
High-dimensional coordinate lists that capture the semantic meanings of words, sentences, or images.
Last reviewed: July 25, 2026
Vector embeddings are the numeric representations that make semantic search, retrieval-augmented generation, and recommendation systems possible. An embedding model converts a piece of text, an image, or another input into a fixed-length list of floating-point numbers — typically somewhere between 256 and 4096 dimensions — positioned in a high-dimensional space such that inputs with similar meaning end up close together, and dissimilar inputs end up far apart.
Why This Works
Embedding models are trained (often via contrastive learning) so that semantically related pairs are pulled closer together in vector space and unrelated pairs are pushed apart. The resulting geometry captures meaning in a way that raw text matching cannot: the embeddings for “car” and “automobile” end up close together even though the words share no characters, while “car” and “carpet” — which share a prefix but not a meaning — end up far apart.
How They’re Used
Once text is converted to embeddings, similarity between any two pieces of text becomes a geometric calculation — typically cosine similarity or dot product — rather than a string comparison. This is the foundation of vector databases and semantic search: a user’s query is embedded, then compared against a pre-computed index of document embeddings to find the closest matches, forming the retrieval half of retrieval-augmented generation pipelines. The same technique underlies recommendation systems (finding items with embeddings similar to what a user has liked) and deduplication (finding near-duplicate content by embedding distance).
Common embedding models include OpenAI’s text-embedding-3 family, Cohere’s embed models, and open-weight options like BGE and Nomic Embed, which vary in dimension count, language coverage, and the maximum input length they can embed in a single pass.
Embedding Model Selection Considerations
Not all embedding models are interchangeable for a given use case: they vary in output dimension (which affects storage and search cost), maximum input length (how much text can be embedded in a single call before truncation), language coverage, and whether they’re optimized for symmetric search (matching similar documents to each other) versus asymmetric search (matching short queries to longer documents, the more common RAG scenario). Mixing embeddings from different models within the same vector index is generally not meaningful, since each model’s embedding space has its own learned geometry — a query embedded with one model cannot be reliably compared against document embeddings produced by a different model, even if both models were trained on similar underlying data.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.