Embeddings and Vector Similarity

Turning text into vectors of numbers that capture meaning, and comparing them with cosine similarity and dot product - the foundation of semantic search and RAG.

What is it?

A vector is simply a list of numbers, like [0.2, -1.3, 0.7]. You can picture a vector with two numbers as an arrow (or a point) on a flat map, three numbers as a point in 3D space, and a vector with 384 or 1024 numbers as a point in a space with that many dimensions - impossible to picture, but the maths is the same.

An embedding is a vector produced by a model (an embedding model) to represent a piece of text - a word, sentence, paragraph or document - such that texts with similar meaning get vectors that are close together. 'How do I reset my password?' and 'I forgot my login credentials' share almost no words, but their embeddings are near each other; 'best pizza in Naples' is far away from both.

This is what makes semantic search possible: instead of matching keywords, you embed every document once, embed the user's query, and return the documents whose vectors are closest to the query vector. It is the retrieval step at the heart of RAG.

Measuring closeness. The common measures:

  • Dot product: multiply matching positions and add them up: a1*b1 + a2*b2 + .... Bigger means more similar, but it is also affected by the vectors' lengths.
  • Cosine similarity: the dot product divided by both vectors' lengths. It measures the angle between vectors and ignores length. It ranges from -1 (opposite) through 0 (unrelated) to 1 (same direction). For text embeddings, values are usually compared relative to each other rather than read as absolute scores.
  • Euclidean distance: the straight-line distance between the points. Smaller means more similar.

Normalization means scaling a vector so its length is exactly 1 (divide every number by the vector's length). For normalized vectors, cosine similarity and dot product give the same number, and ranking by Euclidean distance gives the same order. That is why many systems normalize embeddings once and then use the fast dot product.

Dimensions are how many numbers each vector has. More dimensions can capture more nuance but cost more storage and compute. Each embedding model has a fixed output size, and vectors from different models cannot be compared - a vector only has meaning relative to other vectors from the same model. If you change embedding models, you must re-embed everything.

Where embeddings come from. An embedding model is a neural network (often a transformer, like an LLM but trained differently) trained on pairs of texts that should be close (a question and its answer, a title and its article) and pairs that should be far apart. Training pushes related texts together and unrelated ones apart in the vector space. You do not train one yourself: you use an existing model, either a local open-source one (for example via the sentence-transformers library) or a hosted embeddings API from providers such as Voyage AI, OpenAI or Cohere. Anthropic does not offer its own embedding model.

Embeddings capture meaning well but are weaker at exact matches: product codes, error numbers, rare names. That is why production search often combines embeddings with keyword search (see retrieval-strategies).

Explain like I'm 10

Imagine a huge library where books are not shelved alphabetically but by what they are about: cookbooks in one area, Italian cookbooks in one corner of that area, pasta books right next to each other. An embedding is a book's exact location (its coordinates) in that library. To find books about a question, you work out where the question would be shelved and grab the nearest books. Cosine similarity asks 'are these two books in the same direction from the entrance?' regardless of how far down the aisle each one is.

Examples

Dot product, length, normalization and cosine similarity

const dot = (a, b) => a.reduce((s, x, i) => s + x * b[i], 0);
const length = a => Math.sqrt(dot(a, a));
const normalize = a => a.map(x => x / length(a));
const cosine = (a, b) => dot(a, b) / (length(a) * length(b));

const a = [1, 2, 3];
const b = [2, 4, 6];      // same direction as a, twice as long
const c = [3, -1, 0];     // a different direction

console.log("dot(a,b)    =", dot(a, b));
console.log("cosine(a,b) =", cosine(a, b).toFixed(3), "(same direction -> 1)");
console.log("cosine(a,c) =", cosine(a, c).toFixed(3));

const na = normalize(a), nb = normalize(b);
console.log("normalized a =", na.map(x => x.toFixed(3)).join(", "), "| length", length(na).toFixed(3));
console.log("dot of normalized a,b =", dot(na, nb).toFixed(3), "= cosine");

Once vectors are normalized to length 1, the dot product IS the cosine similarity. This is why vector databases often store normalized vectors and use the cheaper dot product.

Semantic search with hand-made 'meaning' vectors

// Pretend embedding model with 4 dimensions we can read:
// [about-accounts, about-money, about-food, about-travel]
// Real models learn hundreds of dimensions that are not human-readable.
const docs = {
  "How to reset your password":            [0.9, 0.1, 0.0, 0.0],
  "Changing the email on your login":      [0.8, 0.0, 0.0, 0.1],
  "Refund and billing questions":          [0.2, 0.9, 0.0, 0.0],
  "Best pasta restaurants near the office": [0.0, 0.1, 0.9, 0.2],
  "Booking flights for business trips":    [0.0, 0.3, 0.1, 0.9],
};
const queries = {
  "I forgot my login credentials": [0.85, 0.05, 0.0, 0.0],
  "Where can I get lunch?":        [0.0, 0.1, 0.95, 0.1],
};

const dot = (a, b) => a.reduce((s, x, i) => s + x * b[i], 0);
const cosine = (a, b) => dot(a, b) / Math.sqrt(dot(a, a) * dot(b, b));

for (const [q, qv] of Object.entries(queries)) {
  const ranked = Object.entries(docs)
    .map(([title, v]) => [title, cosine(qv, v)])
    .sort((x, y) => y[1] - x[1]);
  console.log("Query:", q);
  for (const [title, score] of ranked.slice(0, 3)) console.log("  " + score.toFixed(3), title);
}

'I forgot my login credentials' shares no keywords with 'How to reset your password', yet it ranks first because their vectors point the same way. A real embedding model produces such vectors automatically from the text.

Bag-of-words vectors: why learned embeddings are needed

// The simplest 'embedding': count words over a fixed vocabulary.
const docs = ["reset your password", "forgot my login credentials", "change your password today"];
const query = "I forgot my password";

const tokenize = s => s.toLowerCase().split(/\W+/).filter(Boolean);
const vocab = [...new Set([...docs, query].flatMap(tokenize))];
const embed = s => {
  const words = tokenize(s);
  return vocab.map(w => words.filter(x => x === w).length);
};
const dot = (a, b) => a.reduce((s, x, i) => s + x * b[i], 0);
const cosine = (a, b) => { const d = Math.sqrt(dot(a, a) * dot(b, b)); return d ? dot(a, b) / d : 0; };

console.log("vocabulary (" + vocab.length + " dims):", vocab.join(", "));
const q = embed(query);
for (const d of docs) console.log(cosine(q, embed(d)).toFixed(3), d);
console.log("'login credentials' vs 'password':", cosine(embed("login credentials"), embed("password")).toFixed(3));

Word counts only match exact words: 'login credentials' and 'password' score 0 even though they mean nearly the same thing. Learned embeddings fix this by placing synonyms and paraphrases close together. Word-count ideas are still useful, though - BM25 keyword search builds on them (retrieval-strategies).

Real embeddings with sentence-transformers (Python)

# pip install sentence-transformers numpy
import numpy as np
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("all-MiniLM-L6-v2")   # small local model, 384-dimensional vectors

docs = [
    "How to reset your password",
    "Refund and billing questions",
    "Best pasta restaurants near the office",
    "Booking flights for business trips",
]
doc_vecs = model.encode(docs, normalize_embeddings=True)       # shape (4, 384)
print("shape:", doc_vecs.shape)

query = "I forgot my login credentials"
q_vec = model.encode([query], normalize_embeddings=True)[0]   # shape (384,)

scores = doc_vecs @ q_vec        # dot product = cosine similarity (vectors are normalized)
for i in np.argsort(-scores):
    print(f"{scores[i]:.3f}  {docs[i]}")

Embed documents once (store the vectors), embed each query at search time, and rank by dot product. Exact scores depend on the model; what matters is the ranking. The same pattern scales to millions of documents with a vector database (vector-databases).

How it works

An embedding model reads the text (as tokens), runs it through its network, and combines the token representations into one fixed-size vector (for example by averaging them). The output is usually a few hundred to a few thousand numbers.

Training uses contrastive learning: show the model a text and a related text (a question and its answer), plus unrelated texts. The loss rewards high similarity for related pairs and low similarity for unrelated ones. After training on many such pairs, meaning is encoded as direction in the vector space.

Search with embeddings: (1) split documents into chunks, (2) embed each chunk and store vector + text + metadata, (3) embed the query with the same model, (4) find the nearest stored vectors (top-k), (5) return their texts. For small collections you can compare against every vector (exact search); for large ones, vector databases use approximate nearest-neighbour indexes to stay fast.

Some embedding models expect different handling for queries versus documents (for example an instruction prefix or a separate input type). Read the model's documentation and apply it consistently.

     meaning space (2 of many dimensions)
   accounts ^
            |  * reset password
            |  * forgot login   <- query lands here
            |
            |               * refunds/billing
            |
            +------------------------------> money
          * pasta places      * flights

 query vector . each doc vector -> scores
 top-k closest -> returned as search results

Why does it exist?

Keyword search fails when users describe things in different words than the documents use. Embeddings give computers a numeric notion of meaning, so search, recommendation, clustering, deduplication and classification can work on what text is about rather than the exact words it contains.

When to use it

Use embeddings for semantic search and RAG retrieval, finding similar items (similar tickets, duplicate questions), clustering documents by topic, and as features for lightweight classifiers. Normalize vectors and use dot product or cosine consistently with how your vector store is configured.

When not to use it

Do not rely on embeddings alone for exact identifiers (SKUs, error codes, names) - combine with keyword search or filters. Do not compare vectors from different embedding models. Do not use embeddings when a structured database query answers the question exactly ('orders over $100 last week').

Common mistakes

  • Mixing vectors from two different embedding models (or versions) in the same index.

  • Forgetting to re-embed all documents after switching embedding models.

  • Using dot product on un-normalized vectors when cosine similarity was intended.

  • Treating a similarity score as an absolute probability ('0.8 means 80% relevant') instead of a relative ranking signal.

  • Embedding whole long documents as one vector, which blurs many topics together; chunk them first.

  • Expecting embeddings to match exact codes or rare names reliably.

  • Ignoring a model's query/document conventions (prefixes or input types) where the model requires them.

Practice exercises

  1. Easy:

    By hand, compute the cosine similarity of [1, 0] and [1, 1]. Then check it with the demo code.

  2. Easy:

    Add two documents and one query to the hand-made semantic search demo and predict the ranking before running it.

  3. Medium:

    Write euclidean(a, b) and show that for normalized vectors, ranking by smallest Euclidean distance gives the same order as ranking by largest cosine similarity.

  4. Medium:

    Using sentence-transformers, embed 20 FAQ questions and find the closest pair (likely duplicates). Print the top 3 pairs with their scores.

  5. Hard:

    Build a tiny semantic search CLI in Python: load a text file of paragraphs, embed them once and save the vectors with numpy, then answer queries by printing the top-3 paragraphs with scores.

Interview questions

What is an embedding?

A fixed-length vector produced by a model to represent an input such as text, arranged so that semantically similar inputs have nearby vectors. It enables similarity search over meaning rather than keywords.

Cosine similarity vs dot product - what's the difference?

Dot product depends on both direction and length; cosine similarity divides by the lengths, so it depends only on direction. For normalized (unit-length) vectors they are identical, which is why systems often normalize and use dot product.

Why can't you compare embeddings from two different models?

Each model defines its own vector space; the dimensions have no shared meaning across models, and sizes often differ. Similarity is only meaningful between vectors from the same model.

Where do embeddings fall short?

Exact matching of identifiers, rare names and numbers, negation and fine-grained details. Hybrid search (embeddings plus keyword/BM25) and metadata filters address this.

How are embedding models trained, roughly?

With contrastive learning on pairs of related texts (and unrelated negatives), so related texts are pulled together and unrelated ones pushed apart in vector space.

What happens to your index if you upgrade the embedding model?

All stored vectors must be regenerated with the new model, and queries must use the same new model; otherwise similarities are meaningless. Plan re-indexing as a migration.