What are embeddings in AI? Text embeddings, vector similarity, choosing a model, and how embeddings power RAG and semantic search

An embedding turns text, an image or audio into a vector so that similar things are close together. How embedding models are trained, what dimensions and similarity mean, how embeddings differ from tokens, how to choose and evaluate a model with MTEB, and how they are used in RAG, search, clustering and recommendation.

Pranjul Rathour
Pranjul Rathour

Key takeaways

  • An embedding is a fixed-length vector representing meaning. Two texts about the same thing produce vectors close together.
  • Embeddings are not tokens and not the LLM's output; a separate model produces them, and vectors from different models cannot be compared.
  • Cosine similarity on normalised vectors is the standard comparison. Dimensions trade storage and speed against nuance.
  • Choose a model by retrieval benchmarks for your language and domain on MTEB, then verify on your own queries. Re-embed everything when you switch.
  • Embeddings power RAG retrieval, semantic search, deduplication, clustering, classification and recommendations, but they miss exact terms, so pair them with keyword search.

What are embeddings in AI?

An embedding is a list of numbers, a vector, that a model produces to represent the meaning of an input: a sentence, a paragraph, an image, a sound clip. The model is trained so that inputs with similar meaning produce vectors that are close together and dissimilar inputs produce vectors far apart. Once text is a vector, meaning becomes geometry: you can measure similarity with a distance, search by nearest neighbour, cluster, and feed vectors to other models.

Also asked as: what are embeddings in ai · what is embedding in machine learning · what is text embedding · what is vector embedding · embeddings explained · what is an embedding model · embedding meaning in ai · what are embeddings in llm · word embeddings vs sentence embeddings

Word2vec showed in 2013 that simple training objectives produce vectors where analogies work as arithmetic [1]. BERT produced contextual vectors for every token [2], and Sentence-BERT trained networks so that a whole sentence's vector could be compared with cosine similarity directly, which is the form every RAG system uses [3]. Dense passage retrieval then showed embeddings beating keyword search for question answering [4].

In every retrieval system I have built, the embedding model is chosen once, early, and it decides what "similar" means for the life of the index. This page is how I make that choice, and what I tell students who confuse embeddings with tokens.

An embedding model is a definition of similarity, frozen into numbers. Pick the one whose idea of similar matches your users' idea of relevant. Pranjul Rathour

How do embedding models work?

A transformer reads the input and produces a vector for every token; a pooling step, usually the mean or a special token's vector, collapses them to one vector for the whole input; a final normalisation makes it unit length. Training uses pairs: a query and a passage that answers it, or two sentences that mean the same, pushed together, with unrelated pairs pushed apart, a contrastive objective. Millions of such pairs teach the model which differences matter and which do not.

Also asked as: how do embeddings work · how are embeddings created · how embedding models are trained · contrastive learning embeddings · sentence transformers explained · how does text embedding work · embedding layer vs embedding model

Two terms collide here. An "embedding layer" inside any neural network maps token ids to vectors as the first step of processing; an "embedding model" is a whole network whose output is a vector for the input. The second is what RAG and search use.

What is the difference between embeddings and tokens?

Tokens are the pieces of text a model reads, word fragments mapped to integer ids; there are as many tokens as pieces. An embedding is one vector representing an entire input's meaning. A 500-word passage might be 650 tokens and exactly one embedding. Tokens are how language models consume text; embeddings are how retrieval systems compare it. People search for this distinction often because both are "numbers that represent text", and they are not the same numbers.

Also asked as: embeddings vs tokens · are embeddings and tokens the same · difference between tokens and embeddings · token embedding vs sentence embedding · tokenization vs embedding · embeddings vs vectors · are embeddings and vectors the same thing

What do dimensions mean, and how many do I need?

Dimensions are the length of the vector, commonly 384, 768, 1,024 or 1,536, sometimes 3,072 or more. More dimensions can capture more nuance and cost more storage and search time. The right number is whatever your chosen model produces; it is not a knob to maximise. Matryoshka-trained models let you truncate a long vector to a shorter prefix with graceful quality loss, which is useful when storage matters [7]. Test truncation on your own retrieval set before relying on it.

Also asked as: embedding dimensions · what is embedding dimension · how many dimensions for embeddings · 768 vs 1536 embeddings · embedding size · reduce embedding dimensions · matryoshka embeddings · embedding vector length

A 1,536-dimension vector for a million chunks is about six gigabytes in 32-bit floats before any index overhead. A 384-dimension model is a quarter of that. Whether it retrieves worse for your documents is an empirical question, and it is often closer than the size suggests.

How is similarity between embeddings measured?

By a distance or similarity function between two vectors. Cosine similarity, the cosine of the angle between them, is the standard for text; on unit-length vectors it equals the dot product and ranges from minus one to one. Euclidean distance is common in image search. Use the metric the model's documentation specifies, normalise vectors at index time, and never mix metrics or models within one index.

Also asked as: cosine similarity embeddings · how to compare embeddings · cosine similarity vs dot product · euclidean distance vs cosine similarity · embedding similarity score · what is a good cosine similarity score · semantic similarity

There is no universal "good" cosine score. The same model can put unrelated texts at 0.6 and paraphrases at 0.85; another model spreads them differently. Thresholds come from your own labelled pairs, not from a blog post.

Which embedding model should I use?

Shortlist by the retrieval task on the MTEB leaderboard, filtered for your language and a size you can run [5][6]. Prefer models trained on query-to-passage pairs for RAG. Then evaluate two or three on twenty of your own questions against your own documents, measuring recall at k. Choose the smallest model that meets your target. For Hindi and mixed-language content, check multilingual models specifically; English-only leaders can collapse.

Also asked as: best embedding model · best embedding model for rag · which embedding model to use · openai embeddings vs open source · best open source embedding model · best embedding model for semantic search · multilingual embedding model · embedding model for hindi · embedding model comparison · text-embedding-3 vs bge

The MTEB paper explains the benchmark's tasks and why a single average hides large differences between tasks [5]. Sentence Transformers is the library most open models ship with [10]; commercial APIs document their models and dimensions [11]. My own decision notes on switching models in a production RAG system are on the portfolio [13].

Pranjul Rathour
Pranjul Rathour

At index time, each document chunk is embedded and stored in a vector index. At query time the question is embedded with the same model and the nearest chunks are retrieved, then reranked and passed to the language model as context. Semantic search is the same retrieval without the generation step: the results are the chunks or documents themselves. Embeddings make "how do I reset my password" match a passage titled "Credential recovery" even though they share no words.

Also asked as: embeddings in rag · how rag uses embeddings · semantic search with embeddings · embedding based search · vector search embeddings · how to use embeddings for search · embeddings for chatbot · semantic search vs keyword search

That table is why every retrieval system I ship is hybrid: dense retrieval for meaning, BM25 for exact terms, fused with reciprocal rank fusion [12].

What else can I do with embeddings?

Deduplicate near-identical documents by finding pairs above a similarity threshold. Cluster support tickets or feedback into themes. Classify text by nearest labelled examples, which needs no training run. Recommend items whose descriptions are close to what a user liked. Detect topic drift in a conversation. Match resumes to job descriptions by section. Anything that reduces to "which of these is most like that" is an embedding problem.

Also asked as: embedding use cases · what can you do with embeddings · embeddings for classification · embeddings for clustering · embeddings for recommendation · embeddings for deduplication · applications of embeddings

What are multimodal, image and code embeddings?

Multimodal models place images and text in the same vector space, so a text query can retrieve images and an image can retrieve captions; CLIP established the approach [8]. Image embeddings power visual search and duplicate detection. Code embeddings, trained on code and natural language pairs, let you search a codebase by describing what a function does. Each is a separate model with its own space; a text embedding and an image embedding from unrelated models cannot be compared.

Also asked as: image embeddings · multimodal embeddings · clip embeddings · code embeddings · audio embeddings · embeddings for images and text · visual search embeddings

What are the limitations of embeddings?

They compress meaning, so they lose detail, especially exact strings, numbers and negation: "no refund" and "refund" can land close. They inherit their training data's biases and domain, so a general model can misjudge medical or legal text. They are model-specific, so switching models means re-embedding everything. They cannot be reversed into the original text, which is a privacy advantage and an interpretability limit. Late-interaction methods such as ColBERT keep per-token vectors to recover some lost precision, at higher storage cost [9].

Also asked as: limitations of embeddings · problems with embeddings · embeddings and negation · embedding bias · can embeddings be reversed · embedding drift · when not to use embeddings

What are common embedding interview questions?

Explain what an embedding is to a non-engineer. Explain why cosine similarity is used and why vectors are normalised. Compare tokens and embeddings. Describe how you would choose an embedding model for a RAG system and how you would evaluate it. Explain why keyword search is still needed. Say what happens when you switch models. If you have built one retrieval system, these are all stories rather than definitions.

Also asked as: embedding interview questions · vector embeddings interview · rag interview questions embeddings · nlp interview questions embeddings

Where should I start with embeddings?

Install Sentence Transformers, embed a hundred sentences you wrote, and print the nearest neighbours of a few. Then embed a real document set, ask twenty questions, and measure how often the right chunk is in the top five. Swap the model and measure again. That afternoon teaches everything above. If your college wants a hands-on session on retrieval, embeddings and RAG, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Mikolov et al., Efficient Estimation of Word Representations in Vector Space (word2vec, 2013)arxiv.org
  2. Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2018)arxiv.org
  3. Reimers & Gurevych, Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks (2019)arxiv.org
  4. Karpukhin et al., Dense Passage Retrieval for Open-Domain Question Answering (2020)arxiv.org
  5. Muennighoff et al., MTEB: Massive Text Embedding Benchmark (2022)arxiv.org
  6. MTEB leaderboard, Hugging Facehuggingface.co
  7. Kusupati et al., Matryoshka Representation Learning (2022)arxiv.org
  8. Radford et al., Learning Transferable Visual Models From Natural Language Supervision (CLIP, 2021)arxiv.org
  9. Khattab & Zaharia, ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT (2020)arxiv.org
  10. Sentence Transformers documentationsbert.net
  11. OpenAI embeddings guideplatform.openai.com
  12. Cormack, Clarke & Buettcher, Reciprocal Rank Fusion (SIGIR 2009)plg.uwaterloo.ca
  13. Choosing an embedding model for RAG, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on RAG & retrieval

All RAG & retrieval guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur