What is a vector database? How vector search works, when you need one, and FAISS vs pgvector vs Qdrant vs Pinecone

A vector database stores embeddings and finds the nearest ones fast. How vector search and ANN indexes work, when a library or Postgres is enough, and how to choose between FAISS, pgvector, Qdrant, Chroma, Weaviate, Pinecone and Milvus.

Presenting to a room
Presenting to a room

Key takeaways

  • A vector database stores embedding vectors and answers nearest-neighbour queries quickly, usually with an approximate index such as HNSW.
  • You do not need one for a prototype. FAISS in memory or pgvector inside the Postgres you already run covers most projects under a few million vectors.
  • Choose by operations, not benchmarks. Filtering, multi-tenancy, backups and who is on call matter more than a few milliseconds.
  • Vector search alone misses exact terms. Pair it with keyword search and fuse the results.
  • Vector databases are not dead and not magic. They are an index. Retrieval quality still comes from chunking, embeddings and reranking.

What is a vector database?

A vector database is a system that stores embedding vectors, lists of numbers that represent the meaning of text, images or audio, and answers one question fast: which stored vectors are closest to this query vector? It does this with an approximate nearest-neighbour index, plus the ordinary database features around it: metadata filters, updates, persistence, backups and an API. It is the retrieval layer under most RAG systems, semantic search products and recommendation engines.

Also asked as: what is vector file · scalar vs vector · what is vector format · what is a vector quantity · what is a vector in physics · what is vector database and how does it work · scalar vs vector quantities in physics · scalar vs vector quantity · vector database examples · vector database roadmap · what is vector database and how it works · what is vector database example · what is vector database in simple terms · what is vector database and embedding

Also asked as: what is vector database · what is vector database in ai · what is vector database in llm · what is a vector db · vector database meaning · vector database explained · vector store vs vector database · why vector database

The word "database" flatters some of the products in the category. Several are libraries with a file on disk; others are full distributed systems. What they share is the index. Understanding that index, and what it cannot do, is what lets you pick the boring option and spend your time on retrieval quality instead.

I have shipped vector search four ways in production: FAISS inside a Python process, pgvector inside Postgres, a managed service, and a hybrid where the keyword index did half the work. The pgvector setup is my default for client work, and this page explains why and when I deviate.

The vector database is the least important choice in a RAG system and the one people spend the longest on. Chunking and reranking decide whether the right passage comes back. The index just decides how fast. Pranjul Rathour, from building RAG.NextUpgrad

How does vector search work?

Every item is embedded once into a vector. At query time the question is embedded with the same model, and the system finds the stored vectors with the smallest distance or largest cosine similarity to it. Exact search compares against every vector, which is fine up to a few hundred thousand items. Beyond that, an approximate index such as HNSW or IVF narrows the search to a small neighbourhood, trading a little recall for a large speed-up.

Also asked as: how does vector database search work · does faiss support hybrid search · course on vector similarity search and faiss · how to speed up similarity search in recommender systems (using faiss · similarity search algorithms: pinecone vector index vs. faiss vs. nmslib · benchmarks of similarity search algorithms: pinecone vs. faiss vs. nmslib

Also asked as: what is vector search · how vector search works · what is vector search in ai · how do vector databases work · vector search vs semantic search · vector similarity search explained · what is nearest neighbor search · how does similarity search work

Dense retrieval, the practice of finding passages by embedding similarity, was shown by Karpukhin et al. to beat keyword retrieval for open-domain question answering [4], and it is the reason the category exists. Vector search and semantic search are used interchangeably in marketing; strictly, vector search is the mechanism and semantic search is the goal.

Which distance metric should I use?

Use the metric the embedding model was trained for, which its documentation states. Most modern text-embedding models are trained for cosine similarity or, with normalised vectors, dot product, which are equivalent once vectors are unit length. Euclidean distance is common in image search. Mixing metrics silently ruins ranking, so set it once and normalise vectors at index time.

Also asked as: cosine similarity vs dot product · cosine vs euclidean vector search · which distance metric for embeddings · vector similarity metrics

What is an ANN index, and what are HNSW and IVF?

An approximate nearest-neighbour index organises vectors so a query does not have to compare against all of them. HNSW builds a layered graph of vectors linked to their neighbours and searches by hopping through it, which gives high recall with fast queries at the cost of memory. IVF clusters vectors into cells and searches only the closest few cells, which uses less memory and suits very large collections. Product quantisation compresses vectors so more fit in memory, at a small accuracy cost.

Also asked as: what is faiss index · how faiss index works · where to store faiss index · does faiss support hnsw · difference between faiss and hnsw · faiss vs hnsw · how to add index to python faiss incrementally · how to write a faiss index to memory

Also asked as: what is hnsw · hnsw vs ivf · what is ann search · approximate nearest neighbor explained · hnsw index explained · ivf flat vs hnsw · product quantization vector search

HNSW comes from Malkov and Yashunin [2], product quantisation from Jégou, Douze and Schmid [3], and FAISS, which implements all of these and more, from Johnson, Douze and Jégou [1]. ANN-Benchmarks publishes reproducible recall-versus-throughput curves for the major algorithms, and it is the place to look before believing any vendor's chart [5].

Do I need a vector database at all?

Not for a prototype, and often not for production. Up to a few hundred thousand chunks, a flat or HNSW index in FAISS inside your application, or a pgvector column in the Postgres you already run, gives you vector search with no new service to operate. You need a dedicated vector database when the index outgrows one machine's memory, when you need rich filtering with multi-tenancy at scale, or when several services must share the index with strict latency targets.

Also asked as: are we at peak vector database · do we really need a specialized vector database · vecstore vs pinecone: when you don't need a raw vector database · what is a vector database (and do you actually need one · where is my faiss vectorstore saved at · do you need a vector database? postgres vs dedicated in 2026 · how can i store an objectindex in my file system so i don't need to recreate it every time with llamaindex

Also asked as: do i need a vector database · do i need a vector database for rag · when to use vector database · vector database vs traditional database · vector database vs sql · can i use postgres as vector database · are vector databases necessary · vector database alternatives

The size of RAG.NextUpgrad is the useful example: a production RAG platform with hybrid retrieval, reranking, streaming answers and a test suite runs in about 220 MB of memory on a free tier [16]. Nothing about RAG requires a heavy retrieval service; the heavy services are for heavy data.

Is Postgres with pgvector good enough for RAG?

For most RAG products, yes. pgvector adds a vector column type, exact search, and HNSW and IVF indexes to PostgreSQL, so your chunks, their metadata, your users and your permissions live in one database and one transaction. Filtering is a WHERE clause. Backups are the backups you already take. It stops being enough when you need to shard across machines or when tail latency under heavy write load matters more than simplicity.

Also asked as: is faiss good for production · faiss vs pgvector

Also asked as: pgvector vs pinecone · pgvector vs qdrant · is pgvector good enough · postgres vector search · pgvector performance · pgvector hnsw · supabase vector search

pgvector's documentation covers index types, distance operators and tuning [6]. Supabase exposes it on their free tier, which makes it the easiest path for a student project that wants a real database from day one.

Which vector database should I use?

Choose by how the system will be operated. FAISS if the index lives inside one application and you control the code. pgvector if you run Postgres and want one database. Chroma for local prototypes with a friendly Python API. Qdrant, Weaviate or Milvus when you need a standalone service with filtering, multi-tenancy and horizontal scale, self-hosted or managed. Pinecone when you want a fully managed service and are fine paying for it. MongoDB Atlas or Elasticsearch when your data already lives there.

Also asked as: why use faiss vector database · what vector database does openai use · what vector database to use · how to use vector database with llm · can i use mongodb as vector database · can i use mysql as vector database · can i use redis as vector database · can i use sqlite as a vector database · how to use vector database in n8n · how to use pinecone vector database · which vector database should i use? a comparison cheatsheet · postgresql as a vector database: when to use pgvector vs pinecone vs weaviate · how to use qdrant vector database in n8n · what jpa + hibernate data type should i use to support the vector extension in a postgresql database

Also asked as: best vector database · best vector database for rag · best vector database for llm · best open source vector database · best free vector database · faiss vs qdrant vs pgvector · qdrant vs pinecone · chroma vs faiss · weaviate vs qdrant · milvus vs qdrant · pinecone vs weaviate · vector database comparison

Each product's own documentation is the source for the cells above and the place to check current features, since this category changes fast [6][7][8][9][10][11][12]. MongoDB Atlas Vector Search and Elasticsearch's kNN search are the two "you already have it" options for teams on those stacks [13][14].

Is Pinecone free? Which vector databases are free?

FAISS, pgvector, Chroma, Qdrant, Weaviate and Milvus are open source and free to self-host. Pinecone, Qdrant Cloud, Weaviate Cloud and Zilliz Cloud offer free tiers with limits on vectors, storage or projects that are enough for a student project or a prototype. Check the current limits on each pricing page; they change, and this page will not be updated faster than the vendors.

Also asked as: which is better faiss or pinecone · difference between faiss and pinecone · faiss vs pinecone · how faiss differs from other vector databases

Also asked as: is pinecone free · free vector database · best free vector database · qdrant free tier · vector database pricing · cheapest vector database

How do I create and query a vector database?

Embed your chunks with one model, store each vector with its text and metadata, build an index, and at query time embed the question with the same model and ask for the top k neighbours. With pgvector that is a table with a vector column, an HNSW index and an ORDER BY on the distance operator; with FAISS it is an index object and a search call. The work is in the embedding and the metadata, not the calls.

Also asked as: how to create vector database in postgresql · does faiss create embeddings · can i embed my graphql schema into a vector database like pinecone and get llm to generate the query/mutations

Also asked as: how to create vector database · how to build a vector database · how to query vector database · how to use faiss vector database · how to use qdrant vector database · vector database tutorial · vector database python · how to store embeddings in database

The pgvector version, which is the whole thing:

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE chunks (
  id        bigserial PRIMARY KEY,
  doc_id    text NOT NULL,
  tenant_id text NOT NULL,
  content   text NOT NULL,
  embedding vector(768)          -- dimensions of your embedding model
);
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);

-- query: nearest chunks for one tenant
SELECT id, doc_id, content, 1 - (embedding <=> $1) AS score
FROM chunks
WHERE tenant_id = $2
ORDER BY embedding <=> $1
LIMIT 30;

And the FAISS version for an in-process index:

import faiss, numpy as np

vectors = np.asarray(embeddings, dtype="float32")     # shape (n, d), L2-normalised
index = faiss.IndexFlatIP(vectors.shape[1])            # exact inner product = cosine on unit vectors
index.add(vectors)
scores, ids = index.search(np.asarray([query_vec], dtype="float32"), k=30)

Two rules that save weeks: store the embedding model's name next to the vectors, because vectors from different models cannot be compared, and never re-embed on every query what you could have embedded once at index time.

How do I filter vector search by metadata?

Filter before the search when the filter is selective, for example by tenant, so the index only searches that user's vectors; filter after when the filter is loose. Dedicated vector databases handle this with payload indexes; in pgvector a WHERE clause and a suitable B-tree index do it, with the caveat that HNSW plus a strict filter can return fewer than k results, which is why over-fetching and re-checking is common.

Also asked as: can faiss store metadata · does faiss support metadata filtering · is there a useful technique to map faiss ids to the appropriate metadata

Also asked as: vector database metadata filtering · filtered vector search · multi tenant vector database · pre filtering vs post filtering vector search · vector search with filters

Why does vector search miss obvious results?

Because embeddings capture meaning, not exact strings. A product code, a person's name, an error number or a rare acronym has weak meaning to the model and blurs into neighbours. Keyword search finds those exactly. The fix is hybrid retrieval: run BM25 and vector search on every query and fuse the ranked lists, most simply with reciprocal rank fusion, which needs no score normalisation [15]. Chunking that strips context, and mixing embedding models, are the other two common causes.

Also asked as: vector search not working · vector search returns wrong results · limitations of vector database · vector search vs keyword search · hybrid search vector database · problems with vector databases · vector database drawbacks

In RAG.NextUpgrad the hybrid path exists for exactly this reason: support questions arrive as "error E-4021" as often as "why did my upload fail", and only one of those is a semantic query [16].

On the mic
On the mic

Are vector databases dead, or still relevant?

Still relevant, and less special than in 2023. Long context windows let you paste more text into a prompt but do not tell you which text to paste; retrieval does. What changed is that vector search became a feature of ordinary databases, Postgres, MongoDB, Elasticsearch, so the standalone category has to justify itself on scale and operations. For most teams, the answer is "use the vector feature of the database you have".

Also asked as: are vector databases dead · are vector databases still relevant · future of vector databases · do we still need vector databases · vector database hype

What are common vector database interview questions?

Explain what an embedding is and why cosine similarity is used. Explain exact versus approximate search and the recall trade-off. Compare HNSW and IVF. Explain why keyword search is still needed. Describe how you would filter by tenant safely. Say when you would choose pgvector over a dedicated service, and defend it. The last one separates candidates who have shipped from candidates who have read.

Also asked as: vector database interview questions · vector search interview questions · embeddings interview questions · rag interview questions vector database

Where should I start?

Build a small RAG system over a document set you know, embed with one open model, index in pgvector on a free Supabase project or in FAISS locally, and measure recall on twenty of your own questions. Then add keyword search and watch recall rise. That one experiment teaches more about vector databases than any comparison article, including this one.

Also asked as: vector database tutorial for beginners · learn vector databases · vector database course · vector database project ideas · how to learn vector search

If your college club wants a hands-on session on retrieval systems, that is one of the talks I offer. Email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Johnson, Douze & Jégou, Billion-scale similarity search with GPUs (FAISS, 2017)arxiv.org
  2. Malkov & Yashunin, Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs (HNSW, 2016)arxiv.org
  3. Jégou, Douze & Schmid, Product Quantization for Nearest Neighbor Search (2011)ieeexplore.ieee.org
  4. Karpukhin et al., Dense Passage Retrieval for Open-Domain Question Answering (2020)arxiv.org
  5. ANN-Benchmarks: benchmarking approximate nearest neighbour algorithmsann-benchmarks.com
  6. pgvector: open-source vector similarity search for Postgresgithub.com
  7. FAISS wiki and documentationgithub.com
  8. Qdrant documentationqdrant.tech
  9. Chroma documentationdocs.trychroma.com
  10. Weaviate documentationweaviate.io
  11. Milvus documentationmilvus.io
  12. Pinecone documentationdocs.pinecone.io
  13. MongoDB Atlas Vector Search documentationmongodb.com
  14. Elasticsearch dense vector field and kNN search documentationelastic.co
  15. Cormack, Clarke & Buettcher, Reciprocal Rank Fusion (SIGIR 2009)plg.uwaterloo.ca
  16. RAG.NextUpgrad source code, Pranjul Rathourgithub.com
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on RAG & retrieval

All RAG & retrieval guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur