Updated Sep 10, 2026

Vector Database

A database that stores text as numerical embeddings and finds entries by meaning rather than by keyword.

Share

Listen to this lesson

Free preview · first 0:30
0:00 / 0:30

Unlock audio and more

Audio streaming, downloadable PDFs and certificates come with Plus and Pro.

What it means

An embedding turns a piece of text into a long list of numbers positioned so that similar meanings land near each other. "How do I reset my password?" and "I forgot my login" end up close together despite sharing almost no words. A vector database stores those numbers and answers one question very efficiently: what is nearest to this?

That is semantic search, and it is the retrieval half of RAG. Keyword search fails when the user's words differ from the document's words; embedding search handles that naturally. In exchange it can miss exact matches — product codes, names, precise identifiers — which is why serious systems usually combine both, an approach called hybrid search.

Why it matters

Vector storage went from a specialist tool to standard infrastructure very quickly, because every RAG system needs it. Notably, it no longer requires adopting a new database: the major general-purpose databases have added vector support, so for most teams this is a feature to switch on rather than a system to procure.

What people get wrong

That you need a dedicated one. This was true when vector search was specialist infrastructure and is mostly not true now. The major general-purpose databases have added vector support, so for most teams this is a feature to switch on in the database they already run rather than a system to procure, operate and back up separately. A dedicated vector database earns its place at large scale or under demanding latency requirements — not by default.

That semantic search replaces keyword search. It fails in the opposite direction from keyword search, which is why serious systems run both. Matching on meaning is exactly what you want for "how do I cancel" finding a document about terminating a subscription — and exactly what you do not want for an order number, a part code, an error string or a proper name, where a semantically adjacent result is a wrong result. Hybrid search is not a hedge; it is an acknowledgment that the two methods fail on different inputs.

That embeddings are interchangeable. Vectors from different embedding models are not comparable, so changing models means re-embedding your entire corpus. That makes the embedding model one of the more consequential and least reversible choices in the system — store which model produced your vectors, because a future migration will need to know.

In practice

For most applications the pragmatic starting point is vector support in the database you already run. A dedicated vector database earns its place at large scale or with demanding latency requirements — not by default.

Where this shows up

Tools and models in our catalog.

Related terms