A vector database is a specialized data store that keeps numerical representations—"vectors"—of items (text, images, audio, etc.) and lets you retrieve the most similar items quickly. The similarity is measured with a distance metric such as cosine similarity or Euclidean distance. In practice, you feed raw data through an encoder (often a transformer model), get a fixed‑size embedding, and write that embedding to the database. When you later query with another embedding, the engine returns the top‑k nearest vectors.

How Vector Search Works Under the Hood

1. Embedding Generation

  • Encoder – a pre‑trained model (e.g., BERT, CLIP) turns a document or image into a dense vector.
  • Dimension – typical sizes range from 128 to 1,024 floats; the choice balances expressiveness and index size.

2. Indexing Structures

Index typeTypical use caseStrengthsWeaknesses
Inverted File (IVF)Large static corporaGood recall, moderate memoryBuild time can be high
Hierarchical Navigable Small World (HNSW)Low‑latency, high‑recall needsSub‑millisecond queries, dynamic insertsHigher memory footprint
Product Quantization (PQ)Very large collections on limited RAMTiny index size, decent recallApproximation error can hurt accuracy

These structures partition the vector space so that only a fraction of vectors need to be examined for each query. The engine computes distances only for candidates in the relevant partitions, then returns the best matches.

Core Trade‑offs

  • Latency vs. Recall – Faster indexes (e.g., HNSW with low ef) give lower latency but may miss some close neighbors. Increasing the search parameter improves recall at the cost of time.
  • Memory vs. Accuracy – PQ compresses vectors heavily, saving RAM but introducing quantization error. IVF keeps full vectors, using more memory but preserving exact distances.
  • Static vs. Dynamic – Some indexes (IVF‑flat) are cheap to rebuild but slow to update. HNSW supports incremental inserts, making it better for streaming data.
  • Consistency – Most vector stores are eventually consistent; writes may not be visible to reads for a short window, which matters when you need real‑time freshness.

A Concrete Example

Imagine a product‑search feature for an e‑commerce site. Each product description is passed through a sentence‑transformer to get a 768‑dimensional vector. Those vectors are stored in a vector DB such as Milvus or Pinecone. When a user types "water‑proof hiking boots", the query string is encoded, and the DB returns the ten most similar product vectors. The application then fetches the full product records from a relational store and displays them. The whole pipeline—encoding, vector lookup, join—runs in under a second for typical traffic.

Typical Interview Questions

  1. What is a vector database and why not just use a regular relational DB? Focus on similarity search, high‑dimensional indexing, and performance.
  2. How does the index affect latency and recall? Explain the trade‑off between search parameters (ef, nprobe) and the quality‑speed curve.
  3. What are the challenges of keeping vectors in sync with source data? Talk about update latency, stale embeddings, and strategies like periodic re‑encoding or change‑data‑capture pipelines.
  4. How would you choose between IVF, HNSW, and PQ for a given workload? Mention dataset size, query latency SLA, and available memory.
  5. Can you combine vector search with traditional filters? Describe hybrid queries where you first filter by metadata (e.g., category) then run the vector search on the subset.

60‑Second Spoken Answer

"A vector database stores dense embeddings—numeric representations of text, images, or audio—so you can quickly find similar items. You generate embeddings with a model like BERT, then write the vectors into a specialized index. The index (often IVF, HNSW, or PQ) partitions the space, letting the engine examine only a small fraction of vectors for each query, which yields sub‑second latency even on millions of items. The main trade‑offs are latency versus recall and memory usage versus accuracy; you tune parameters like ef for HNSW or nprobe for IVF to hit your SLA. In practice, you might use a vector DB to power a product‑search feature: encode product descriptions, store them, and at query time retrieve the nearest vectors, then join back to the relational store for full details. Interviewers usually dig into how you’d pick an index, keep embeddings fresh, and combine vector search with filters."

How to Practice This

  1. Build a mini pipeline – Encode a small text corpus with a public transformer, load the vectors into an open‑source vector store (e.g., Milvus), and run similarity queries.
  2. Swap index types – Re‑index the same data with IVF, HNSW, and PQ; measure latency vs. recall to feel the trade‑offs.
  3. Mock interview – Use Call Assistant to rehearse the 60‑second answer aloud, then ask follow‑up questions about scaling or consistency and let the assistant keep the conversation on track.

FAQ

  • Q: Do I need a GPU to run a vector database? A: The database itself runs on CPU; only the embedding step typically benefits from a GPU. Many production setups keep encoding on a separate service.
  • Q: How large can a vector collection be? A: With efficient indexing, billions of vectors are feasible, especially when using compression techniques like PQ.
  • Q: Are vector databases ACID compliant? A: Most provide eventual consistency for writes; strong transactional guarantees are rare because the primary goal is fast similarity search.
  • Q: Can I store non‑numeric data directly? A: The DB stores only vectors and optional metadata. The original document lives elsewhere; you join on a key after the vector lookup.

Frequently asked questions

What is the difference between IVF and HNSW indexes?

IVF partitions vectors into coarse clusters and scans a few clusters per query; it balances memory use and recall. HNSW builds a navigable graph that can reach any vector in a few hops, offering lower latency but higher memory consumption.

When should I use product quantization?

Use PQ when you need to fit millions of vectors into limited RAM and can tolerate a modest drop in recall. It's common for offline recommendation systems where storage cost is a bigger concern than micro‑second latency.

How do I keep embeddings up to date after data changes?

Set up a pipeline that re‑encodes changed records (via CDC or scheduled jobs) and either overwrites the old vector or inserts a new version, depending on the DB's update semantics.

Can vector search be combined with traditional filters?

Yes. Most vector stores allow you to filter by metadata (e.g., category, date) before performing the similarity search, enabling hybrid queries that narrow the candidate set.

#concept#vector databases#interview#machine learning#search