Go deeper:
- For lexical search and the inverted index, see Elasticsearch and Elasticsearch vs Postgres Full-Text.
- For the full retrieval system around the index, see Search Engine.
Why This Matters#
A vector database is not a smarter database. It is an approximate index with a recall dial: it returns probably the nearest neighbours to a query embedding, trading a few percent of correctness for orders of magnitude less latency and compute. Every design decision is a point on one surface: recall vs latency vs memory, with filtering, freshness and cost bolted on. The hard production problems are rarely the algorithm. They are filtered queries that silently return 3 results instead of 10, an index that no longer fits in RAM at 40 M vectors, and a model upgrade that makes every stored embedding incompatible with every new query.
That is why "we'll put the embeddings in a vector database" is a sentence interviewers press on. The L5 candidate draws a box labelled "vector DB" next to the LLM. The L6 candidate says "documents are chunked to ~500 tokens and embedded at 1,024 dimensions; vectors live in pgvector with an HNSW index (m 16, ef_construction 64), tenant filtering uses iterative scans so filtered queries still return k results, retrieval is hybrid with BM25 fused by reciprocal rank, and I'll measure recall@10 against exact search on a 1,000-query golden set before and after every index change." The L7 candidate asks what happens to 300 M stored vectors when the embedding model is replaced next year, and who pays for re-embedding the corpus.
The L5 → L6 gap is not knowing what HNSW stands for. It is knowing that a vector index is approximate by design, so recall must be measured, not assumed, and that the embedding model, not the database, is the schema.
The L5 → L6 → L7 Contrast#
| Behavior | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| First move | "Store embeddings in Pinecone / a vector DB" | "How many vectors, how many dimensions, what filters, what freshness, what recall target? At under ~10 M vectors with relational filters, pgvector next to the data is usually enough." | "Is retrieval a shared platform or per-team? Which embedding models do we standardise on, and what does a model migration cost the org?" |
| Correctness | "It finds similar items" | Recall@k measured against brute force on a golden set; ef_search / nprobe tuned to a target (e.g. 0.95) at a latency budget | Retrieval quality as an SLO with owners; offline eval gates every model or index change |
| Filtering | "Add a WHERE clause" | Knows post-filtering starves results; chooses pre-filter, filtered traversal, iterative scans or partitions by filter selectivity | Designs tenancy so the most common filter is a partition boundary, not a predicate |
| Scale | "It scales horizontally" | Sizes RAM: vectors × dims × bytes plus graph overhead; quantization (int8, PQ, binary) when it doesn't fit | Prices the memory tier vs disk-based indexes vs object-storage vector stores across 3 scales |
| Change | "Re-index when needed" | Model change = new index: dual-write, backfill, shadow-compare recall, cut over, keep rollback | Treats the embedding model as a versioned contract with a deprecation policy and a re-embedding budget |
| Ownership | "The ML team owns it" | Ingest pipeline owner, index owner and eval owner named; freshness SLO on embedding lag | Platform owns engines and SLOs; product teams own chunking, filters and relevance |
Why "Correctness" separates levels
Approximate nearest neighbour (ANN) search does not return the true top-k. It returns a candidate set whose overlap with the exact top-k, recall@k, depends on index parameters. HNSW with a small ef_search may return 85% of the true neighbours; raising it to 200 may reach 99% at 3–5× the latency. Nothing in the query response tells you which you got. The Senior answer trusts the index. The Staff answer builds a golden set (say 1,000 representative queries), computes exact neighbours with a brute-force scan offline, and tracks recall@10 alongside p99 latency on every parameter or data change. The Principal answer ties that to end-task quality (answer correctness, click-through) because recall@10 of 0.99 on a bad embedding model is still bad retrieval.
Why "Filtering" separates levels
The naive plan is "find the 10 nearest vectors, then apply WHERE tenant_id = 42." If tenant 42 owns 0.1% of the corpus, the 10 nearest neighbours almost never include its rows, and the query returns 0–2 results with no error. Over-fetching (top 1,000, then filter) helps until the filter is selective enough to need top 100,000. The Staff answer picks by selectivity: a separate index or partition per large tenant, filtered graph traversal or iterative scans for medium selectivity, and exact search over a pre-filtered set when the filter leaves only a few thousand rows. pgvector added iterative index scans in 0.8.0 for exactly this case (pgvector).
The 60-Second Pitch#
"I'd embed documents in ~500-token chunks at 1,024 dimensions and store them in Postgres with pgvector, next to the metadata and permissions they need to be filtered by. At 8 M chunks that's ~33 GB of float32 vectors; stored as halfvec it's ~16 GB, and the HNSW index, which holds its own copy of the vectors plus ~1–2 GB of graph links, is ~18 GB, which fits in RAM on one memory-optimised instance with a replica. HNSW with m 16 and ef_construction 64; ef_search tuned to recall@10 of 0.95 at a p99 under 30 ms. Tenant filters use iterative scans, and our five largest tenants get their own partitions. Retrieval is hybrid: BM25 and vector results fused with reciprocal rank fusion, then a cross-encoder re-ranks the top 50. Every embedding carries a model version; a model change builds a new index in parallel, shadow-compares recall on a golden set and cuts over behind a flag. If we pass ~50 M vectors or need multi-node scale-out, I'd move the index to a dedicated engine and keep Postgres as the source of truth."
The Three Intents#
| Intent | Constraint | Strategy | Failure Mode | Correctness Bar |
|---|---|---|---|---|
| RAG / semantic document search | Per-tenant permissions, freshness in minutes, millions of chunks | pgvector or a dedicated engine; hybrid with BM25; re-ranker; ACL filters | Filtered queries return too few results; stale or leaked chunks | Never return a chunk the user can't read; recall@10 ≥ 0.9 on golden set |
| Recommendations / similarity at scale | 100 M–10 B item vectors, high QPS, loose filters | Dedicated ANN engine or library (HNSW, IVF-PQ), sharded, quantized | RAM cost explodes; index rebuilds take hours | Recall traded for throughput; measured by online metrics |
| Deduplication / clustering / analytics | Batch, offline, whole-corpus | Faiss-style library jobs, GPU, IVF-PQ; no online database | Treating a batch job as an online service | Throughput per dollar; near-dup precision |
🎯 Staff Move: "I'll design for RAG over tenant documents. The permission filter is the hardest part, so the store has to filter correctly, not just search fast. That pushes me toward keeping vectors next to the metadata, and I'll revisit a dedicated engine only when the vector count or QPS forces it."
The Staff Positions#
| Position | Rationale |
|---|---|
| Recall is measured, never assumed | ANN is approximate by design; a golden set with exact neighbours is the only way to know where you are on the dial. |
| The embedding model is the schema | Vectors from different models live in different spaces; a model change is a full re-embed and re-index. |
| Start with pgvector unless numbers say otherwise | Below ~10–50 M vectors, transactional filters, permissions and joins beat a second system to keep in sync. |
| Filters decide the index design | The most selective common filter should be a partition or separate index, not a post-filter. |
| Hybrid beats pure vector for user-facing search | Exact terms, IDs, names and rare words are where embeddings are weakest and BM25 is strongest. |
| Size RAM first | Vectors × dimensions × bytes, plus graph overhead, decides instance type, shard count and whether you need quantization. |
| The source of truth is not the vector index | Indexes are derived and rebuildable; the documents and metadata live in the primary database or object storage. |
Architecture & Internals#
Five internals change design decisions: embeddings and distance, the ANN index family, quantization, filtering, and the write path.
Embeddings and Distance#
An embedding model maps text, images or users to a fixed-length vector (commonly 384–3,072 dimensions). Similarity is a distance: cosine, inner product (equal to cosine on normalised vectors) or L2. Use the metric the model was trained for; normalise once at write time and use inner product for speed.
Storage per vector = dimensions x bytes per component
1,536 dims x 4 bytes (float32) = 6,144 bytes (~6 KB)
1,536 dims x 2 bytes (float16) = 3,072 bytes
1,536 dims x 1 byte (int8) = 1,536 bytes
1,536 dims x 1 bit (binary) = 192 bytes
PQ with 96 sub-quantizers x 1 B = 96 bytes (~64x smaller than float32)
100 M vectors x 6 KB = ~614 GB raw float32, before any index overhead
The ANN Index Family#
| Index | How it works | Strengths | Weaknesses | Typical knobs |
|---|---|---|---|---|
| Flat (exact) | Scan every vector | Perfect recall; no build | O(N) per query; ~1 M vectors is the practical limit for online use | None |
| HNSW | Multi-layer proximity graph; greedy search from a top-layer entry point | Best recall/latency trade in RAM; incremental inserts | Memory-hungry; slow builds; deletes leave tombstones | M (links per node), ef_construction, ef_search |
| IVF (inverted file) | k-means clusters ("lists"); search the nprobe closest lists | Smaller, faster to build; pairs with PQ | Needs training data; recall degrades as data drifts from centroids | lists, nprobe |
| IVF-PQ | IVF plus product-quantized codes | Billion-scale in modest RAM | Lower recall; usually needs re-ranking with full vectors | lists, nprobe, PQ code size |
| Disk-based graph (DiskANN-style) | Graph on SSD, compressed vectors in RAM | 5–10× cheaper per vector than all-RAM | Higher latency (SSD reads per hop); tuning heavy | Beam width, cache size |
HNSW was introduced by Malkov and Yashunin (arXiv 1603.09320); it is the default in most engines today.
Why it matters in design: HNSW query cost grows roughly logarithmically with N, but memory grows linearly with N × (vector size + links), and a graph that spills out of RAM onto disk falls off a latency cliff. That is why the sizing math, not the QPS, usually picks the hardware.
The Recall Dial#
HNSW, illustrative behaviour on 10 M x 768-dim vectors, single node, in RAM:
ef_search recall@10 p50 latency p99 latency
20 ~0.85 ~1 ms ~3 ms
40 ~0.92 ~2 ms ~5 ms <- pgvector default is 40
100 ~0.97 ~4 ms ~10 ms
400 ~0.995 ~12 ms ~30 ms
(shape is typical; measure on your data, it varies with dimensionality and distribution)
Raising M or ef_construction improves the graph (better recall at the same ef_search) at the cost of build time and memory; raising ef_search improves recall at query time at the cost of latency. pgvector's defaults are m 16, ef_construction 64 and hnsw.ef_search 40; IVFFlat defaults to probes 1, with suggested lists of rows/1,000 up to 1 M rows and √rows beyond (pgvector).
🎯 Staff Insight: "Recall and latency are one dial, and memory is the price of the dial's range. I pick a recall target with the product owner, say 0.95 at k=10, then find the cheapest ef_search that meets it under the p99 budget, and re-check both every time the corpus doubles."
Quantization: Paying for RAM With Recall#
| Technique | Size vs float32 | Recall impact | Typical use |
|---|---|---|---|
| float16 / halfvec | 1/2 | Negligible for most models | Default storage choice when supported |
| Scalar int8 | 1/4 | Small; re-rank fixes most of it | Large HNSW indexes |
| Product quantization (PQ) | 1/16–1/64 | Noticeable; re-rank required | Billion-scale IVF-PQ |
| Binary quantization | 1/32 | Large alone; good as a first-pass filter with re-rank | Very high-dimensional models built for it |
The standard pattern is two-stage: search the compressed index for the top 100–500 candidates, then re-score them with full-precision vectors fetched from disk or the primary store. Meta's Faiss work showed PQ codes letting 1 B vectors be indexed in under 30 GB of RAM (Faiss).
Filtering Strategies#
In pgvector 0.8.0+, hnsw.iterative_scan = relaxed_order keeps walking the graph until enough rows pass the filter, bounded by hnsw.max_scan_tuples (20,000 by default). Partial indexes work for a handful of low-cardinality filter values; partitioning works when there are many. Dedicated engines implement filter-aware graph traversal natively; check how each handles very selective filters before trusting it.
The Write Path#
| Concern | What happens | Design consequence |
|---|---|---|
| Inserts into HNSW | Each insert searches the graph to pick neighbours, ~ms each | Bulk-load then build the index; streaming inserts cap ingest at thousands/s per node |
| Updates | Usually delete + insert | Frequent re-embedding of the same rows churns the graph |
| Deletes | Tombstones; space and recall recovered only on rebuild or vacuum | Plan periodic rebuilds for high-churn corpora |
| Index build | CPU- and memory-heavy; fastest when the graph fits in maintenance_work_mem in Postgres | Build on a replica or new table, then swap |
| Freshness | Embedding is an external call (10–100 ms per batch to a model) | Async pipeline with an embedding-lag SLO, e.g. p99 < 5 min |
Data Modeling / Core Usage — "The Entire Game"#
In Kafka the game is the partition key. In a vector store it is the retrieval contract: what a vector represents (the chunk), which model produced it (the version), and which filters must hold when it's returned (permissions, tenant, freshness). The index type is a tuning decision; these three are design decisions.
Step 1: Decide What a Vector Represents#
| Unit | Typical size | Good for | Trap |
|---|---|---|---|
| Whole document | 1 vector per doc | Short items: products, titles, profiles | Long documents average into mush |
| Chunk (fixed window) | 200–800 tokens, 10–20% overlap | RAG over manuals, tickets, wikis | Chunks cut sentences and tables; chunk count multiplies storage |
| Structural chunk | Section, paragraph, function | Docs with headings, code | Uneven sizes; very long sections still need splitting |
| Multiple vectors per item | Title + body + image | Multi-modal and multi-field retrieval | 3× storage and write cost |
Chunk math: 2 M documents x ~4,000 tokens / 500-token chunks = ~16 M chunks
16 M x 1,024 dims x 2 bytes (halfvec) = ~33 GB of vectors
Change chunking to 250 tokens -> 32 M chunks -> double everything
Step 2: Put the Model Version on Every Row#
create extension if not exists vector;
create table chunks (
chunk_id bigserial primary key,
doc_id bigint not null references documents(doc_id),
tenant_id bigint not null,
acl_groups bigint[] not null, -- who may see it
content text not null,
embedding halfvec(1024) not null,
model_version text not null, -- e.g. 'embed-v3-1024'
updated_at timestamptz not null default now()
) partition by list (tenant_id); -- big tenants get their own partitions
create index on chunks_default using hnsw (embedding halfvec_ip_ops)
with (m = 16, ef_construction = 64);
create index on chunks_default (tenant_id);
-- Query: tenant-scoped, permission-filtered, model-matched
set hnsw.ef_search = 100;
set hnsw.iterative_scan = relaxed_order; -- keep scanning until k rows pass filters
select chunk_id, doc_id, content
from chunks
where tenant_id = $1
and acl_groups && $2 -- user's groups
and model_version = 'embed-v3-1024'
order by embedding <#> $3 -- negative inner product
limit 10;
Keeping the vector next to tenant_id and acl_groups means the permission check happens in the same query, inside one transaction boundary. With a separate vector store, the ACL has to be copied into its metadata and kept in sync, and a lagging sync is a data leak, not a stale result.
🎯 Staff Move: "Permissions are filtered inside the retrieval query, not after it. If I move to a dedicated engine, ACLs replicate into it through the same change stream as the vectors, and I'll alert on replication lag because a stale ACL is a security bug."
Step 3: Design for Hybrid Retrieval#
Embeddings are strong on paraphrase and weak on exact tokens: product SKUs, error codes, names, rare terms. BM25 is the reverse. Reciprocal rank fusion needs no score calibration between the two lists, which is why it's the default; the constant 60 comes from the original RRF paper. Lexical search mechanics are in Elasticsearch and Elasticsearch vs Postgres Full-Text; the broader ranking stack is in Search Engine.
Step 4: Build the Golden Set Before the Index#
golden_set: 1,000 real queries (sampled from logs, stratified by tenant size and query type)
for each query: exact top-10 via brute-force scan (offline, once per corpus snapshot)
metrics per index config:
recall@10 = |ANN top-10 ∩ exact top-10| / 10, averaged
filtered_fill = fraction of queries that returned the full k after filters
p50 / p99 latency at production QPS
gate: no index, parameter or model change ships if recall@10 drops > 0.02
or filtered_fill < 0.99
filtered_fill is the metric teams forget: it catches the post-filter starvation bug that recall on unfiltered queries hides.
The Tunable Tradeoff — Recall × Latency × Memory#
Every vector-search decision moves along three linked axes. You can have any two cheaply; the third costs money.
| Setting | Cheap / fast end | Accurate end | Who pays at the cheap end |
|---|---|---|---|
ef_search / nprobe | Low (20, 1) | High (200+, 32+) | Users, through missed results |
Graph degree M | 8–12 | 32–64 | Recall ceiling; product relevance |
| Vector precision | Binary or PQ | float32 | Relevance, unless re-ranked |
| Storage medium | Disk-based or object-storage index | All in RAM | Latency: SSD hops add ms per query |
| Re-ranking | None | Full-precision + cross-encoder | Relevance at the top of the list |
| Freshness | Nightly batch rebuild | Streaming inserts | Users who search for what they just wrote |
RAM sizing (HNSW, rule of thumb):
bytes ≈ N x (d x bytes_per_dim + M x 2 x ~8 bytes for links) x 1.2 overhead
20 M x 768 dims, float32, M=16:
vectors 20 M x 3,072 B = 61 GB
links 20 M x 256 B = 5 GB
total ~66 GB x 1.2 = ~80 GB -> a 128 GB node, or 2 shards on 64 GB nodes
same corpus, int8 + re-rank from disk: ~20 GB in RAM
Check your own numbers with the Back-of-Envelope Calculator and, when the query sits inside a user request, against the Latency Budget.
🎯 Staff Move: "I'll quantize to int8 and re-rank the top 200 with full-precision vectors. That cuts RAM about 4× and costs us roughly a point of recall, which the re-rank wins back. I'd rather spend money on a re-ranker than on RAM for float32 vectors nobody reads at full precision."
Who Pays for Each Choice#
| Choice | What Works | What Breaks | Who Pays |
|---|---|---|---|
Default ef_search, never measured | Fast, simple | Recall unknown, often 0.85–0.9 | Users and the LLM answering from wrong context |
| Post-filtering | Easy to implement | Selective filters return too few results, silently | Small tenants, who get worse search |
| Separate vector store | Scale-out, rich ANN features | ACL and metadata sync, two systems to operate | Security on sync lag; on-call for two systems |
| All-in-RAM float32 | Best latency and recall | Memory bill grows linearly | Finance |
| Aggressive quantization without re-rank | Cheapest RAM | Relevance drops noticeably | Product quality |
| Streaming inserts into HNSW | Fresh results | Ingest throughput limits, graph fragmentation | Ingest pipeline on-call |
Anti-Patterns — What Kills Vector Database Deployments#
1. Never Measuring Recall#
The index is tuned once with defaults and never evaluated. Corpus grows 10×, recall slides from 0.95 to 0.8, and the team blames the LLM. Fix: a golden set with exact neighbours, recall@k and filtered fill tracked on every change.
2. Post-Filtering Selective Queries#
"ANN top 10, then WHERE tenant_id" returns empty pages for small tenants. Fix: partitions or per-tenant indexes for large tenants, iterative scans or filtered traversal for medium selectivity, exact search for tiny filtered sets. See Multi-Tenancy.
3. Mixing Embedding Models in One Index#
New documents embedded with v2, old ones still v1, queries embedded with v2. Distances across models are meaningless, so old documents silently stop being found. Fix: model_version on every row, one index per model, a migration plan for every model change.
4. The Vector Store as Source of Truth#
Documents exist only as chunks in the vector index. A model change, chunking change or index corruption means re-ingesting from wherever the data originally came from, if it still exists. Fix: canonical documents in a database or object storage; the index is derived and rebuildable.
5. Pure Vector Search for Exact-Match Queries#
Users search "ERR_4021" or an order ID and get semantically similar but wrong results. Fix: hybrid retrieval with BM25, or route ID-shaped queries to an exact lookup.
6. Unbounded Chunk Multiplication#
Halving chunk size, adding overlap and adding a title vector each multiply storage; together they turn 10 M vectors into 60 M. Fix: treat chunking as a costed schema decision with an eval behind it.
7. Building the Index in Place During Peak#
HNSW builds are CPU- and memory-heavy and can starve queries on the same node for hours. Fix: build on a replica or a new table, then swap; schedule rebuilds off-peak.
8. A Dedicated Engine for 200,000 Vectors#
A second database, sync pipeline and on-call rotation for a corpus a flat scan or pgvector handles in milliseconds. Fix: start next to the data; move when numbers force it. See Buy or Build.
The Technology Landscape — Head-to-Head Comparison#
| Dimension | pgvector (Postgres) | Dedicated engines (Milvus, Qdrant, Weaviate, Vespa) | Managed vector services | Search engines with kNN (Elasticsearch, OpenSearch) | Object-storage vector indexes (e.g. S3 Vectors) | ANN libraries (Faiss, hnswlib, ScaNN) |
|---|---|---|---|---|---|---|
| Model | Index type inside a relational DB | Purpose-built vector DB, sharded | Hosted API, serverless or pods | Inverted index plus HNSW fields | Vector indexes in object storage, API queries | In-process library, you build the service |
| Sweet spot | Up to ~10–50 M vectors with relational filters | 50 M–billions, high QPS | Teams without ops capacity | Hybrid lexical + vector in one engine | Large, infrequently queried corpora at low cost | Batch jobs, custom serving, research |
| Filtering | Full SQL; iterative scans for ANN + filter | Native filter-aware traversal | Metadata filters | Full query DSL | Metadata filters (limited size) | Yours to build |
| Transactions with source data | Yes | No; sync pipeline | No | No | No | No |
| Ops burden | Low if you already run Postgres | Medium–high | Low | Medium | Low | High (you own serving) |
| Latency profile | ms in RAM | ms in RAM; tiered options | ms, network hop | ms–tens of ms | Sub-second, not single-digit ms | µs–ms in process |
S3 Vectors, for example, documents up to 2 billion vectors per index and 1–4,096 dimensions, with write rates capped per index (S3 Vectors limits), which makes it a fit for cheap, large, lower-QPS retrieval rather than a hot path.
🎯 Staff Insight: "The real comparison isn't pgvector versus a vector database. It's one system with transactional filters versus two systems with a sync pipeline. I pay for the second system only when vector count, QPS or ANN features force me to, and I write down which number triggered it."
Patterns#
Pattern 1: Vectors Next to the Data (pgvector)#
Embeddings live in the same Postgres as documents and ACLs; HNSW index; partitions for large tenants; read replicas for query scale. The default for RAG under a few tens of millions of chunks. See PostgreSQL.
Pattern 2: Source of Truth + Derived Vector Index#
The primary database stays the truth; a change stream (see Transactional Outbox and Kafka) drives embedding and indexing; deletes and ACL changes flow through the same stream with priority. The query service re-checks permissions against the primary for the final page if the domain is sensitive.
Pattern 3: Two-Stage Retrieval (Compressed Recall, Precise Rank)#
Quantized ANN for the top 200–1,000 candidates, full-precision re-score, then a cross-encoder for the top 20–50. Each stage is 10× smaller and 10× more expensive per item than the last.
Pattern 4: Blue/Green Index for Model Migration#
New model → new embeddings column or index built in the background → shadow queries compare recall and task metrics → cut over by flag per tenant → keep the old index for a rollback window → drop. Same discipline as Online Migrations.
Pattern 5: Per-Tenant Partitions for Isolation#
Large tenants get dedicated partitions or collections; the long tail shares one index with a tenant filter. Removes the post-filter problem for the tenants who matter most and makes tenant deletion a partition drop.
Pattern 6: Semantic Cache#
Embed the incoming query; if a previous query's embedding is within a tight distance threshold, return its cached answer. Saves LLM calls on repetitive traffic; needs a conservative threshold and tenant scoping, or users see each other's answers. See Caching Fundamentals.
Scaling#
The Numbers#
| Resource | Documented limit or typical value | Design note |
|---|---|---|
| pgvector indexable dimensions | vector 2,000; halfvec 4,000; bit 64,000; sparsevec 1,000 non-zero | 3,072-dim models need halfvec to be indexed |
| pgvector HNSW defaults | m 16, ef_construction 64, ef_search 40 | Tune ef_search per query; it's a session setting |
| pgvector iterative scan bound | hnsw.max_scan_tuples 20,000 | Very selective filters can still under-fill; use partitions |
| Flat scan practical limit | ~1 M vectors per query online | Fine for small tenants and pre-filtered sets |
| HNSW insert cost | ~ms per insert per node (typical) | Bulk-load, then build |
| float32 vector at 1,536 dims | ~6 KB | 100 M vectors ≈ 614 GB before index overhead |
| Object-storage vector index (S3 Vectors) | Up to 2 B vectors per index, 1–4,096 dims, ~1,000 write requests/s per index | Cheap at rest; not a single-digit-ms path |
Sharding and Tail Latency#
ANN indexes shard by random or hash assignment (each shard holds a slice of all vectors, and every query fans out to every shard) or by a partition key such as tenant (queries go to one shard). Fan-out multiplies tail latency: the query is as slow as its slowest shard.
Fan-out math: 16 shards, each with p99 = 10 ms
P(query avoids every shard's slowest 1%) = 0.99^16 ≈ 0.85
-> ~15% of queries see at least one shard at its p99
-> query p99 is well above 10 ms; mitigate with hedged requests,
fewer bigger shards, or tenant-routed shards
Scaling Moves in Order#
- Measure first: golden-set recall and filtered fill at current parameters, so every later step has a baseline.
- Shrink the vectors: halfvec, then int8 or PQ with re-ranking. Often a 2–4× RAM cut for under a point of recall.
- Reduce the corpus: deduplicate near-identical chunks, drop boilerplate, revisit chunk overlap.
- Partition by the dominant filter (tenant, language, region) so most queries hit one smaller index.
- Add read replicas for QPS while the index still fits one node.
- Shard or move to a dedicated engine once a single node can't hold the index in RAM at your recall target, or rebuilds take longer than your freshness SLO allows.
- Tier cold data to disk-based or object-storage indexes for large, rarely queried corpora.
🎯 Staff Move: "I scale a vector index by making it smaller before making it bigger: halfvec, quantization with re-rank, dedupe, then partitions by tenant. Sharding comes last because fan-out costs me tail latency on every query."
Failure Modes & Recovery#
1. Silent Recall Decay#
- Symptom: Answer quality complaints rise; latency and error dashboards are green.
- Root cause: Corpus grew 5–10× at fixed
ef_search; IVF centroids trained on old data; tombstones from deletes. - Detection: Nightly golden-set
recall_at_10andfiltered_fill; online proxy metrics (click-through, answer acceptance). - Fix: Raise
ef_search/nprobeto restore the target; retrain IVF or rebuild HNSW. - Prevention: Recall gate in CI for index changes; scheduled rebuilds tied to corpus growth.
2. Filtered Queries Return Too Few Results#
- Symptom: Small tenants report "search finds nothing"; large tenants are fine.
- Root cause: Post-filtering or a scan bound reached before k rows pass a selective filter.
- Detection:
filtered_fillby tenant size bucket; count of queries returning fewer than k. - Fix: Enable iterative scans, raise the scan bound, or route small tenants to exact search.
- Prevention: Partition design by filter selectivity; per-tenant eval slices.
3. Index Outgrows RAM#
- Symptom: p99 jumps from ~10 ms to hundreds of ms as the index starts hitting disk; CPU waits on I/O.
- Root cause: Corpus growth or a chunking change pushed the index past memory.
- Detection: Index size vs available memory; buffer cache hit ratio; page-fault and disk-read rates on the query nodes.
- Fix: Quantize, move to a larger node, or shard; temporarily shed low-priority query traffic.
- Prevention: Capacity alert at 70% of RAM; sizing review for every chunking or model change.
4. Model Mismatch After a Partial Migration#
- Symptom: Relevance collapses for older documents; newer ones are fine.
- Root cause: Queries embedded with the new model hit vectors from the old one, or a backfill stopped halfway.
- Detection: Count of rows by
model_version; recall per model slice; query-time assertion that query and index versions match. - Fix: Route queries to the index matching the stored version until backfill completes; resume backfill.
- Prevention: One index per model version; blue/green cutover; version check enforced in the query service.
5. Ingest Lag and Stale or Leaked Results#
- Symptom: New documents not searchable for hours; worse, deleted or permission-revoked documents still returned.
- Root cause: Embedding provider throttling, a stuck consumer, or ACL updates queued behind bulk re-embeds.
- Detection: Embedding lag p99; queue age; count of results whose ACL check fails at the primary.
- Fix: Prioritise deletes and ACL changes over inserts; scale workers; final-page permission re-check against the primary.
- Prevention: Separate priority lanes; embedding-lag SLO with alerting; deletes applied as filters immediately even before index compaction.
Operational Reality Matrix#
| Failure | Detection Signal | Blast Radius | Mitigation | Owner |
|---|---|---|---|---|
| Recall decay | Golden-set recall, online quality proxies | All queries, gradually | Retune, rebuild | Search/retrieval team |
| Filter starvation | Filtered fill by tenant size | Small tenants | Iterative scans, partitions | Retrieval team; product for tenancy model |
| Index exceeds RAM | Memory headroom, cache hit ratio | All queries on the node | Quantize, scale, shard | Platform/DB team |
| Model mismatch | Rows by model version | Older content | Version routing, finish backfill | ML platform + retrieval team |
| Ingest lag / ACL lag | Embedding lag, ACL re-check failures | Freshness; security if ACLs lag | Priority lanes, re-check at primary | Ingest owner; security reviews |
When to Use vs. Alternatives#
| Need | Pick | Why |
|---|---|---|
| RAG over < ~10–50 M chunks with permissions | pgvector in the existing Postgres | One system, transactional filters, SQL joins |
| Hundreds of millions to billions of vectors, high QPS | Dedicated vector engine | Sharding, quantization, filter-aware traversal |
| Lexical + vector in one query with rich text features | Elasticsearch / OpenSearch kNN | Hybrid and BM25 in one engine |
| Large, rarely queried corpus where cost dominates | Object-storage vector index | Pay mostly for storage, not RAM |
| Offline dedupe, clustering, candidate generation | Faiss / ScaNN in batch jobs | No database needed; GPUs help |
| Exact-match, IDs, structured filters only | No vector index: B-tree, inverted index | Embeddings add cost without benefit |
When NOT to Use a Vector Database#
- The corpus is small. Under ~100K–1 M vectors, an exact scan in memory or in Postgres is fast enough and has perfect recall.
- The queries are exact. IDs, SKUs, codes and structured filters belong in B-tree or inverted indexes; see Database Indexes.
- There's no evaluation. Without a golden set and a quality metric, you can't tell a good index from a bad one, and adding one adds risk without measurable value.
- Permissions are complex and the engine can't filter them. Copying fine-grained ACLs into a separate store with lag is a leak waiting to happen.
- As the system of record. The index is derived; keep the documents elsewhere.
Operational Concerns#
What the On-Call Actually Does#
- Watches query p99 and memory headroom on index nodes; the RAM cliff is the most common performance incident.
- Watches ingest: embedding lag, provider throttling, queue age, and the ACL/delete lane separately from bulk inserts.
- Reviews nightly recall and filtered-fill reports and treats a drop of more than 0.02 as an incident ticket.
- Runs rebuilds on replicas or shadow indexes, then swaps; never rebuilds in place at peak.
- Drives model migrations: backfill progress by
model_version, shadow comparison, per-tenant cutover, rollback window.
Key Metrics & Alerts#
| Metric | Healthy | Alert |
|---|---|---|
| Query latency p99 | Within budget (e.g. < 30 ms) | 2× budget for 10 min |
| Index size / RAM | < 70% | > 85% (page before the cliff) |
| Golden-set recall@10 | ≥ target (e.g. 0.95) | Drop > 0.02 day over day |
| Filtered fill (queries returning k) | ≥ 0.99 | < 0.97 in any tenant bucket |
| Embedding lag p99 | < 5 min | > 30 min |
| ACL/delete propagation lag | < 1 min | > 5 min (security ticket) |
Rows not on current model_version | 0 outside migrations | Any growth after cutover |
Interview Application — Staff-Level Plays#
Which Case Studies Use Vector Search#
| Case Study | How Vectors Are Used | Key Pattern |
|---|---|---|
| Search Engine | Semantic retrieval stage beside BM25 | Hybrid retrieval, fusion, re-ranking |
| News Feed | Candidate generation from user and item embeddings | ANN for recall, ranker for precision |
| Dating & Matching | Similar-profile candidates under hard filters | Filter selectivity drives index design |
| Typeahead | Semantic suggestions for long-tail queries | Prefix index first; vectors only as a fallback |
| Web Crawler | Near-duplicate page detection | Batch ANN or hashing, not an online DB |
Every System Design Question Has a Vector Moment#
- Support chatbot: "Tickets and docs are chunked and embedded into pgvector with tenant and ACL columns; retrieval is hybrid and filtered in one query; recall@10 on a golden set gates every change."
- E-commerce search: "SKUs and brand names go through BM25; descriptive queries get vector results; RRF merges them; a learned ranker orders the top 100."
- Recommendations: "Item embeddings in a sharded ANN index produce 500 candidates in ~10 ms; the ranker does the expensive work on those 500."
- Content moderation: "Known-bad images are embedded; new uploads are checked for near-duplicates above a tuned similarity threshold, with precision measured on labelled pairs."
What Interviewers Probe#
| After You Say... | They Will Ask... | What They're Evaluating |
|---|---|---|
| "Store embeddings in a vector DB" | "Why not Postgres? How many vectors?" | Sizing and avoiding a second system without cause |
| "HNSW index" | "How do you know the results are right?" | Recall measurement against exact search |
| "Filter by tenant" | "What happens for a tenant with 0.1% of the data?" | Post-filter starvation and partition design |
| "We'll upgrade the embedding model" | "What happens to the 200 M stored vectors?" | Re-embedding, blue/green index, version tagging |
| "It scales horizontally" | "What does fan-out do to p99?" | Tail latency across shards |
| "Semantic search for everything" | "What about an exact order ID?" | Hybrid retrieval |
Common Interview Mistakes#
| What Candidates Say | What Interviewers Hear | What Staff Engineers Say |
|---|---|---|
| "The vector DB returns the nearest neighbours" | Doesn't know ANN is approximate | "It returns approximate neighbours; I measure recall@10 against exact search." |
| "Filter the results by user" | Post-filter starvation | "Filtering happens inside the search, with partitions for big tenants." |
| "We'll use a dedicated vector DB" (at 1 M vectors) | Adds a system for no reason | "pgvector next to the data until vector count or QPS forces a move." |
| "Re-index when we change models" | Hasn't priced it | "Blue/green: new index, backfill, shadow-compare, cut over per tenant." |
| "More dimensions are better" | Ignores memory and cost | "Dimensions are RAM; I'd test a smaller model or truncation against the golden set." |
L5 vs L6 vs L7 Responses#
| Scenario | L5 Answer | L6 / Staff Answer | L7 / Principal Answer |
|---|---|---|---|
| "Add semantic search to our docs" | Embed everything into a managed vector DB | pgvector with tenant/ACL filters, hybrid BM25, golden-set eval, embedding-lag SLO | Shared retrieval platform with standard models, eval harness and cost per query |
| "Search quality dropped" | Try a bigger model | Check recall decay, filtered fill, model-version mix before touching the model | Quality SLO with owners; eval gate on every model and index change |
| "Index doesn't fit in RAM" | Bigger instance | halfvec / int8 with re-rank, dedupe, partitions, then shard | Tiering policy: hot RAM index, cold disk or object-storage index, priced per tenant |
| "Switch embedding providers" | Re-embed overnight | Blue/green index, shadow traffic, per-tenant cutover, rollback window | Model deprecation policy, re-embedding budget, contract terms to avoid lock-in |
The Staff Vector Database Checklist#
- Size it: "16 M chunks × 1,024 dims × 2 bytes ≈ 33 GB; plus graph links; one node with a replica."
- Define the unit: "500-token chunks with 15% overlap; model version on every row."
- Filters first: "Tenant and ACL filters inside the query; partitions for the largest tenants."
- Measure: "Golden set of 1,000 queries; recall@10 ≥ 0.95 and filtered fill ≥ 0.99 at p99 < 30 ms."
- Hybrid: "BM25 plus vectors, RRF, cross-encoder on the top 50."
- Plan the migration: "A model change is a blue/green index with shadow comparison and per-tenant cutover."
🎯 Staff Insight: Don't use a vector index as the system of record, as a substitute for exact lookups, or without an evaluation set. The strongest signal is saying how you know the approximate answers are good enough, and what happens to every stored vector when the model changes.
Evaluation Rubric#
| Dimension | Senior (L5) | Staff (L6) | Principal (L7) |
|---|---|---|---|
| Correctness | Trusts the index | Recall@k and filtered fill on a golden set | Task-level quality SLO and eval gates org-wide |
| Sizing | "It scales" | RAM math, quantization, partitioning, shard trade-offs | Tiered storage and cost per query across products |
| Filtering | WHERE clause | Selectivity-driven strategy; ACLs inside the query | Tenancy model designed around retrieval |
| Change | "Re-index" | Blue/green model migration with shadow compare | Model lifecycle policy and re-embedding budget |
| Choice of engine | Dedicated DB by default | pgvector until a stated trigger | Platform standard, exceptions process, build vs buy |
Beyond Staff: The Principal View#
Why L7 Sees This Problem Differently#
At Staff level a vector database is an index you size and tune. At Principal level retrieval is shared infrastructure every AI feature depends on, and the embedding model is a company-wide schema. When five teams pick five models, five engines and five chunking schemes, the org pays five times for storage, five times for re-embedding, and has no common way to say whether retrieval got better or worse. The L7 question is "which models do we standardise on, who owns the eval harness, and what's the plan when the model we chose is deprecated?"
🧭 Principal Move: "Before we choose an engine, I want two standards: a small set of approved embedding models with a version and deprecation policy, and one evaluation harness every team uses. Engines are replaceable; ungoverned embeddings in a dozen places are not."
The Org-Level Fault Line#
Shared retrieval platform vs per-team vector stores.
| Option | What Works | What Breaks | Who Pays |
|---|---|---|---|
| Every team picks its own store and model | Fast experiments | Duplicate corpora, inconsistent ACL handling, N migrations per model change | Security and finance |
| One central vector platform for everything | Shared ops, consistent evals | Becomes a bottleneck; one-size indexing fits nobody | Product velocity |
| Paved road: pgvector by default, platform engine above a threshold, shared eval harness | Simple default, scale path, common quality bar | Needs a platform team and clear triggers | Platform headcount (2–4 engineers) |
The Principal default: a paved road. Postgres with pgvector for most teams; a platform-run dedicated engine for corpora above a stated size or QPS; approved models with version tags; one eval harness; ACL propagation as a platform guarantee with a lag SLO.
Cost Model#
Assumptions: memory-optimised nodes ~$8–12 per GB of RAM per month including replicas; embedding at ~$0.02–0.10 per million tokens (varies widely by provider and model); loaded engineer ~$250K/year. Directional only.
| Scale | Vectors | RAM needed (halfvec + graph, 2 replicas) | Infra/month | Re-embed corpus once | Headcount | Total/month |
|---|---|---|---|---|---|---|
| Startup | 2 M × 1,024 | ~10 GB | ~$300–600 (inside existing Postgres) | ~$20–100 | 0.1 FTE | ~$2–3K |
| Growth | 100 M × 1,024 | ~450 GB | ~$4–6K | ~$1–5K | 1 FTE (~$21K) | ~$25–30K |
| Enterprise | 5 B × 768, int8 + re-rank | ~10 TB | ~$90–130K | ~$50–250K per migration | 4–6 FTE (~$100K) | ~$200–250K |
At growth and enterprise scale, people and migrations rival infrastructure. A model change at 5 B vectors is a quarter-long programme, so the model standard is the biggest cost lever, followed by quantization.
The 3-Year Evolution Path#
One-Way Doors vs Two-Way Doors#
| Decision | Reversibility | Cost to Reverse |
|---|---|---|
| Embedding model for a large corpus | One-way-ish | Full re-embed and re-index; weeks at billions of vectors |
| Storing documents only in the vector store | One-way if sources disappear | Re-ingest from origin, if possible |
| Chunking scheme | One-way-ish | Re-chunk, re-embed, re-index |
| Engine choice with a CDC-fed derived index | Two-way | Rebuild index in the new engine from the source of truth |
HNSW parameters, ef_search, quantization | Two-way | Rebuild or session setting |
| Hybrid fusion and re-ranking | Two-way | Code change |
The Standard I'd Write#
RFC-RETR-001: Vector Retrieval Baseline
Scope: Every production feature that retrieves by embedding similarity.
MUST
1. Use an approved embedding model; store model_version on every vector.
2. Keep canonical documents in a primary database or object storage; vector
indexes are derived and rebuildable from them.
3. Enforce tenant and permission filters inside the retrieval query; ACL and
delete propagation lag p99 < 1 minute where a separate store is used.
4. Maintain a golden set (>= 500 queries) and report recall@10 and filtered
fill; changes that drop recall by > 0.02 do not ship.
5. Migrate models blue/green with shadow comparison and a rollback window.
SHOULD
6. Use pgvector below 50 M vectors unless an exception is approved.
7. Combine lexical and vector retrieval for user-facing search.
8. Quantize with full-precision re-ranking above 100 GB of index.
Exceptions: retrieval platform lead approval, reviewed each quarter.
Success metrics: zero cross-tenant retrieval incidents; every index on a supported
model; recall reported for 100% of production indexes; cost per 1K queries per team.
What I'd Tell the VP#
"Every AI feature we ship depends on finding the right documents first, and today each team does that differently, with different models and no shared way to measure whether it works. That costs us in duplicated infrastructure and, more importantly, makes every future model upgrade five separate projects. I'm proposing a small set of approved models, one quality test every team runs, and a default storage approach in the database we already operate, with a dedicated system only for the largest workloads. That keeps costs predictable, makes upgrades a routine instead of a crisis, and closes the risk of one customer's documents appearing in another's answers. The cost is a platform team of two to three engineers."
Principal Interview Signals#
| Signal | What It Sounds Like |
|---|---|
| Model as schema | "The embedding model is a versioned contract; changing it is a migration, not a deploy." |
| Quality as an SLO | "Recall@10 and filtered fill are reported like latency, with an owner." |
| Prices migrations | "At 5 B vectors a model change is a quarter of work; that's why we standardise." |
| Security in retrieval | "ACLs propagate with a lag SLO; a stale ACL is a leak." |
| Defaults with triggers | "pgvector until 50 M vectors or a QPS we can't serve from replicas." |
Staff answers that L7 interviewers find insufficient:
- "We'll benchmark engines and pick the fastest" without who owns the eval harness or the migration path.
- "Use the best embedding model" without its deprecation risk and re-embedding cost.
- "Filter by tenant" without the org's tenancy and ACL propagation guarantees.
How Real Companies Built It#
Meta — Faiss and Billion-Scale Similarity Search#
Meta's research team released Faiss in 2017 as a library for similarity search over billions of vectors. The announcement describes building what it calls the first k-nearest-neighbour graph on 1 billion high-dimensional vectors, being 8.5× faster than the previously reported state of the art, GPU implementations typically 5–10× faster than CPU, and product-quantized indexes that fit 1 billion vectors in under 30 GB of RAM (Engineering at Meta).
Staff insight: Compression is what made billion-scale search practical. When an interviewer pushes the corpus to billions, the answer is quantization plus re-ranking, not more RAM.
Spotify — From Annoy to Voyager#
Spotify open-sourced Annoy in 2013 and used it for a decade behind features like Discover Weekly. In 2023 it released Voyager, an HNSW-based library built on hnswlib, reporting more than 10× the speed of Annoy at the same recall, up to 50% more accuracy at the same speed, up to 4× less memory than Annoy, and production use across Spotify teams since 2022 (Spotify Engineering).
Staff insight: The index algorithm changed under a decade-old product without changing the product. That's the argument for keeping the index derived and swappable behind a stable retrieval interface.
Google — ScaNN and Anisotropic Quantization#
Google Research's ScaNN library introduced anisotropic vector quantization, which penalises quantization error parallel to the original vector more heavily because that error distorts the high inner products that matter for ranking. On the glove-100-angular benchmark, Google reported roughly twice the queries per second of the next-fastest library at a given accuracy (Google Research blog).
Staff insight: Quantization isn't neutral compression; how you compress decides which errors you accept. In an interview, "we'll quantize" is incomplete until you say how you'll measure the recall it costs.
Practice Drill#
Prompt: "Your B2B product launched an AI assistant that answers from each customer's documents. It runs on a managed vector service with 60 M chunks. Large customers love it; small customers say it 'can't find anything.' Last week one customer saw a snippet from a document they had deleted. The ML team wants to switch to a new embedding model next quarter. What do you do?"
Staff Answer
Three problems, in priority order. The deleted-document snippet is a security incident, so it goes first. A deleted document still being retrievable means deletes propagate to the vector store asynchronously and lag, or tombstoned vectors are still returned. Immediate fix: re-check every returned chunk against the primary database (document exists, user has access) before it reaches the model, which closes the leak regardless of index state. Then put deletes and ACL changes on a priority lane in the change stream with a propagation-lag SLO of p99 under a minute, and alert on it. Small customers finding nothing is almost certainly post-filter starvation: a small tenant owns ~0.01% of 60 M chunks, so an unfiltered top-k rarely contains its rows. I'd measure filtered fill by tenant size, then fix it by selectivity: tenants under ~50K chunks get exact search over their own rows (a few ms), mid-size tenants use the engine's filtered traversal with a higher candidate budget, and the largest tenants get their own collections. The model change becomes a blue/green migration: tag every vector with model_version, build the new index from the source documents in the background (60 M chunks is a few days of embedding at provider rate limits), shadow-query both on a 1,000-query golden set per tenant segment comparing recall@10 and answer acceptance, cut over per tenant behind a flag, keep the old index for two weeks. Metrics I'd report weekly: ACL/delete lag, filtered fill by tenant bucket, recall@10, p99 latency and cost per 1K queries.
Why this is L6:
- Treats the stale-delete as a security bug with an immediate, index-independent fix (final permission check at the primary).
- Diagnoses small-tenant failures as filter selectivity, and chooses a strategy per selectivity band.
- Turns the model change into a measured blue/green migration instead of an overnight re-embed.
What L7 adds:
- Asks why ACLs live in two systems at all: for a B2B product with fine-grained permissions, pgvector next to the source data may remove the sync problem entirely.
- Sets an org standard: approved models with version tags, a shared eval harness, and a re-embedding budget per model change.
- Prices the migration (embedding cost, double index for the overlap window, engineering weeks) and uses it to negotiate model and vendor terms.
❌ Common L5 Trap
"Increase top-k to 100 so small customers get results, rebuild the index to purge deleted documents, and re-embed everything with the new model over a weekend."
Why this misses: Over-fetching only helps until a filter is more selective than k/N, so the smallest tenants still starve. A rebuild purges today's deletes but not tomorrow's, leaving the leak open without a propagation SLO or a final permission check. And an in-place weekend re-embed mixes model versions mid-flight with no recall comparison or rollback.
Quick Reference Card#
Core idea: approximate nearest neighbours; recall is a dial, not a given
Storage: dims x bytes; 1,536 x float32 = ~6 KB; 100 M vectors ~614 GB raw
Precision: float16 1/2, int8 1/4, PQ 1/16-1/64, binary 1/32 (re-rank after)
HNSW: graph; knobs M, ef_construction, ef_search; RAM-hungry; tombstone deletes
pgvector defaults: m 16, ef_construction 64, ef_search 40; IVFFlat probes 1
pgvector limits: index vector <= 2,000 dims, halfvec <= 4,000, bit <= 64,000
Iterative scans: pgvector 0.8.0+; hnsw.iterative_scan; max_scan_tuples 20,000 default
IVF: k-means lists; lists ~ rows/1,000 (<= 1 M rows), sqrt(rows) above
Filtering: tiny set -> exact scan; medium -> filtered traversal / iterative scan;
dominant key -> partition; broad -> ANN then filter
Hybrid: BM25 + vectors, RRF score = sum 1/(60 + rank); cross-encoder top 50
Evaluation: golden set, exact top-k offline; recall@10 and filtered fill
Model change: new index per model_version; backfill, shadow, cut over, rollback
Fan-out: query p99 = slowest shard; fewer shards or tenant routing
RED FLAGS
- No recall measurement
- Post-filtering selective tenant or ACL filters
- Mixed embedding models in one index
- Vector store as the only copy of the documents
- ACLs copied to a second store with no lag SLO
- Dedicated vector DB for a corpus a flat scan handles
- Index sized past available RAM