Hiring BarSupport

Vector Databases

Technology guide40 min read5 diagrams

Go deeper:

Why This Matters#

A vector database is not a smarter database. It is an approximate index with a recall dial: it returns probably the nearest neighbours to a query embedding, trading a few percent of correctness for orders of magnitude less latency and compute. Every design decision is a point on one surface: recall vs latency vs memory, with filtering, freshness and cost bolted on. The hard production problems are rarely the algorithm. They are filtered queries that silently return 3 results instead of 10, an index that no longer fits in RAM at 40 M vectors, and a model upgrade that makes every stored embedding incompatible with every new query.

That is why "we'll put the embeddings in a vector database" is a sentence interviewers press on. The L5 candidate draws a box labelled "vector DB" next to the LLM. The L6 candidate says "documents are chunked to ~500 tokens and embedded at 1,024 dimensions; vectors live in pgvector with an HNSW index (m 16, ef_construction 64), tenant filtering uses iterative scans so filtered queries still return k results, retrieval is hybrid with BM25 fused by reciprocal rank, and I'll measure recall@10 against exact search on a 1,000-query golden set before and after every index change." The L7 candidate asks what happens to 300 M stored vectors when the embedding model is replaced next year, and who pays for re-embedding the corpus.

The L5 → L6 gap is not knowing what HNSW stands for. It is knowing that a vector index is approximate by design, so recall must be measured, not assumed, and that the embedding model, not the database, is the schema.

The L5 → L6 → L7 Contrast#

BehaviorSenior (L5)Staff (L6)Principal (L7)
First move"Store embeddings in Pinecone / a vector DB""How many vectors, how many dimensions, what filters, what freshness, what recall target? At under ~10 M vectors with relational filters, pgvector next to the data is usually enough.""Is retrieval a shared platform or per-team? Which embedding models do we standardise on, and what does a model migration cost the org?"
Correctness"It finds similar items"Recall@k measured against brute force on a golden set; ef_search / nprobe tuned to a target (e.g. 0.95) at a latency budgetRetrieval quality as an SLO with owners; offline eval gates every model or index change
Filtering"Add a WHERE clause"Knows post-filtering starves results; chooses pre-filter, filtered traversal, iterative scans or partitions by filter selectivityDesigns tenancy so the most common filter is a partition boundary, not a predicate
Scale"It scales horizontally"Sizes RAM: vectors × dims × bytes plus graph overhead; quantization (int8, PQ, binary) when it doesn't fitPrices the memory tier vs disk-based indexes vs object-storage vector stores across 3 scales
Change"Re-index when needed"Model change = new index: dual-write, backfill, shadow-compare recall, cut over, keep rollbackTreats the embedding model as a versioned contract with a deprecation policy and a re-embedding budget
Ownership"The ML team owns it"Ingest pipeline owner, index owner and eval owner named; freshness SLO on embedding lagPlatform owns engines and SLOs; product teams own chunking, filters and relevance
Why "Correctness" separates levels

Approximate nearest neighbour (ANN) search does not return the true top-k. It returns a candidate set whose overlap with the exact top-k, recall@k, depends on index parameters. HNSW with a small ef_search may return 85% of the true neighbours; raising it to 200 may reach 99% at 3–5× the latency. Nothing in the query response tells you which you got. The Senior answer trusts the index. The Staff answer builds a golden set (say 1,000 representative queries), computes exact neighbours with a brute-force scan offline, and tracks recall@10 alongside p99 latency on every parameter or data change. The Principal answer ties that to end-task quality (answer correctness, click-through) because recall@10 of 0.99 on a bad embedding model is still bad retrieval.

Why "Filtering" separates levels

The naive plan is "find the 10 nearest vectors, then apply WHERE tenant_id = 42." If tenant 42 owns 0.1% of the corpus, the 10 nearest neighbours almost never include its rows, and the query returns 0–2 results with no error. Over-fetching (top 1,000, then filter) helps until the filter is selective enough to need top 100,000. The Staff answer picks by selectivity: a separate index or partition per large tenant, filtered graph traversal or iterative scans for medium selectivity, and exact search over a pre-filtered set when the filter leaves only a few thousand rows. pgvector added iterative index scans in 0.8.0 for exactly this case (pgvector).

The 60-Second Pitch#

"I'd embed documents in ~500-token chunks at 1,024 dimensions and store them in Postgres with pgvector, next to the metadata and permissions they need to be filtered by. At 8 M chunks that's ~33 GB of float32 vectors; stored as halfvec it's ~16 GB, and the HNSW index, which holds its own copy of the vectors plus ~1–2 GB of graph links, is ~18 GB, which fits in RAM on one memory-optimised instance with a replica. HNSW with m 16 and ef_construction 64; ef_search tuned to recall@10 of 0.95 at a p99 under 30 ms. Tenant filters use iterative scans, and our five largest tenants get their own partitions. Retrieval is hybrid: BM25 and vector results fused with reciprocal rank fusion, then a cross-encoder re-ranks the top 50. Every embedding carries a model version; a model change builds a new index in parallel, shadow-compares recall on a golden set and cuts over behind a flag. If we pass ~50 M vectors or need multi-node scale-out, I'd move the index to a dedicated engine and keep Postgres as the source of truth."

The Three Intents#

IntentConstraintStrategyFailure ModeCorrectness Bar
RAG / semantic document searchPer-tenant permissions, freshness in minutes, millions of chunkspgvector or a dedicated engine; hybrid with BM25; re-ranker; ACL filtersFiltered queries return too few results; stale or leaked chunksNever return a chunk the user can't read; recall@10 ≥ 0.9 on golden set
Recommendations / similarity at scale100 M–10 B item vectors, high QPS, loose filtersDedicated ANN engine or library (HNSW, IVF-PQ), sharded, quantizedRAM cost explodes; index rebuilds take hoursRecall traded for throughput; measured by online metrics
Deduplication / clustering / analyticsBatch, offline, whole-corpusFaiss-style library jobs, GPU, IVF-PQ; no online databaseTreating a batch job as an online serviceThroughput per dollar; near-dup precision

🎯 Staff Move: "I'll design for RAG over tenant documents. The permission filter is the hardest part, so the store has to filter correctly, not just search fast. That pushes me toward keeping vectors next to the metadata, and I'll revisit a dedicated engine only when the vector count or QPS forces it."

The Staff Positions#

PositionRationale
Recall is measured, never assumedANN is approximate by design; a golden set with exact neighbours is the only way to know where you are on the dial.
The embedding model is the schemaVectors from different models live in different spaces; a model change is a full re-embed and re-index.
Start with pgvector unless numbers say otherwiseBelow ~10–50 M vectors, transactional filters, permissions and joins beat a second system to keep in sync.
Filters decide the index designThe most selective common filter should be a partition or separate index, not a post-filter.
Hybrid beats pure vector for user-facing searchExact terms, IDs, names and rare words are where embeddings are weakest and BM25 is strongest.
Size RAM firstVectors × dimensions × bytes, plus graph overhead, decides instance type, shard count and whether you need quantization.
The source of truth is not the vector indexIndexes are derived and rebuildable; the documents and metadata live in the primary database or object storage.

Architecture & Internals#

Five internals change design decisions: embeddings and distance, the ANN index family, quantization, filtering, and the write path.

Embeddings and Distance#

An embedding model maps text, images or users to a fixed-length vector (commonly 384–3,072 dimensions). Similarity is a distance: cosine, inner product (equal to cosine on normalised vectors) or L2. Use the metric the model was trained for; normalise once at write time and use inner product for speed.

Storage per vector = dimensions x bytes per component
  1,536 dims x 4 bytes (float32)  = 6,144 bytes  (~6 KB)
  1,536 dims x 2 bytes (float16)  = 3,072 bytes
  1,536 dims x 1 byte  (int8)     = 1,536 bytes
  1,536 dims x 1 bit   (binary)   =   192 bytes
  PQ with 96 sub-quantizers x 1 B =    96 bytes  (~64x smaller than float32)

100 M vectors x 6 KB = ~614 GB raw float32, before any index overhead

The ANN Index Family#

IndexHow it worksStrengthsWeaknessesTypical knobs
Flat (exact)Scan every vectorPerfect recall; no buildO(N) per query; ~1 M vectors is the practical limit for online useNone
HNSWMulti-layer proximity graph; greedy search from a top-layer entry pointBest recall/latency trade in RAM; incremental insertsMemory-hungry; slow builds; deletes leave tombstonesM (links per node), ef_construction, ef_search
IVF (inverted file)k-means clusters ("lists"); search the nprobe closest listsSmaller, faster to build; pairs with PQNeeds training data; recall degrades as data drifts from centroidslists, nprobe
IVF-PQIVF plus product-quantized codesBillion-scale in modest RAMLower recall; usually needs re-ranking with full vectorslists, nprobe, PQ code size
Disk-based graph (DiskANN-style)Graph on SSD, compressed vectors in RAM5–10× cheaper per vector than all-RAMHigher latency (SSD reads per hop); tuning heavyBeam width, cache size

HNSW was introduced by Malkov and Yashunin (arXiv 1603.09320); it is the default in most engines today.

Diagram: The ANN Index Family

Why it matters in design: HNSW query cost grows roughly logarithmically with N, but memory grows linearly with N × (vector size + links), and a graph that spills out of RAM onto disk falls off a latency cliff. That is why the sizing math, not the QPS, usually picks the hardware.

The Recall Dial#

HNSW, illustrative behaviour on 10 M x 768-dim vectors, single node, in RAM:
  ef_search    recall@10    p50 latency    p99 latency
     20          ~0.85         ~1 ms          ~3 ms
     40          ~0.92         ~2 ms          ~5 ms     <- pgvector default is 40
    100          ~0.97         ~4 ms         ~10 ms
    400          ~0.995       ~12 ms         ~30 ms
  (shape is typical; measure on your data, it varies with dimensionality and distribution)

Raising M or ef_construction improves the graph (better recall at the same ef_search) at the cost of build time and memory; raising ef_search improves recall at query time at the cost of latency. pgvector's defaults are m 16, ef_construction 64 and hnsw.ef_search 40; IVFFlat defaults to probes 1, with suggested lists of rows/1,000 up to 1 M rows and √rows beyond (pgvector).

🎯 Staff Insight: "Recall and latency are one dial, and memory is the price of the dial's range. I pick a recall target with the product owner, say 0.95 at k=10, then find the cheapest ef_search that meets it under the p99 budget, and re-check both every time the corpus doubles."

Quantization: Paying for RAM With Recall#

TechniqueSize vs float32Recall impactTypical use
float16 / halfvec1/2Negligible for most modelsDefault storage choice when supported
Scalar int81/4Small; re-rank fixes most of itLarge HNSW indexes
Product quantization (PQ)1/16–1/64Noticeable; re-rank requiredBillion-scale IVF-PQ
Binary quantization1/32Large alone; good as a first-pass filter with re-rankVery high-dimensional models built for it

The standard pattern is two-stage: search the compressed index for the top 100–500 candidates, then re-score them with full-precision vectors fetched from disk or the primary store. Meta's Faiss work showed PQ codes letting 1 B vectors be indexed in under 30 GB of RAM (Faiss).

Filtering Strategies#

Diagram: Filtering Strategies

In pgvector 0.8.0+, hnsw.iterative_scan = relaxed_order keeps walking the graph until enough rows pass the filter, bounded by hnsw.max_scan_tuples (20,000 by default). Partial indexes work for a handful of low-cardinality filter values; partitioning works when there are many. Dedicated engines implement filter-aware graph traversal natively; check how each handles very selective filters before trusting it.

The Write Path#

ConcernWhat happensDesign consequence
Inserts into HNSWEach insert searches the graph to pick neighbours, ~ms eachBulk-load then build the index; streaming inserts cap ingest at thousands/s per node
UpdatesUsually delete + insertFrequent re-embedding of the same rows churns the graph
DeletesTombstones; space and recall recovered only on rebuild or vacuumPlan periodic rebuilds for high-churn corpora
Index buildCPU- and memory-heavy; fastest when the graph fits in maintenance_work_mem in PostgresBuild on a replica or new table, then swap
FreshnessEmbedding is an external call (10–100 ms per batch to a model)Async pipeline with an embedding-lag SLO, e.g. p99 < 5 min

Data Modeling / Core Usage — "The Entire Game"#

In Kafka the game is the partition key. In a vector store it is the retrieval contract: what a vector represents (the chunk), which model produced it (the version), and which filters must hold when it's returned (permissions, tenant, freshness). The index type is a tuning decision; these three are design decisions.

Step 1: Decide What a Vector Represents#

UnitTypical sizeGood forTrap
Whole document1 vector per docShort items: products, titles, profilesLong documents average into mush
Chunk (fixed window)200–800 tokens, 10–20% overlapRAG over manuals, tickets, wikisChunks cut sentences and tables; chunk count multiplies storage
Structural chunkSection, paragraph, functionDocs with headings, codeUneven sizes; very long sections still need splitting
Multiple vectors per itemTitle + body + imageMulti-modal and multi-field retrieval3× storage and write cost
Chunk math: 2 M documents x ~4,000 tokens / 500-token chunks = ~16 M chunks
            16 M x 1,024 dims x 2 bytes (halfvec) = ~33 GB of vectors
            Change chunking to 250 tokens -> 32 M chunks -> double everything

Step 2: Put the Model Version on Every Row#

create extension if not exists vector;

create table chunks (
  chunk_id      bigserial primary key,
  doc_id        bigint not null references documents(doc_id),
  tenant_id     bigint not null,
  acl_groups    bigint[] not null,          -- who may see it
  content       text not null,
  embedding     halfvec(1024) not null,
  model_version text not null,              -- e.g. 'embed-v3-1024'
  updated_at    timestamptz not null default now()
) partition by list (tenant_id);            -- big tenants get their own partitions

create index on chunks_default using hnsw (embedding halfvec_ip_ops)
  with (m = 16, ef_construction = 64);
create index on chunks_default (tenant_id);
-- Query: tenant-scoped, permission-filtered, model-matched
set hnsw.ef_search = 100;
set hnsw.iterative_scan = relaxed_order;     -- keep scanning until k rows pass filters

select chunk_id, doc_id, content
from chunks
where tenant_id = $1
  and acl_groups && $2                         -- user's groups
  and model_version = 'embed-v3-1024'
order by embedding <#> $3                      -- negative inner product
limit 10;

Keeping the vector next to tenant_id and acl_groups means the permission check happens in the same query, inside one transaction boundary. With a separate vector store, the ACL has to be copied into its metadata and kept in sync, and a lagging sync is a data leak, not a stale result.

🎯 Staff Move: "Permissions are filtered inside the retrieval query, not after it. If I move to a dedicated engine, ACLs replicate into it through the same change stream as the vectors, and I'll alert on replication lag because a stale ACL is a security bug."

Step 3: Design for Hybrid Retrieval#

Diagram: Step 3: Design for Hybrid Retrieval

Embeddings are strong on paraphrase and weak on exact tokens: product SKUs, error codes, names, rare terms. BM25 is the reverse. Reciprocal rank fusion needs no score calibration between the two lists, which is why it's the default; the constant 60 comes from the original RRF paper. Lexical search mechanics are in Elasticsearch and Elasticsearch vs Postgres Full-Text; the broader ranking stack is in Search Engine.

Step 4: Build the Golden Set Before the Index#

golden_set: 1,000 real queries (sampled from logs, stratified by tenant size and query type)
for each query: exact top-10 via brute-force scan (offline, once per corpus snapshot)
metrics per index config:
  recall@10      = |ANN top-10 ∩ exact top-10| / 10, averaged
  filtered_fill  = fraction of queries that returned the full k after filters
  p50 / p99 latency at production QPS
gate: no index, parameter or model change ships if recall@10 drops > 0.02
      or filtered_fill < 0.99

filtered_fill is the metric teams forget: it catches the post-filter starvation bug that recall on unfiltered queries hides.


The Tunable Tradeoff — Recall × Latency × Memory#

Every vector-search decision moves along three linked axes. You can have any two cheaply; the third costs money.

SettingCheap / fast endAccurate endWho pays at the cheap end
ef_search / nprobeLow (20, 1)High (200+, 32+)Users, through missed results
Graph degree M8–1232–64Recall ceiling; product relevance
Vector precisionBinary or PQfloat32Relevance, unless re-ranked
Storage mediumDisk-based or object-storage indexAll in RAMLatency: SSD hops add ms per query
Re-rankingNoneFull-precision + cross-encoderRelevance at the top of the list
FreshnessNightly batch rebuildStreaming insertsUsers who search for what they just wrote
RAM sizing (HNSW, rule of thumb):
  bytes ≈ N x (d x bytes_per_dim + M x 2 x ~8 bytes for links) x 1.2 overhead

  20 M x 768 dims, float32, M=16:
    vectors  20 M x 3,072 B   = 61 GB
    links    20 M x 256 B     = 5 GB
    total    ~66 GB x 1.2     = ~80 GB  -> a 128 GB node, or 2 shards on 64 GB nodes
  same corpus, int8 + re-rank from disk:  ~20 GB in RAM

Check your own numbers with the Back-of-Envelope Calculator and, when the query sits inside a user request, against the Latency Budget.

🎯 Staff Move: "I'll quantize to int8 and re-rank the top 200 with full-precision vectors. That cuts RAM about 4× and costs us roughly a point of recall, which the re-rank wins back. I'd rather spend money on a re-ranker than on RAM for float32 vectors nobody reads at full precision."

Who Pays for Each Choice#

ChoiceWhat WorksWhat BreaksWho Pays
Default ef_search, never measuredFast, simpleRecall unknown, often 0.85–0.9Users and the LLM answering from wrong context
Post-filteringEasy to implementSelective filters return too few results, silentlySmall tenants, who get worse search
Separate vector storeScale-out, rich ANN featuresACL and metadata sync, two systems to operateSecurity on sync lag; on-call for two systems
All-in-RAM float32Best latency and recallMemory bill grows linearlyFinance
Aggressive quantization without re-rankCheapest RAMRelevance drops noticeablyProduct quality
Streaming inserts into HNSWFresh resultsIngest throughput limits, graph fragmentationIngest pipeline on-call

Anti-Patterns — What Kills Vector Database Deployments#

1. Never Measuring Recall#

The index is tuned once with defaults and never evaluated. Corpus grows 10×, recall slides from 0.95 to 0.8, and the team blames the LLM. Fix: a golden set with exact neighbours, recall@k and filtered fill tracked on every change.

2. Post-Filtering Selective Queries#

"ANN top 10, then WHERE tenant_id" returns empty pages for small tenants. Fix: partitions or per-tenant indexes for large tenants, iterative scans or filtered traversal for medium selectivity, exact search for tiny filtered sets. See Multi-Tenancy.

3. Mixing Embedding Models in One Index#

New documents embedded with v2, old ones still v1, queries embedded with v2. Distances across models are meaningless, so old documents silently stop being found. Fix: model_version on every row, one index per model, a migration plan for every model change.

4. The Vector Store as Source of Truth#

Documents exist only as chunks in the vector index. A model change, chunking change or index corruption means re-ingesting from wherever the data originally came from, if it still exists. Fix: canonical documents in a database or object storage; the index is derived and rebuildable.

5. Pure Vector Search for Exact-Match Queries#

Users search "ERR_4021" or an order ID and get semantically similar but wrong results. Fix: hybrid retrieval with BM25, or route ID-shaped queries to an exact lookup.

6. Unbounded Chunk Multiplication#

Halving chunk size, adding overlap and adding a title vector each multiply storage; together they turn 10 M vectors into 60 M. Fix: treat chunking as a costed schema decision with an eval behind it.

7. Building the Index in Place During Peak#

HNSW builds are CPU- and memory-heavy and can starve queries on the same node for hours. Fix: build on a replica or a new table, then swap; schedule rebuilds off-peak.

8. A Dedicated Engine for 200,000 Vectors#

A second database, sync pipeline and on-call rotation for a corpus a flat scan or pgvector handles in milliseconds. Fix: start next to the data; move when numbers force it. See Buy or Build.


The Technology Landscape — Head-to-Head Comparison#

Dimensionpgvector (Postgres)Dedicated engines (Milvus, Qdrant, Weaviate, Vespa)Managed vector servicesSearch engines with kNN (Elasticsearch, OpenSearch)Object-storage vector indexes (e.g. S3 Vectors)ANN libraries (Faiss, hnswlib, ScaNN)
ModelIndex type inside a relational DBPurpose-built vector DB, shardedHosted API, serverless or podsInverted index plus HNSW fieldsVector indexes in object storage, API queriesIn-process library, you build the service
Sweet spotUp to ~10–50 M vectors with relational filters50 M–billions, high QPSTeams without ops capacityHybrid lexical + vector in one engineLarge, infrequently queried corpora at low costBatch jobs, custom serving, research
FilteringFull SQL; iterative scans for ANN + filterNative filter-aware traversalMetadata filtersFull query DSLMetadata filters (limited size)Yours to build
Transactions with source dataYesNo; sync pipelineNoNoNoNo
Ops burdenLow if you already run PostgresMedium–highLowMediumLowHigh (you own serving)
Latency profilems in RAMms in RAM; tiered optionsms, network hopms–tens of msSub-second, not single-digit msµs–ms in process

S3 Vectors, for example, documents up to 2 billion vectors per index and 1–4,096 dimensions, with write rates capped per index (S3 Vectors limits), which makes it a fit for cheap, large, lower-QPS retrieval rather than a hot path.

🎯 Staff Insight: "The real comparison isn't pgvector versus a vector database. It's one system with transactional filters versus two systems with a sync pipeline. I pay for the second system only when vector count, QPS or ANN features force me to, and I write down which number triggered it."


Patterns#

Pattern 1: Vectors Next to the Data (pgvector)#

Embeddings live in the same Postgres as documents and ACLs; HNSW index; partitions for large tenants; read replicas for query scale. The default for RAG under a few tens of millions of chunks. See PostgreSQL.

Pattern 2: Source of Truth + Derived Vector Index#

Diagram: Pattern 2: Source of Truth + Derived Vector Index

The primary database stays the truth; a change stream (see Transactional Outbox and Kafka) drives embedding and indexing; deletes and ACL changes flow through the same stream with priority. The query service re-checks permissions against the primary for the final page if the domain is sensitive.

Pattern 3: Two-Stage Retrieval (Compressed Recall, Precise Rank)#

Quantized ANN for the top 200–1,000 candidates, full-precision re-score, then a cross-encoder for the top 20–50. Each stage is 10× smaller and 10× more expensive per item than the last.

Pattern 4: Blue/Green Index for Model Migration#

New model → new embeddings column or index built in the background → shadow queries compare recall and task metrics → cut over by flag per tenant → keep the old index for a rollback window → drop. Same discipline as Online Migrations.

Pattern 5: Per-Tenant Partitions for Isolation#

Large tenants get dedicated partitions or collections; the long tail shares one index with a tenant filter. Removes the post-filter problem for the tenants who matter most and makes tenant deletion a partition drop.

Pattern 6: Semantic Cache#

Embed the incoming query; if a previous query's embedding is within a tight distance threshold, return its cached answer. Saves LLM calls on repetitive traffic; needs a conservative threshold and tenant scoping, or users see each other's answers. See Caching Fundamentals.


Scaling#

The Numbers#

ResourceDocumented limit or typical valueDesign note
pgvector indexable dimensionsvector 2,000; halfvec 4,000; bit 64,000; sparsevec 1,000 non-zero3,072-dim models need halfvec to be indexed
pgvector HNSW defaultsm 16, ef_construction 64, ef_search 40Tune ef_search per query; it's a session setting
pgvector iterative scan boundhnsw.max_scan_tuples 20,000Very selective filters can still under-fill; use partitions
Flat scan practical limit~1 M vectors per query onlineFine for small tenants and pre-filtered sets
HNSW insert cost~ms per insert per node (typical)Bulk-load, then build
float32 vector at 1,536 dims~6 KB100 M vectors ≈ 614 GB before index overhead
Object-storage vector index (S3 Vectors)Up to 2 B vectors per index, 1–4,096 dims, ~1,000 write requests/s per indexCheap at rest; not a single-digit-ms path

Sharding and Tail Latency#

ANN indexes shard by random or hash assignment (each shard holds a slice of all vectors, and every query fans out to every shard) or by a partition key such as tenant (queries go to one shard). Fan-out multiplies tail latency: the query is as slow as its slowest shard.

Fanning a query out to many shards makes its latency the maximum of all shard latencies, so p99 grows with shard count.
Fan-out math: 16 shards, each with p99 = 10 ms
  P(query avoids every shard's slowest 1%) = 0.99^16 ≈ 0.85
  -> ~15% of queries see at least one shard at its p99
  -> query p99 is well above 10 ms; mitigate with hedged requests,
     fewer bigger shards, or tenant-routed shards

Scaling Moves in Order#

  1. Measure first: golden-set recall and filtered fill at current parameters, so every later step has a baseline.
  2. Shrink the vectors: halfvec, then int8 or PQ with re-ranking. Often a 2–4× RAM cut for under a point of recall.
  3. Reduce the corpus: deduplicate near-identical chunks, drop boilerplate, revisit chunk overlap.
  4. Partition by the dominant filter (tenant, language, region) so most queries hit one smaller index.
  5. Add read replicas for QPS while the index still fits one node.
  6. Shard or move to a dedicated engine once a single node can't hold the index in RAM at your recall target, or rebuilds take longer than your freshness SLO allows.
  7. Tier cold data to disk-based or object-storage indexes for large, rarely queried corpora.

🎯 Staff Move: "I scale a vector index by making it smaller before making it bigger: halfvec, quantization with re-rank, dedupe, then partitions by tenant. Sharding comes last because fan-out costs me tail latency on every query."


Failure Modes & Recovery#

1. Silent Recall Decay#

  • Symptom: Answer quality complaints rise; latency and error dashboards are green.
  • Root cause: Corpus grew 5–10× at fixed ef_search; IVF centroids trained on old data; tombstones from deletes.
  • Detection: Nightly golden-set recall_at_10 and filtered_fill; online proxy metrics (click-through, answer acceptance).
  • Fix: Raise ef_search / nprobe to restore the target; retrain IVF or rebuild HNSW.
  • Prevention: Recall gate in CI for index changes; scheduled rebuilds tied to corpus growth.

2. Filtered Queries Return Too Few Results#

  • Symptom: Small tenants report "search finds nothing"; large tenants are fine.
  • Root cause: Post-filtering or a scan bound reached before k rows pass a selective filter.
  • Detection: filtered_fill by tenant size bucket; count of queries returning fewer than k.
  • Fix: Enable iterative scans, raise the scan bound, or route small tenants to exact search.
  • Prevention: Partition design by filter selectivity; per-tenant eval slices.

3. Index Outgrows RAM#

  • Symptom: p99 jumps from ~10 ms to hundreds of ms as the index starts hitting disk; CPU waits on I/O.
  • Root cause: Corpus growth or a chunking change pushed the index past memory.
  • Detection: Index size vs available memory; buffer cache hit ratio; page-fault and disk-read rates on the query nodes.
  • Fix: Quantize, move to a larger node, or shard; temporarily shed low-priority query traffic.
  • Prevention: Capacity alert at 70% of RAM; sizing review for every chunking or model change.

4. Model Mismatch After a Partial Migration#

  • Symptom: Relevance collapses for older documents; newer ones are fine.
  • Root cause: Queries embedded with the new model hit vectors from the old one, or a backfill stopped halfway.
  • Detection: Count of rows by model_version; recall per model slice; query-time assertion that query and index versions match.
  • Fix: Route queries to the index matching the stored version until backfill completes; resume backfill.
  • Prevention: One index per model version; blue/green cutover; version check enforced in the query service.

5. Ingest Lag and Stale or Leaked Results#

  • Symptom: New documents not searchable for hours; worse, deleted or permission-revoked documents still returned.
  • Root cause: Embedding provider throttling, a stuck consumer, or ACL updates queued behind bulk re-embeds.
  • Detection: Embedding lag p99; queue age; count of results whose ACL check fails at the primary.
  • Fix: Prioritise deletes and ACL changes over inserts; scale workers; final-page permission re-check against the primary.
  • Prevention: Separate priority lanes; embedding-lag SLO with alerting; deletes applied as filters immediately even before index compaction.

Operational Reality Matrix#

FailureDetection SignalBlast RadiusMitigationOwner
Recall decayGolden-set recall, online quality proxiesAll queries, graduallyRetune, rebuildSearch/retrieval team
Filter starvationFiltered fill by tenant sizeSmall tenantsIterative scans, partitionsRetrieval team; product for tenancy model
Index exceeds RAMMemory headroom, cache hit ratioAll queries on the nodeQuantize, scale, shardPlatform/DB team
Model mismatchRows by model versionOlder contentVersion routing, finish backfillML platform + retrieval team
Ingest lag / ACL lagEmbedding lag, ACL re-check failuresFreshness; security if ACLs lagPriority lanes, re-check at primaryIngest owner; security reviews

When to Use vs. Alternatives#

NeedPickWhy
RAG over < ~10–50 M chunks with permissionspgvector in the existing PostgresOne system, transactional filters, SQL joins
Hundreds of millions to billions of vectors, high QPSDedicated vector engineSharding, quantization, filter-aware traversal
Lexical + vector in one query with rich text featuresElasticsearch / OpenSearch kNNHybrid and BM25 in one engine
Large, rarely queried corpus where cost dominatesObject-storage vector indexPay mostly for storage, not RAM
Offline dedupe, clustering, candidate generationFaiss / ScaNN in batch jobsNo database needed; GPUs help
Exact-match, IDs, structured filters onlyNo vector index: B-tree, inverted indexEmbeddings add cost without benefit

When NOT to Use a Vector Database#

  • The corpus is small. Under ~100K–1 M vectors, an exact scan in memory or in Postgres is fast enough and has perfect recall.
  • The queries are exact. IDs, SKUs, codes and structured filters belong in B-tree or inverted indexes; see Database Indexes.
  • There's no evaluation. Without a golden set and a quality metric, you can't tell a good index from a bad one, and adding one adds risk without measurable value.
  • Permissions are complex and the engine can't filter them. Copying fine-grained ACLs into a separate store with lag is a leak waiting to happen.
  • As the system of record. The index is derived; keep the documents elsewhere.

Operational Concerns#

What the On-Call Actually Does#

  1. Watches query p99 and memory headroom on index nodes; the RAM cliff is the most common performance incident.
  2. Watches ingest: embedding lag, provider throttling, queue age, and the ACL/delete lane separately from bulk inserts.
  3. Reviews nightly recall and filtered-fill reports and treats a drop of more than 0.02 as an incident ticket.
  4. Runs rebuilds on replicas or shadow indexes, then swaps; never rebuilds in place at peak.
  5. Drives model migrations: backfill progress by model_version, shadow comparison, per-tenant cutover, rollback window.

Key Metrics & Alerts#

MetricHealthyAlert
Query latency p99Within budget (e.g. < 30 ms)2× budget for 10 min
Index size / RAM< 70%> 85% (page before the cliff)
Golden-set recall@10≥ target (e.g. 0.95)Drop > 0.02 day over day
Filtered fill (queries returning k)≥ 0.99< 0.97 in any tenant bucket
Embedding lag p99< 5 min> 30 min
ACL/delete propagation lag< 1 min> 5 min (security ticket)
Rows not on current model_version0 outside migrationsAny growth after cutover

Interview Application — Staff-Level Plays#

Case StudyHow Vectors Are UsedKey Pattern
Search EngineSemantic retrieval stage beside BM25Hybrid retrieval, fusion, re-ranking
News FeedCandidate generation from user and item embeddingsANN for recall, ranker for precision
Dating & MatchingSimilar-profile candidates under hard filtersFilter selectivity drives index design
TypeaheadSemantic suggestions for long-tail queriesPrefix index first; vectors only as a fallback
Web CrawlerNear-duplicate page detectionBatch ANN or hashing, not an online DB

Every System Design Question Has a Vector Moment#

  • Support chatbot: "Tickets and docs are chunked and embedded into pgvector with tenant and ACL columns; retrieval is hybrid and filtered in one query; recall@10 on a golden set gates every change."
  • E-commerce search: "SKUs and brand names go through BM25; descriptive queries get vector results; RRF merges them; a learned ranker orders the top 100."
  • Recommendations: "Item embeddings in a sharded ANN index produce 500 candidates in ~10 ms; the ranker does the expensive work on those 500."
  • Content moderation: "Known-bad images are embedded; new uploads are checked for near-duplicates above a tuned similarity threshold, with precision measured on labelled pairs."

What Interviewers Probe#

After You Say...They Will Ask...What They're Evaluating
"Store embeddings in a vector DB""Why not Postgres? How many vectors?"Sizing and avoiding a second system without cause
"HNSW index""How do you know the results are right?"Recall measurement against exact search
"Filter by tenant""What happens for a tenant with 0.1% of the data?"Post-filter starvation and partition design
"We'll upgrade the embedding model""What happens to the 200 M stored vectors?"Re-embedding, blue/green index, version tagging
"It scales horizontally""What does fan-out do to p99?"Tail latency across shards
"Semantic search for everything""What about an exact order ID?"Hybrid retrieval

Common Interview Mistakes#

What Candidates SayWhat Interviewers HearWhat Staff Engineers Say
"The vector DB returns the nearest neighbours"Doesn't know ANN is approximate"It returns approximate neighbours; I measure recall@10 against exact search."
"Filter the results by user"Post-filter starvation"Filtering happens inside the search, with partitions for big tenants."
"We'll use a dedicated vector DB" (at 1 M vectors)Adds a system for no reason"pgvector next to the data until vector count or QPS forces a move."
"Re-index when we change models"Hasn't priced it"Blue/green: new index, backfill, shadow-compare, cut over per tenant."
"More dimensions are better"Ignores memory and cost"Dimensions are RAM; I'd test a smaller model or truncation against the golden set."

L5 vs L6 vs L7 Responses#

ScenarioL5 AnswerL6 / Staff AnswerL7 / Principal Answer
"Add semantic search to our docs"Embed everything into a managed vector DBpgvector with tenant/ACL filters, hybrid BM25, golden-set eval, embedding-lag SLOShared retrieval platform with standard models, eval harness and cost per query
"Search quality dropped"Try a bigger modelCheck recall decay, filtered fill, model-version mix before touching the modelQuality SLO with owners; eval gate on every model and index change
"Index doesn't fit in RAM"Bigger instancehalfvec / int8 with re-rank, dedupe, partitions, then shardTiering policy: hot RAM index, cold disk or object-storage index, priced per tenant
"Switch embedding providers"Re-embed overnightBlue/green index, shadow traffic, per-tenant cutover, rollback windowModel deprecation policy, re-embedding budget, contract terms to avoid lock-in

The Staff Vector Database Checklist#

  1. Size it: "16 M chunks × 1,024 dims × 2 bytes ≈ 33 GB; plus graph links; one node with a replica."
  2. Define the unit: "500-token chunks with 15% overlap; model version on every row."
  3. Filters first: "Tenant and ACL filters inside the query; partitions for the largest tenants."
  4. Measure: "Golden set of 1,000 queries; recall@10 ≥ 0.95 and filtered fill ≥ 0.99 at p99 < 30 ms."
  5. Hybrid: "BM25 plus vectors, RRF, cross-encoder on the top 50."
  6. Plan the migration: "A model change is a blue/green index with shadow comparison and per-tenant cutover."

🎯 Staff Insight: Don't use a vector index as the system of record, as a substitute for exact lookups, or without an evaluation set. The strongest signal is saying how you know the approximate answers are good enough, and what happens to every stored vector when the model changes.

Evaluation Rubric#

DimensionSenior (L5)Staff (L6)Principal (L7)
CorrectnessTrusts the indexRecall@k and filtered fill on a golden setTask-level quality SLO and eval gates org-wide
Sizing"It scales"RAM math, quantization, partitioning, shard trade-offsTiered storage and cost per query across products
FilteringWHERE clauseSelectivity-driven strategy; ACLs inside the queryTenancy model designed around retrieval
Change"Re-index"Blue/green model migration with shadow compareModel lifecycle policy and re-embedding budget
Choice of engineDedicated DB by defaultpgvector until a stated triggerPlatform standard, exceptions process, build vs buy

Beyond Staff: The Principal View#

Why L7 Sees This Problem Differently#

At Staff level a vector database is an index you size and tune. At Principal level retrieval is shared infrastructure every AI feature depends on, and the embedding model is a company-wide schema. When five teams pick five models, five engines and five chunking schemes, the org pays five times for storage, five times for re-embedding, and has no common way to say whether retrieval got better or worse. The L7 question is "which models do we standardise on, who owns the eval harness, and what's the plan when the model we chose is deprecated?"

🧭 Principal Move: "Before we choose an engine, I want two standards: a small set of approved embedding models with a version and deprecation policy, and one evaluation harness every team uses. Engines are replaceable; ungoverned embeddings in a dozen places are not."

The Org-Level Fault Line#

Shared retrieval platform vs per-team vector stores.

OptionWhat WorksWhat BreaksWho Pays
Every team picks its own store and modelFast experimentsDuplicate corpora, inconsistent ACL handling, N migrations per model changeSecurity and finance
One central vector platform for everythingShared ops, consistent evalsBecomes a bottleneck; one-size indexing fits nobodyProduct velocity
Paved road: pgvector by default, platform engine above a threshold, shared eval harnessSimple default, scale path, common quality barNeeds a platform team and clear triggersPlatform headcount (2–4 engineers)

The Principal default: a paved road. Postgres with pgvector for most teams; a platform-run dedicated engine for corpora above a stated size or QPS; approved models with version tags; one eval harness; ACL propagation as a platform guarantee with a lag SLO.

Cost Model#

Assumptions: memory-optimised nodes ~$8–12 per GB of RAM per month including replicas; embedding at ~$0.02–0.10 per million tokens (varies widely by provider and model); loaded engineer ~$250K/year. Directional only.

ScaleVectorsRAM needed (halfvec + graph, 2 replicas)Infra/monthRe-embed corpus onceHeadcountTotal/month
Startup2 M × 1,024~10 GB~$300–600 (inside existing Postgres)~$20–1000.1 FTE~$2–3K
Growth100 M × 1,024~450 GB~$4–6K~$1–5K1 FTE (~$21K)~$25–30K
Enterprise5 B × 768, int8 + re-rank~10 TB~$90–130K~$50–250K per migration4–6 FTE (~$100K)~$200–250K

At growth and enterprise scale, people and migrations rival infrastructure. A model change at 5 B vectors is a quarter-long programme, so the model standard is the biggest cost lever, followed by quantization.

The 3-Year Evolution Path#

Diagram: The 3-Year Evolution Path

One-Way Doors vs Two-Way Doors#

DecisionReversibilityCost to Reverse
Embedding model for a large corpusOne-way-ishFull re-embed and re-index; weeks at billions of vectors
Storing documents only in the vector storeOne-way if sources disappearRe-ingest from origin, if possible
Chunking schemeOne-way-ishRe-chunk, re-embed, re-index
Engine choice with a CDC-fed derived indexTwo-wayRebuild index in the new engine from the source of truth
HNSW parameters, ef_search, quantizationTwo-wayRebuild or session setting
Hybrid fusion and re-rankingTwo-wayCode change

The Standard I'd Write#

RFC-RETR-001: Vector Retrieval Baseline

Scope: Every production feature that retrieves by embedding similarity.

MUST
  1. Use an approved embedding model; store model_version on every vector.
  2. Keep canonical documents in a primary database or object storage; vector
     indexes are derived and rebuildable from them.
  3. Enforce tenant and permission filters inside the retrieval query; ACL and
     delete propagation lag p99 < 1 minute where a separate store is used.
  4. Maintain a golden set (>= 500 queries) and report recall@10 and filtered
     fill; changes that drop recall by > 0.02 do not ship.
  5. Migrate models blue/green with shadow comparison and a rollback window.
SHOULD
  6. Use pgvector below 50 M vectors unless an exception is approved.
  7. Combine lexical and vector retrieval for user-facing search.
  8. Quantize with full-precision re-ranking above 100 GB of index.

Exceptions: retrieval platform lead approval, reviewed each quarter.
Success metrics: zero cross-tenant retrieval incidents; every index on a supported
  model; recall reported for 100% of production indexes; cost per 1K queries per team.

What I'd Tell the VP#

"Every AI feature we ship depends on finding the right documents first, and today each team does that differently, with different models and no shared way to measure whether it works. That costs us in duplicated infrastructure and, more importantly, makes every future model upgrade five separate projects. I'm proposing a small set of approved models, one quality test every team runs, and a default storage approach in the database we already operate, with a dedicated system only for the largest workloads. That keeps costs predictable, makes upgrades a routine instead of a crisis, and closes the risk of one customer's documents appearing in another's answers. The cost is a platform team of two to three engineers."

Principal Interview Signals#

SignalWhat It Sounds Like
Model as schema"The embedding model is a versioned contract; changing it is a migration, not a deploy."
Quality as an SLO"Recall@10 and filtered fill are reported like latency, with an owner."
Prices migrations"At 5 B vectors a model change is a quarter of work; that's why we standardise."
Security in retrieval"ACLs propagate with a lag SLO; a stale ACL is a leak."
Defaults with triggers"pgvector until 50 M vectors or a QPS we can't serve from replicas."

Staff answers that L7 interviewers find insufficient:

  • "We'll benchmark engines and pick the fastest" without who owns the eval harness or the migration path.
  • "Use the best embedding model" without its deprecation risk and re-embedding cost.
  • "Filter by tenant" without the org's tenancy and ACL propagation guarantees.

How Real Companies Built It#

Meta's research team released Faiss in 2017 as a library for similarity search over billions of vectors. The announcement describes building what it calls the first k-nearest-neighbour graph on 1 billion high-dimensional vectors, being 8.5× faster than the previously reported state of the art, GPU implementations typically 5–10× faster than CPU, and product-quantized indexes that fit 1 billion vectors in under 30 GB of RAM (Engineering at Meta).

Staff insight: Compression is what made billion-scale search practical. When an interviewer pushes the corpus to billions, the answer is quantization plus re-ranking, not more RAM.

Spotify — From Annoy to Voyager#

Spotify open-sourced Annoy in 2013 and used it for a decade behind features like Discover Weekly. In 2023 it released Voyager, an HNSW-based library built on hnswlib, reporting more than 10× the speed of Annoy at the same recall, up to 50% more accuracy at the same speed, up to 4× less memory than Annoy, and production use across Spotify teams since 2022 (Spotify Engineering).

Staff insight: The index algorithm changed under a decade-old product without changing the product. That's the argument for keeping the index derived and swappable behind a stable retrieval interface.

Google — ScaNN and Anisotropic Quantization#

Google Research's ScaNN library introduced anisotropic vector quantization, which penalises quantization error parallel to the original vector more heavily because that error distorts the high inner products that matter for ranking. On the glove-100-angular benchmark, Google reported roughly twice the queries per second of the next-fastest library at a given accuracy (Google Research blog).

Staff insight: Quantization isn't neutral compression; how you compress decides which errors you accept. In an interview, "we'll quantize" is incomplete until you say how you'll measure the recall it costs.


Practice Drill#

Prompt: "Your B2B product launched an AI assistant that answers from each customer's documents. It runs on a managed vector service with 60 M chunks. Large customers love it; small customers say it 'can't find anything.' Last week one customer saw a snippet from a document they had deleted. The ML team wants to switch to a new embedding model next quarter. What do you do?"

Staff Answer

Three problems, in priority order. The deleted-document snippet is a security incident, so it goes first. A deleted document still being retrievable means deletes propagate to the vector store asynchronously and lag, or tombstoned vectors are still returned. Immediate fix: re-check every returned chunk against the primary database (document exists, user has access) before it reaches the model, which closes the leak regardless of index state. Then put deletes and ACL changes on a priority lane in the change stream with a propagation-lag SLO of p99 under a minute, and alert on it. Small customers finding nothing is almost certainly post-filter starvation: a small tenant owns ~0.01% of 60 M chunks, so an unfiltered top-k rarely contains its rows. I'd measure filtered fill by tenant size, then fix it by selectivity: tenants under ~50K chunks get exact search over their own rows (a few ms), mid-size tenants use the engine's filtered traversal with a higher candidate budget, and the largest tenants get their own collections. The model change becomes a blue/green migration: tag every vector with model_version, build the new index from the source documents in the background (60 M chunks is a few days of embedding at provider rate limits), shadow-query both on a 1,000-query golden set per tenant segment comparing recall@10 and answer acceptance, cut over per tenant behind a flag, keep the old index for two weeks. Metrics I'd report weekly: ACL/delete lag, filtered fill by tenant bucket, recall@10, p99 latency and cost per 1K queries.

Why this is L6:

  • Treats the stale-delete as a security bug with an immediate, index-independent fix (final permission check at the primary).
  • Diagnoses small-tenant failures as filter selectivity, and chooses a strategy per selectivity band.
  • Turns the model change into a measured blue/green migration instead of an overnight re-embed.

What L7 adds:

  • Asks why ACLs live in two systems at all: for a B2B product with fine-grained permissions, pgvector next to the source data may remove the sync problem entirely.
  • Sets an org standard: approved models with version tags, a shared eval harness, and a re-embedding budget per model change.
  • Prices the migration (embedding cost, double index for the overlap window, engineering weeks) and uses it to negotiate model and vendor terms.
❌ Common L5 Trap

"Increase top-k to 100 so small customers get results, rebuild the index to purge deleted documents, and re-embed everything with the new model over a weekend."

Why this misses: Over-fetching only helps until a filter is more selective than k/N, so the smallest tenants still starve. A rebuild purges today's deletes but not tomorrow's, leaving the leak open without a propagation SLO or a final permission check. And an in-place weekend re-embed mixes model versions mid-flight with no recall comparison or rollback.


Quick Reference Card#

Core idea:         approximate nearest neighbours; recall is a dial, not a given
Storage:           dims x bytes; 1,536 x float32 = ~6 KB; 100 M vectors ~614 GB raw
Precision:         float16 1/2, int8 1/4, PQ 1/16-1/64, binary 1/32 (re-rank after)
HNSW:              graph; knobs M, ef_construction, ef_search; RAM-hungry; tombstone deletes
pgvector defaults: m 16, ef_construction 64, ef_search 40; IVFFlat probes 1
pgvector limits:   index vector <= 2,000 dims, halfvec <= 4,000, bit <= 64,000
Iterative scans:   pgvector 0.8.0+; hnsw.iterative_scan; max_scan_tuples 20,000 default
IVF:               k-means lists; lists ~ rows/1,000 (<= 1 M rows), sqrt(rows) above
Filtering:         tiny set -> exact scan; medium -> filtered traversal / iterative scan;
                   dominant key -> partition; broad -> ANN then filter
Hybrid:            BM25 + vectors, RRF score = sum 1/(60 + rank); cross-encoder top 50
Evaluation:        golden set, exact top-k offline; recall@10 and filtered fill
Model change:      new index per model_version; backfill, shadow, cut over, rollback
Fan-out:           query p99 = slowest shard; fewer shards or tenant routing

RED FLAGS
  - No recall measurement
  - Post-filtering selective tenant or ACL filters
  - Mixed embedding models in one index
  - Vector store as the only copy of the documents
  - ACLs copied to a second store with no lag SLO
  - Dedicated vector DB for a corpus a flat scan handles
  - Index sized past available RAM
  1. Loading the index…