Learning pathsA
GUIDED PRACTICE

Vector Databases

Vector retrieval finds items whose embeddings are close under a chosen distance function.

Vector retrieval finds items whose embeddings are close under a chosen distance function. Embeddings are model-produced coordinates, not truth or authorization. A useful system versions the embedding model, filters by current access, measures recall and latency, and has a plan for re-embedding when either the model or source content changes.

The mechanism at a glance

Source records → Embedding worker (content revision); Embedding worker → Vector index (model-versioned vector); Candidate search → Vector index (nearest candidates); Vector index → ACL authority (current access check); ACL authority → Reranker (allowed candidates); Reranker → Candidate search (ranked results)
Scroll to inspect the diagram, or open it at full size.

Figure — Source records → Embedding worker (content revision); Embedding worker → Vector index (model-versioned vector); Candidate search → Vector index (nearest candidates); Vector index → ACL authority (current access check); ACL authority → Reranker (allowed candidates); Reranker → Candidate search (ranked results)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. 1 · Frame

Define the retrieval task, corpus size, embedding dimension, update/delete rate, metadata filters, latency target, and quality metric. Clarify whether nearest neighbors are global or tenant-scoped and whether exact results are required. Similarity metrics such as cosine or dot product are only meaningful under the model’s training and normalization assumptions.

2. 2 · Model

Store object ID, vector, model/version, source revision, tenant, and filterable metadata. An approximate nearest neighbor index such as graph-based or inverted-file search narrows candidates; a reranker may compare a larger candidate set. Keep source content and ACL authority elsewhere and join only after access checks.

3. 3 · Scale

Tune search breadth against recall, tail latency, and memory. Partition or filter by tenant only if it preserves enough candidate quality. Frequent deletes may leave tombstones or require segment rebuilds. Use a stable model version per index; dual-write or backfill a new version and compare quality before switching traffic.

4. 4 · Recover

On a query, embed with the matching model version, retrieve candidates, apply authorization and metadata filters, then rank and return safe source references. Delete or permission changes must reach the response boundary immediately even if index cleanup lags. Monitor recall on labeled queries, empty-result rate, index age, and model-version mix.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
VectorRecord(object_id, tenant_id, model_version, vector, source_revision)
Metadata(object_id, attributes, access_scope, deleted_version)
ANN index: metric + build parameters + segment version
Query: embed -> candidates -> current ACL/filter -> rerank

Worked example

A new embedding model improves semantic recall but changes vector geometry. Build a parallel index from the same source revision, evaluate it on a fixed labeled set, and shadow queries. Switch only after recall and latency meet the target. During rollout, query vectors must match the index model; comparing different model spaces is invalid.

Failure walkthrough

A document becomes private while its vector remains searchable for several minutes. Candidate retrieval can still find its ID, but the authorization join must suppress the result before title, snippet, or score explanation is returned. Deletion propagation and cache invalidation follow. Index recall tests alone cannot detect this privacy failure.

Source revision changes → New model embeds it → Parallel index receives version → Query retrieves candidates → ACL check removes revoked record → Canary results are compared
Scroll to inspect the diagram, or open it at full size.

Figure — Source revision changes → New model embeds it → Parallel index receives version → Query retrieves candidates → ACL check removes revoked record → Canary results are compared

Decisions and trade-offs

DecisionImprovesCosts
Larger ANN search breadthRecallLatency and compute
Metadata prefilterTenant isolation and speedSmaller candidate pool may reduce recall
Exact rerankOrdering qualityExtra fetch and compute
Parallel model indexSafe model migrationTemporary duplicate storage and writes

Check your understanding

Plan a model-version migration while maintaining deletes and access restrictions. Which evaluation gates must pass before routing queries to the new index?

Show answer and explanation

Answer: Backfill from versioned source records while tailing changes and deletions. Verify index completeness and model compatibility; compare recall on a fixed labeled query set, latency percentiles, filter behavior, and deletion/ACL tests. Shadow or canary traffic, then switch routing with rollback. Keep the old index until the new version is stable.

Primary documentation

Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.

Continue the connection

Study ChatGPT and explain which guarantee from this lesson carries into that topic.

Trade recall against index cost

An embedding maps content into a numeric vector. Similarity may use cosine, dot product or Euclidean distance; normalization and the model’s training objective determine which comparison is meaningful. Vectors from different model versions are not automatically comparable. Store model identity with every vector and rebuild or dual-serve deliberately during migration.

Exact search compares the query against every candidate and is a useful correctness baseline. HNSW navigates a proximity graph, trading graph memory and search breadth for speed and recall. IVF first selects coarse clusters and searches a subset; probing more clusters usually improves recall at more work. Quantization reduces memory with approximation. LSH uses hashing designed for similar items, while tree-based approaches such as Annoy offer another read-oriented tradeoff. Choose by measured workload, not acronym count.

Filtering is central. Retrieving global top 20 then removing unauthorized records may return too few useful results and can leak through metadata. Prefer authorization-aware candidate retrieval and verify current permissions before disclosure. Hybrid retrieval combines lexical and vector candidates, then deduplicates and reranks a bounded set. Measure recall@K against exact results and task quality, alongside p95 latency and memory.

A decision worksheet for Vector Databases: read the mechanism and its guarantee together.
Scroll to inspect the diagram, or open it at full size.

Figure — A decision worksheet for Vector Databases: read the mechanism and its guarantee together.

Operational sketch

Contract / pseudocode
query -> embed with model v3
ACL/tenant scope -> candidate retrieval
lexical + vector candidates -> dedup -> rerank
current authorization -> return evidence
record index generation and source document version

A tempting mistake

High similarity is not factual correctness. A retrieved passage may be outdated, unauthorized or irrelevant to the exact question. A new embedding model can silently damage quality if only query vectors are upgraded.

Transfer exercise

What should you measure before reducing HNSW search breadth?

Show answer and explanation

Answer: Recall@K against an exact baseline on representative queries, including filtered and rare cases, plus latency and resource cost. Faster results are not useful if the required evidence disappears.

Candidate search budgets can exclude nearby vectors; compare recall against exact search.
Scroll to inspect the diagram, or open it at full size.

Figure — Candidate search budgets can exclude nearby vectors; compare recall against exact search.

8:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

Plan a model-version migration while maintaining deletes and access restrictions. Which evaluation gates must pass before routing queries to the new index?

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.