Elasticsearch
Search engines organize text and other fields for retrieval rather than ordinary transactional ownership.
Search engines organize text and other fields for retrieval rather than ordinary transactional ownership. Elasticsearch uses inverted-index concepts to find documents containing analyzed terms and supports structured and geographic queries. Treat the index as a projection unless you have explicitly designed its role as the source of truth. Search visibility and database commit are separate moments.
Learning goals
Inverted indexes; analyzers; mappings; shards; refresh lag; replicas; query/filter distinctions; reindexing; visibility controls; source of truth.
The mechanism at a glance
Figure — Authoritative document → Change stream (revision event); Change stream → Analyzer + mapping (upsert); Analyzer + mapping → Versioned index (tokens + fields); Versioned index → Search alias (validated version); Query → Search alias (search)
The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.
Step-by-step reasoning
1. Understand analysis and postings
An analyzer transforms text into tokens. An inverted index maps tokens to matching documents, enabling retrieval without scanning every full document. A field intended for exact filtering or sorting often needs a different mapping from a field intended for analyzed text. Test punctuation, language, and case behavior against actual queries.
2. Map and partition deliberately
Define field types and shard layout with data volume and query patterns in mind. Too many small shards create overhead; too few can limit distribution and recovery options. Replicas support availability and query capacity but do not remove every write or coordination bottleneck. Avoid uncontrolled dynamic fields that explode mappings.
3. Explain near-real-time visibility
A successful indexing operation does not mean every search immediately sees the document. Refresh makes changes searchable under the engine’s visibility model. Design user-facing read-after-write behavior accordingly: show the authoritative detail record or explicitly wait for a supported visibility boundary when necessary.
4. Evolve through a new index
Changing analysis or mapping often requires reindexing existing data. Build a new versioned index, backfill, catch up changes, validate counts and representative queries, and switch an alias or routing reference. Prevent old update events from overwriting newer document versions. Retain rollback capability until the new view is trusted.
Contracts and state
The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.
Source business v18 -> index event {id, version:18}
Mapping: name/text, category/keyword, location/geo_point
Index v2: backfill + live catch-up
Read alias: businesses_current -> businesses_v2
Source remains authority for ownership and visibility.Worked example
A user renames a cafe and immediately searches for the new name. The database detail page shows the new value while the search index still reflects the previous refresh. This can be expected within a stated freshness budget. If the edit confirmation must be immediate, read the source record; do not falsely promise that every search projection is synchronously updated.
Failure walkthrough
A reindex finishes its backfill while new edits continue. Switching before catch-up can expose stale data or resurrect removed documents. Track a change watermark, apply version checks, validate deletions, and switch only after the new index reaches the required boundary. Rebuilding a projection is a migration with concurrent writes, not merely a bulk copy.
Figure — Create index v2 → Backfill source records → Apply concurrent changes → Validate and catch up → Switch alias; retain rollback
Decisions and trade-offs
| Feature | Purpose | Trade-off |
|---|---|---|
| Analyzed text | Flexible term retrieval | Analysis affects matching |
| Exact field | Filters and sorting | Different representation |
| Refresh | Search visibility | Freshness vs resource cost |
| Versioned reindex | Safe schema evolution | Temporary duplicate storage |
Check your understanding
Why can a committed business update be absent from search for a short time?
Show answer and explanation
Answer: The search projection may not yet have received or refreshed the update. Measure pipeline lag and refresh visibility separately. The authoritative write success and search visibility are different contracts.
Transfer to a new scenario
Add searchable business descriptions and change the analyzer through a new index and alias swap.
Why can a committed document temporarily be absent from search results?
Primary documentation
Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.
Separate indexing freshness from source correctness
An inverted index maps analyzed terms to matching documents. The analyzer affects tokenization, case handling and stemming, so schema choices change search meaning. Exact identifiers usually need keyword-style treatment rather than free-text analysis. Store authoritative business state elsewhere unless the application deliberately adopts search storage as its source with suitable guarantees.
Indexing acknowledgement and search visibility are different events. Near-real-time refresh exposes changes to search; forcing refresh after every write can hurt throughput. Shards distribute indexing and queries, replicas add read capacity and resilience, and every scatter/gather request consumes resources across the selected shards.
Deep pagination using large offsets forces work on many discarded hits. Prefer a stable snapshot and search-after values when consistent paging matters. Reindexing into a new index and switching an alias is useful for incompatible mappings, but capture writes during backfill so cutover does not lose recent changes. Version documents so delayed updates cannot resurrect deleted data.
Figure — A decision worksheet for Elasticsearch: read the mechanism and its guarantee together.
Operational sketch
source transaction -> outbox -> versioned index update
query -> shard candidates -> merged ranking
current authorization -> hydrated response
new mapping -> new index -> catch-up -> alias switchA tempting mistake
Do not trust stale indexed permissions to authorize a sensitive result. Filtering after returning snippets is too late; authorization must precede disclosure.
Transfer exercise
What does a refresh solve, and what does it not solve?
Show answer and explanation
Answer: It makes indexed changes searchable. It does not repair a missed source event, enforce current access control or make a multi-system write atomic.