Learning pathsA
GUIDED PRACTICE

ZooKeeper

ZooKeeper coordinates small amounts of critical shared state using an ordered replicated service and client sessions.

ZooKeeper coordinates small amounts of critical shared state using an ordered replicated service and client sessions. It can support leader election and short-lived ownership signals, but it is not a general database or high-volume message bus. Session expiration means a client may have lost ownership even if its process has not noticed yet.

The mechanism at a glance

Contenders → Ephemeral sequential nodes (create contender); Ephemeral sequential nodes → Session/ensemble (session cleanup); Ephemeral sequential nodes → Watched predecessor (watch predecessor); Watched predecessor → Ephemeral sequential nodes (recheck); Session/ensemble → New owner epoch (elect); New owner epoch → Protected database (fenced write)
Scroll to inspect the diagram, or open it at full size.

Figure — Contenders → Ephemeral sequential nodes (create contender); Ephemeral sequential nodes → Session/ensemble (session cleanup); Ephemeral sequential nodes → Watched predecessor (watch predecessor); Watched predecessor → Ephemeral sequential nodes (recheck); Session/ensemble → New owner epoch (elect); New owner epoch → Protected database (fenced write)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. 1 · Frame

Define the coordination problem, required ordering, number of znodes, and tolerance for session loss. Use it for membership/leader coordination or a small lock namespace, not as the main store for every job. Create ephemeral nodes for session-bound presence and sequential nodes when ordered contenders are needed.

2. 2 · Model

A lock recipe creates an ephemeral sequential contender and watches only its immediate predecessor to avoid a thundering herd. The smallest sequence holder owns the lock. On disconnect/session expiry, its ephemeral node disappears. Watch notifications are hints to re-read state; they can be missed or coalesced between disconnect and reconnect.

3. 3 · Scale

Clients must handle connection loss during create: the server may have created a node while the response was lost. Use a client identity in node data and search for an existing contender before creating another. Keep znode payloads small. Monitor quorum health, request latency, outstanding watches, and session churn. Avoid long application work while assuming a session cannot expire.

4. 4 · Recover

A process pauses after its session expires and later resumes believing it is leader. Ephemeral node removal helps elect a new leader but cannot stop stale external writes. Use a fencing epoch at the protected destination or downstream lease validation. Coordination membership does not magically revoke capabilities already held outside ZooKeeper.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
/locks/job-42/lock-0000000012 (ephemeral sequential)
Client identity is stored in node data
Owner = lowest live sequence
Each waiter watches predecessor and rechecks children
Protected store validates current fencing epoch

Worked example

Clients A, B, and C create sequential contenders 12, 13, and 14. A owns; B watches 12; C watches 13. When A disappears, B wakes and rechecks the child list. C does not wake yet, reducing notification load. If B’s session expires, its node disappears and C can become owner after rechecking.

Failure walkthrough

A paused A loses its session, B is elected, then A resumes and writes to a database. ZooKeeper cannot prevent that external write on its own. B receives a newer fencing epoch; the database rejects A’s stale epoch. If the destination cannot validate ownership, use another coordination protocol appropriate to the side effect.

A holds epoch 7 → A session expires → B becomes owner with epoch 8 → A resumes and writes → Destination rejects epoch 7
Scroll to inspect the diagram, or open it at full size.

Figure — A holds epoch 7 → A session expires → B becomes owner with epoch 8 → A resumes and writes → Destination rejects epoch 7

Decisions and trade-offs

FeatureUseful forCaution
Ephemeral nodeSession-bound membershipSession expiry can surprise paused clients
Sequential nodeOrdering contendersCreate response can be ambiguous
WatchWake up to recheck stateNot a durable event queue
Fencing epochReject stale owner effectsDestination must enforce it

Check your understanding

Why must an elected leader carry a fencing token to the database it protects?

Show answer and explanation

Answer: A former leader can pause, lose its ZooKeeper session, then resume after a new leader is elected. The database must reject writes bearing an older epoch. ZooKeeper can order ownership decisions, but the protected destination enforces that stale authority is no longer valid.

Primary documentation

Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.

Continue the connection

Study Distributed Cache and explain which guarantee from this lesson carries into that topic.

Use coordination state sparingly

Coordination services are useful for membership, configuration and small ownership records. They are not a natural home for high-volume event payloads. A client session can own ephemeral state, but a network interruption and session expiry are different events; applications must handle ambiguous connectivity carefully.

A leader-election recipe can order contenders and observe a predecessor. Winning election grants a role in the coordination system, but the protected database must still reject writes from a former leader. Use an increasing epoch or fencing token accepted by the downstream resource. Without that check, a paused old leader can resume and corrupt state.

Watches notify that state may have changed; clients should reread authoritative state and reestablish watches according to the API semantics. Treat notifications as prompts to reconcile, not an unlimited durable event stream. Keep coordination payloads small and monitor session churn, request latency and ensemble health.

A decision worksheet for ZooKeeper: read the mechanism and its guarantee together.
Scroll to inspect the diagram, or open it at full size.

Figure — A decision worksheet for ZooKeeper: read the mechanism and its guarantee together.

Operational sketch

Contract / pseudocode
contender creates ephemeral sequential node
smallest sequence becomes elected owner
owner obtains monotonic fencing epoch
resource rejects stale epoch
watch event -> reread state -> restore watch

A tempting mistake

Availability of the coordination cluster does not imply availability of the protected data service. Likewise, a lease or ephemeral node cannot physically stop code already running elsewhere.

Transfer exercise

What happens if an old leader reconnects after another leader was elected?

Show answer and explanation

Answer: It must relinquish leadership and its old epoch must be rejected by the resource. Local belief or a cached election result cannot authorize a write.

7:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

Why must an elected leader carry a fencing token to the database it protects?

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.