Kafka
Kafka is a partitioned durable log.
Kafka is a partitioned durable log. Producers append records, and consumer groups track their progress through partition offsets. This model is useful when multiple systems need independent replay of the same events. Its guarantees depend on partitioning, acknowledgment, transaction, and consumer behavior; the word Kafka alone does not make every downstream effect exactly once.
Learning goals
Topics; partitions; keys; consumer groups; ordering scope; offsets; replay; retention; delivery semantics; external effects; skew.
The mechanism at a glance
Figure — Producer → Partition 0 (keyed append); Producer → Partition 1 (keyed append); Partition 0 → Group A workers (own offsets); Partition 1 → Group A workers (own offsets); Partition 0 → Group B rebuild (independent replay); Group A workers → External sink (deduped effect)
The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.
Step-by-step reasoning
1. Choose ordering scope
A topic contains partitions. Records within a partition have an order, while different partitions progress independently. A key commonly determines partition placement. Use a conversation ID when per-conversation order matters, but recognize that one extremely active conversation can become a hot partition.
2. Understand consumer groups
Within a group, a partition is assigned to one consumer at a time under ordinary group consumption. More consumers than partitions do not add parallelism for that topic. A separate group can rebuild analytics without advancing the production group’s offsets. Rebalances change ownership and require careful handling of in-flight work.
3. Place the offset boundary
Committing an offset before an external database update risks losing the effect after a crash. Updating first and committing afterward risks replaying the effect. An idempotent sink or a transactionally coordinated offset-and-effect design resolves this boundary. Kafka transactions can coordinate supported Kafka reads and writes, but arbitrary external APIs need their own protection.
4. Operate retention and replay
Offsets identify log positions, not permanent storage of every event. Retention limits how far back a consumer can replay. Monitor lag in both offsets and time, disk usage, skew, and failure recovery. Schema evolution must keep old retained records readable by the replaying consumer.
Contracts and state
The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.
Record: {event_id, campaign_id, occurred_at, schema_version}
Partition key: campaign_id (ordering vs hot-campaign trade-off)
Group A: production aggregates
Group B: rebuild-v2
Sink key: event_id or deterministic aggregate update identityWorked example
An analytics worker writes a click aggregate and crashes before committing its offset. The replacement reads the same record again. If it simply increments a counter, the click counts twice. If the sink records event_id atomically with the increment, replay can recognize the prior effect. Retaining only a short dedupe window is safe only when it covers the promised replay and redelivery horizon.
Failure walkthrough
A consumer is slow enough that records expire before it reads them. Restarting from its old offset cannot recover data no longer retained. Decide whether to restore from an archive, rebuild from another source, or report an unrecoverable gap. Set retention using recovery objectives and expected outage duration, not just disk convenience.
Figure — Read offset 42 → Commit effect to external DB → Crash before offset commit → Offset 42 delivered again → Sink deduplicates event identity
Decisions and trade-offs
| Concept | Meaning | Common confusion |
|---|---|---|
| Partition order | Order inside one log partition | Global order across topic |
| Consumer group | Independent progress and assignments | A separate copy of every record per worker |
| Offset commit | Recorded read progress | Atomic external side effect |
Check your understanding
Why can a second consumer group rebuild a view without disturbing the live group?
Show answer and explanation
Answer: Each group tracks its own offsets, so replay progress is independent while retained records are shared. Both groups still compete for broker and sink resources; rate-limit rebuilds and ensure retention covers the needed history.
Transfer to a new scenario
Replay a click stream to rebuild aggregates while preserving a separate production consumer group.
Why does committing an offset not make an external database update exactly once?
Primary documentation
Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.
Trace one event through log and consumer state
A topic partition is an ordered append log. Producers choose keys that determine which related events share order. A consumer group distributes partitions among consumers; ordering across different partitions is not one global sequence. Increasing consumer count beyond partition count does not create more active readers in that group for that topic.
Producer acknowledgement and replication settings determine when an append is considered accepted and what failures it can tolerate. Idempotent production addresses certain producer retry duplicates, while transactions can coordinate supported Kafka operations. Neither automatically makes an external email or database write atomic with an offset commit.
A consumer can commit an offset before processing, risking loss after a crash, or after processing, risking repetition. Protect the sink with idempotent writes or a transaction that couples the effect to a consumed identity. Retention is a finite replay window; consumer lag approaching retention is a data-loss risk for that consumer. Rebalancing requires safe ownership transfer and bounded processing times.
Figure — A decision worksheet for Kafka: read the mechanism and its guarantee together.
Operational sketch
partition key = order_id
consume record at offset 420
apply sink effect using event_id
commit safe progress after effect
on replay: same event_id produces no second effectA tempting mistake
A hot key can bottleneck one partition. Randomly salting it improves distribution only if the application no longer requires a single order or adds a later merge step.
Transfer exercise
Can “exactly-once Kafka” guarantee one SMS?
Show answer and explanation
Answer: No. The SMS provider is outside the Kafka transaction. Use provider idempotency where supported and reconciliation for uncertain outcomes.
Why can a second consumer group rebuild a view without disturbing the live group?
Your design draft
Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.
Self-review checklist
Self-guided practice. Automated AI feedback and code execution are not connected.