Learning pathsA
GUIDED PRACTICE

Time Series Databases

A time-series system stores measurements keyed by time and one or more dimensions.

A time-series system stores measurements keyed by time and one or more dimensions. Design around bounded time-range queries, ingestion rate, cardinality, retention, and correction needs. “Millions of points per second” is not enough to pick a store: an unbounded set of unique label combinations can exhaust memory even when byte volume looks manageable.

The mechanism at a glance

Metric producer → Series index (validate bounded labels); Series index → Time chunks (append by shard + time); Time chunks → Rollup compactor (compact older range); Rollup compactor → Query planner (serve coarse history); Time chunks → Query planner (serve recent detail); Rollup compactor → Retention manager (expire after verification)
Scroll to inspect the diagram, or open it at full size.

Figure — Metric producer → Series index (validate bounded labels); Series index → Time chunks (append by shard + time); Time chunks → Rollup compactor (compact older range); Rollup compactor → Query planner (serve coarse history); Time chunks → Query planner (serve recent detail); Rollup compactor → Retention manager (expire after verification)

The numbered components identify responsibilities. Follow the labeled arrows rather than treating the numbers as a global execution order. The scenario later in this lesson shows one concrete sequence.

Step-by-step reasoning

1. 1 · Frame

Clarify sample shape, timestamp source, write rate, query windows, resolution, retention, and late corrections. Separate operational dashboards from forensic or billing records. Define whether a chart may be approximate and how quickly a new sample must appear. Estimate both bytes and active series: series cardinality often drives index and metadata costs.

2. 2 · Model

Use a logical key of measurement plus bounded dimensions, with timestamped values. Reject or normalize unbounded labels such as raw user IDs unless a query need justifies them. Assign event time and ingest time separately. Partition by time and a stable shard key; keep series metadata distinct from compressed sample blocks.

3. 3 · Scale

Append samples into time-ordered chunks, compress them, and use time and label indexes to prune scans. Retention can delete whole old blocks efficiently. Downsampling produces coarser rollups for long ranges, while raw data remains available for the promised interval. Avoid downsampling across reset or counter-wrap boundaries without reset-aware logic.

4. 4 · Recover

Late data may target a sealed block; choose a bounded correction horizon and a repair path. A compaction job should be idempotent and publish a new block version before deleting old inputs. Protect query paths from expensive high-cardinality selections and monitor active series, ingestion lag, dropped samples, and storage growth.

Contracts and state

The following sketch makes the decision boundary concrete. Field names and capacity assumptions are illustrative; adapt them to the stated product contract.

Contract / pseudocode
Series(key, dimensions, unit, schema_version)
Sample(series_id, event_time, ingest_time, value)
Chunk(series_id, time_range, encoding, checksum)
Rollup(series_id, interval, start_time, aggregate, revision)
Query = dimensions + bounded time range + resolution

Worked example

A dashboard asks for one-minute CPU averages over 30 days. The system reads daily rollups rather than scanning every raw point. An incident investigator asks for a five-minute interval from yesterday; raw samples remain available during the declared raw-retention period. Both queries carry the same bounded dimensions, so neither creates a scan over every tenant series.

Failure walkthrough

A client attaches a unique request ID as a metric label, creating a new series per request. Ingestion still succeeds initially, but series indexes and memory grow much faster than sample bytes. Enforce label budgets, drop or route unsuitable dimensions to logs/traces, and make rejected-label behavior observable. Never silently merge semantically distinct metrics.

Samples append into open chunk → Chunk is sealed by time → Rollup is computed with revision → Late sample revises bounded range → Old raw chunk expires by policy
Scroll to inspect the diagram, or open it at full size.

Figure — Samples append into open chunk → Chunk is sealed by time → Rollup is computed with revision → Late sample revises bounded range → Old raw chunk expires by policy

Decisions and trade-offs

ChoiceGood fitCost
Time partitioned chunksAppend-heavy bounded range readsLate corrections and compaction
RollupsLong-range chartsFine detail is irrecoverable after raw expiry
High-cardinality labelsOnly when queries require themIndex and memory growth
Separate event archiveAudits and rebuildsMore storage and delayed queries

Check your understanding

Design metrics retention for raw samples, hourly charts, and one late correction that arrives after a block was compacted.

Show answer and explanation

Answer: Keep raw data for a shorter explicit interval and hourly rollups for the longer chart horizon. Track event and ingest time; late corrections within a bounded horizon revise the relevant rollup version. Corrections outside that horizon require a replay/archive path or are rejected by contract. Publish compaction output atomically and retain source blocks until verification.

Primary documentation

Read the first-party engineering account or official technical reference. Company engineering posts describe the scope and date of that publication; the interview reconstruction and scenarios here are original teaching examples.

Continue the connection

Study Metrics Monitoring and explain which guarantee from this lesson carries into that topic.

Budget series cardinality and retention

A time series is identified by a metric and label set, then contains timestamped values. Millions of unique label combinations can cost more than a high sample rate for a small stable set. User IDs, request IDs and free-form URLs are dangerous labels when they grow without bound.

Time partitioning supports retention and range scans; compression benefits from regular timestamps and related values. Downsampling lowers long-term cost but changes available questions. An hourly average cannot reconstruct a one-minute spike, and averaging percentiles from separate instances does not generally produce the combined percentile. Preserve suitable histogram buckets or raw observations for the required analysis.

Separate ingestion freshness, query latency and alert delay. Late or out-of-order samples need a policy. Missing data is distinct from zero. Delete or compact whole expired blocks where possible rather than generating billions of individual delete operations. Replication and remote archive paths must have clear replay and deduplication behavior.

A decision worksheet for Time Series Databases: read the mechanism and its guarantee together.
Scroll to inspect the diagram, or open it at full size.

Figure — A decision worksheet for Time Series Databases: read the mechanism and its guarantee together.

Operational sketch

Contract / pseudocode
series = metric_name + normalized label set
sample = timestamp + numeric value
raw: 15-second resolution for 7 days
rollup: 5-minute buckets for 90 days
retention policy states which peaks are preserved

A tempting mistake

Downsampling before defining alert requirements can erase the evidence for incidents. A cheap dashboard query is not enough if the data model cannot answer the operational question.

Transfer exercise

Why is endpoint=/orders/{id} a safer label than the full request URL?

Show answer and explanation

Answer: The normalized route has bounded cardinality. Full URLs create a distinct series for many IDs and query strings, increasing index and memory cost.

8:00Self-guided practice timer
The timer resets when you leave this page. Save your design separately.
Your challenge

Design metrics retention for raw samples, hourly charts, and one late correction that arrives after a block was compacted.

Your design draft

Clarify assumptions, explain your approach, and test the difficult cases. Save your draft, then compare it with the study notes.

Read study notes

Self-review checklist

Self-guided practice. Automated AI feedback and code execution are not connected.