Learning pathsA
Foundations

Evaluation and generalization

Measure the mistakes that matter.

Understand the problem

Use holdout data matching the expected serving distribution. Select precision, recall, ranking quality, or calibration based on the action. Slice metrics by relevant cohorts and monitor drift.

Make it concrete

For rare harmful content, accuracy can look excellent even if the model misses most harmful items.

Trade-offs and pitfalls

Offline benchmarks cannot fully predict feedback loops or user behavior.

Check your understanding

Explain when improving recall may create an unacceptable review workload.

Practice this topic

Your study notes