Foundations
Evaluation and generalization
Measure the mistakes that matter.
Understand the problem
Use holdout data matching the expected serving distribution. Select precision, recall, ranking quality, or calibration based on the action. Slice metrics by relevant cohorts and monitor drift.
Make it concrete
For rare harmful content, accuracy can look excellent even if the model misses most harmful items.
Trade-offs and pitfalls
Offline benchmarks cannot fully predict feedback loops or user behavior.
Check your understanding
Explain when improving recall may create an unacceptable review workload.
Practice this topic