Match the metric to the decision
Discrimination measures how well a model separates outcomes; calibration concerns agreement between predicted probabilities and observed outcomes. If users act on a risk threshold, probability quality matters. Evaluate both on data appropriate to the intended use.
Look beyond an overall score
Review calibration across a sensible range of predictions and relevant subgroups, with attention to sample size and uncertainty. Avoid making strong claims from sparse bins. If recalibration is needed, perform it using a defined process and reassess on separate data.
Sources & further reading
Sources checked 20 September 2026. Implementation commentary reflects a practical review perspective; source material may change.
Have a related programming or implementation challenge?
Discuss it with Sai ↗