Match the metric to the decision

Discrimination measures how well a model separates outcomes; calibration concerns agreement between predicted probabilities and observed outcomes. If users act on a risk threshold, probability quality matters. Evaluate both on data appropriate to the intended use.

Look beyond an overall score

Review calibration across a sensible range of predictions and relevant subgroups, with attention to sample size and uncertainty. Avoid making strong claims from sparse bins. If recalibration is needed, perform it using a defined process and reassess on separate data.

Sources & further reading

Sources checked 20 September 2026. Implementation commentary reflects a practical review perspective; source material may change.

Have a related programming or implementation challenge?

Discuss it with Sai ↗