Papers › On the good reliability of an interval-based metric to validate prediction uncertainty...

On the good reliability of an interval-based metric to validate prediction uncertainty for machine learning regression tasks

23 Aug 2024arXiv:2408.13089archive 2025-07-28

Pascal Pernot

This short study presents an opportunistic approach to a (more) reliable validation method for prediction uncertainty average calibration. Considering that variance-based calibration metrics (ZMS, NLL, RCE...) are quite sensitive to the presence of heavy tails in the uncertainty and error distributions, a shift is proposed to an interval-based metric, the Prediction Interval Coverage Probability (PICP). It is shown on a large ensemble of molecular properties datasets that (1) sets of z-scores are well represented by Student's-t(ν) distributions, ν being the number of degrees of freedom; (2) accurate estimation of 95 % prediction intervals can be obtained by the simple 2σ rule for ν>3; and (3) the resulting PICPs are more quickly and reliably tested than variance-based calibration metrics. Overall, this method enables to test 20 % more datasets than ZMS testing. Conditional calibration is also assessed using the PICP approach.

PaperPDFCode

Code

ppernot/2024_picp officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

PredictionPrediction Intervals

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections