Papers › Rethinking Early Stopping: Refine, Then Calibrate

Rethinking Early Stopping: Refine, Then Calibrate

31 Jan 2025arXiv:2501.19195archive 2025-07-28

Eugène Berta, David Holzmüller, Michael I. Jordan, Francis Bach

Machine learning classifiers often produce probabilistic predictions that are critical for accurate and interpretable decision-making in various domains. The quality of these predictions is generally evaluated with proper losses like cross-entropy, which decompose into two components: calibration error assesses general under/overconfidence, while refinement error measures the ability to distinguish different classes. In this paper, we provide theoretical and empirical evidence that these two errors are not minimized simultaneously during training. Selecting the best training epoch based on validation loss thus leads to a compromise point that is suboptimal for both calibration error and, most importantly, refinement error. To address this, we introduce a new metric for early stopping and hyperparameter tuning that makes it possible to minimize refinement error during training. The calibration error is minimized after training, using standard techniques. Our method integrates seamlessly with any architecture and consistently improves performance across diverse classification tasks.

PaperPDFCode

Code

dholzmueller/pytabkit officialmentioned in papermentioned on GitHubpytorch report
eugeneberta/refinethencalibrate-theory officialmentioned in papermentioned on GitHub report
eugeneberta/refinethencalibrate-vision officialmentioned in papermentioned on GitHubpytorch report
dholzmueller/probmetrics officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Decision Making

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Early Stopping

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections