Browse State-of-the-Art › scoring rule
scoring rule
25 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
25 shown of 25 papers with code (90 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Oct 2019 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedNGBoost generalizes gradient boosting to probabilistic regression by treating the parameters of the conditional distribution as targets for a multiparameter boosting algorithm.
-
13 Jan 2021 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWe show that in the context of object detection, training variance networks with negative log likelihood (NLL) can lead to high entropy predictive distributions regardless of the correctness of the output mean.
-
28 Oct 2022 2 repositories listedYet calibration is not enough: even a perfectly calibrated classifier with the best possible accuracy can have confidence scores that are far from the true posterior probabilities.
-
26 Mar 2021 2 repositories listedWe consider frequently used scoring rules for right-censored survival regression models such as time-dependent concordance, survival-CRPS, integrated Brier score and integrated binomial log-likelihood, and prove that…
-
3 Aug 2020 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedSpeech synthesis is an important practical generative modeling problem that has seen great progress over the last few years, with likelihood-based autoregressive neural models now outperforming traditional concatenative…
-
22 Jun 2025 1 repository listedIn this study, we summarize and categorize previous works into three general strategies: intuitively designed methods, binning-based methods, and methods based on formulations of ideal calibration.
-
16 Jun 2025 1 repository listedWe demonstrate that state-of-the-art models do not issue calibrated forecasts for extreme wind speeds, and that the calibration of forecasts for extreme events can be improved by suitable adaptations to the loss…
-
25 Nov 2024 1 repository listedIn many large-scale classification problems, classes are organized in a known hierarchy, typically represented as a tree expressing the inclusion of classes in superclasses.
-
Superior Scoring Rules for Probabilistic Evaluation of Single-Label Multi-Class Classification Tasks25 Jul 2024 1 repository listedTraditional scoring rules like Brier Score and Logarithmic Loss sometimes assign better scores to misclassifications in comparison with correct classifications.
-
23 Jul 2024 1 repository listedBuilding upon the Bayesian approach introduced in \cite{c:07}, we devise a new method for online change point detection in the mean of a univariate time series, which is well suited for real-time applications and is…
-
22 Jul 2024 1 repository listedIn this work we aim to improve statistical post-processing models for probabilistic predictions of extreme wind speeds.
-
2 Jul 2024 1 repository listedAccurate precipitation forecasts have a high socio-economic value due to their role in decision-making in various fields such as transport networks and farming.
-
6 Jun 2024 1 repository listedEvaluating on 8 survival metrics, we assess discrimination, calibration, and overall predictive performance of the tested models.
-
29 May 2024 1 repository listed Syntology ran 5 of 6 samples · 1 unverified · 1 pointer-only (licence)Leveraging this strategy, we train language generation models using two classic strictly proper scoring rules, the Brier score and the Spherical score, as alternatives to the logarithmic score.
-
19 Apr 2023 1 repository listed Syntology ran 0 of 19 samples · 19 unverifiedMultivariate probabilistic time series forecasts are commonly evaluated via proper scoring rules, i.
-
10 Nov 2022 1 repository listedIn this paper we propose a new method for adjusting approximate posterior samples to reduce bias and produce more accurate uncertainty quantification.
-
29 Oct 2022 1 repository listedWhile being an effective framework of learning a shared model across multiple edge devices, federated learning (FL) is generally vulnerable to Byzantine attacks from adversarial edge devices.
-
26 Sep 2022 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedEnsemble weather forecasts based on multiple runs of numerical weather prediction models typically show systematic errors and require post-processing to obtain reliable forecasts.
-
31 May 2022 1 repository listedHowever, generative networks only allow sampling from the parametrized distribution; for this reason, Ramesh et al.
-
4 May 2022 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedWe propose a novel deep dynamics model, Probabilistic Equivariant Continuous COnvolution (PECCO) for probabilistic prediction of multi-agent trajectories.
-
15 Mar 2022 1 repository listedAccurate uncertainty estimates are essential for deploying deep object detectors in safety-critical systems.
-
3 Mar 2022 1 repository listedWe consider the problem of discovering causal relations from independence constraints selection bias in addition to confounding is present.
-
15 Dec 2021 1 repository listedAdversarial-free minimization is possible for some scoring rules; hence, our framework avoids the cumbersome hyperparameter tuning and uncertainty underestimation due to unstable adversarial training, thus unlocking…
-
20 Feb 2020 1 repository listedFirst, we want the learning algorithm to be no-regret with respect to the best fixed expert in hindsight.
-
5 Mar 2019 1 repository listedRobust scatter estimation is a fundamental task in statistics.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections