Papers › The Fishyscapes Benchmark: Measuring Blind Spots in Semantic Segmentation

The Fishyscapes Benchmark: Measuring Blind Spots in Semantic Segmentation

5 Apr 2019arXiv:1904.03215archive 2025-07-28

Hermann Blum, Paul-Edouard Sarlin, Juan Nieto, Roland Siegwart, Cesar Cadena

Deep learning has enabled impressive progress in the accuracy of semantic segmentation. Yet, the ability to estimate uncertainty and detect failure is key for safety-critical applications like autonomous driving. Existing uncertainty estimates have mostly been evaluated on simple tasks, and it is unclear whether these methods generalize to more complex scenarios. We present Fishyscapes, the first public benchmark for uncertainty estimation in a real-world task of semantic segmentation for urban driving. It evaluates pixel-wise uncertainty estimates towards the detection of anomalous objects in front of the vehicle. We~adapt state-of-the-art methods to recent semantic segmentation models and compare approaches based on softmax confidence, Bayesian learning, and embedding density. Our results show that anomaly detection is far from solved even for ordinary situations, while our benchmark allows measuring advancements beyond the state-of-the-art.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Anomaly DetectionAutonomous DrivingSegmentationSemantic Segmentation

Datasets

Introduced by this paper, per the archive.

Fishyscapes

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Anomaly Detection Fishyscapes L&F Dirichlet DeepLab AP 34.28 #13 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Dirichlet DeepLab FPR95 47.43 #13 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Void Classifier AP 10.29 #15 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Void Classifier FPR95 22.11 #15 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Bayesian DeepLab AP 9.8 #16 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Bayesian DeepLab FPR95 38.5 #16 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Learned Embedding Density AP 4.7 #17 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Learned Embedding Density FPR95 24.4 #17 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Softmax Entropy AP 2.9 #18 of 18 Archive leaderboard report
Anomaly Detection Fishyscapes L&F Softmax Entropy FPR95 44.8 #18 of 18 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Softmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections