Home › Datasets › task › Unsupervised Anomaly Detection
Unsupervised Anomaly Detection datasets
archive 2025-07-28
27 datasets carry the task tag "Unsupervised Anomaly Detection" (the task itself: Unsupervised Anomaly Detection), ordered by the archive's paper count. Page 1 of 1: 27 shown of 27. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Unsupervised Anomaly Detection datasets 1–27 of 27
The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
Fashion-MNIST is a dataset comprising of 28×28 grayscale images of 70,000 fashion products from 10 categories, with 7,000 images per category.
3,202 papers · 15 benchmarks
STL-10 (Self-Taught Learning 10)
The STL-10 is an image dataset derived from ImageNet and popularly used to evaluate algorithms of unsupervised feature learning or self-taught learning.
1,092 papers · 18 benchmarks
The Caltech101 dataset contains images from 101 object categories (e.g., “helicopter”, “elephant” and “chair” etc.) and a background category that contains the images not from the 101 object categories.
709 papers · 10 benchmarks
MVTecAD (MVTEC ANOMALY DETECTION DATASET)
MVTec AD is a dataset for benchmarking anomaly detection methods with a focus on industrial inspection.
402 papers · 4 benchmarks
SMAP (Soil Moisture Active Passive)
Soil Moisture Active Passive (SMAP) dataset is a dataset of soil samples and telemetry information using the Mars rover by NASA.
127 papers · 2 benchmarks
The Reuters-21578 dataset is a collection of documents with news articles.
66 papers · 5 benchmarks
Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection (MIMII) is a sound dataset of industrial machine sounds.
40 papers · 0 benchmarks
MVTec Logical Constraints Anomaly Detection (MVTec LOCO AD) dataset is intended for the evaluation of unsupervised anomaly localization algorithms.
32 papers · 1 benchmark
The dataset is constructed from images of defective production items that were provided and annotated by Kolektor Group d.o.o..
15 papers · 1 benchmark
KolektorSDD2 is a surface-defect detection dataset with over 3000 images containing several types of defects, obtained while addressing a real-world industrial problem.
15 papers · 2 benchmarks
ToyADMOS dataset is a machine operating sounds dataset of approximately 540 hours of normal machine operating sounds and over 12,000 samples of anomalous sounds collected with four microphones at a 48kHz sampling rate, prepared by Yuma…
15 papers · 0 benchmarks
AeBAD (Aero-engine Blade Anomaly Detection Dataset)
Unlike previous datasets that focus on detecting the diversity of defect categories (like MVTec AD and VisA), AeBAD is centered on the diversity of domains within the same data category.
11 papers · 3 benchmarks
SMD (Server Machine Dataset)
a dataset of time-series anomaly detection
10 papers · 3 benchmarks
CHAD (Charlotte Anomaly Dataset)
CHAD: Charlotte Anomaly Dataset CHAD is high-resolution, multi-camera dataset for surveillance video anomaly detection.
7 papers · 1 benchmark
UBI-Fights - Concerning a specific anomaly detection and still providing a wide diversity in fighting scenarios, the UBI-Fights dataset is a unique new large-scale dataset of 80 hours of video fully annotated at the frame level.
7 papers · 2 benchmarks
AnoShift (AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection)
AnoShift is a large-scale anomaly detection benchmark, which focuses on splitting the test data based on its temporal distance to the training set, introducing three testing splits: IID, NEAR, and FAR.
6 papers · 1 benchmark
This is a synthetic dataset for defect detection on textured surfaces.
6 papers · 1 benchmark
The original dataset for "ECG5000" is a 20-hour long ECG downloaded from Physionet.
6 papers · 3 benchmarks
This failure dataset contains information on the events collected in the OpenStack cloud computing platform during three different campaigns of fault-injection experiments performed with three different workloads.
2 papers · 0 benchmarks
MIAD contains more than 100K high-resolution color images in various outdoor industrial scenarios, designed for unsupervised anomaly detection.
2 papers · 0 benchmarks
PAD Dataset (Pose-agnostic/Multi-pose Anomaly Detection Dataset)
Multi-pose Anomaly Detection (MAD) dataset, which represents the first attempt to evaluate the performance of pose-agnostic anomaly detection.
2 papers · 1 benchmark
The code to create the dataset is available here.
2 papers · 2 benchmarks
ISP-AD (The Industrial Screen Printing Anomaly Detection Dataset)
The ISP-AD Dataset is a large-scale anomaly detection dataset, representing a real-world industrial use case.
1 paper · 0 benchmarks
ITD (Industrial Textile Dataset)
This dataset aims to provide a color dataset with real industrial fabric defect gathered in a visiting machine with several industrial cameras.
1 paper · 0 benchmarks
PRONTO (PRONTO heterogeneous benchmark dataset)
The PRONTO heterogeneous benchmark dataset is based on an industrial-scale multiphase flow facility.
1 paper · 1 benchmark
SPOT-10 (Animal Pattern Benchmark Dataset for Machine Learning Algorithms)
The SPOTS-10 dataset is an extensive collection of grayscale images showcasing diverse patterns found in ten animal species.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.