Home › Datasets › task › Anomaly Detection

Anomaly Detection datasets

archive 2025-07-28

119 datasets carry the task tag "Anomaly Detection" (the task itself: Anomaly Detection), ordered by the archive's paper count. Page 2 of 3: 48 shown of 119. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Anomaly Detection datasets 49–96 of 119

SMD (Server Machine Dataset)
a dataset of time-series anomaly detection
10 papers · 3 benchmarks
UEA time-series datasets (UEA time-series datasets for series-level anomaly detection)
Five datasets used in NeurTraL-AD paper: \textit{RacketSports (RS).} Accelerometer and gyroscope recording of players playing four different racket sports.
10 papers · 1 benchmark
DAD (Driver Anomaly Detection)
Contains normal driving videos together with a set of anomalous actions in its training set.
9 papers · 0 benchmarks
Contains 4,677 videos with temporal, spatial, and categorical annotations.
9 papers · 0 benchmarks
The Standardized Project Gutenberg Corpus (SPGC) is an open science approach to a curated version of the complete PG data containing more than 50,000 books and more than 3×109 word-tokens.
9 papers · 0 benchmarks
Amazon-Fraud (Multi-relational Graph Dataset for Amazon Fraudulent Account Detection)
Amazon-Fraud is a multi-relational graph dataset built upon the Amazon review dataset, which can be used in evaluating graph-based node classification, fraud detection, and anomaly detection models.
8 papers · 3 benchmarks
The dataset contains transactions made by credit cards in September 2013 by European cardholders.
8 papers · 2 benchmarks
CHAD (Charlotte Anomaly Dataset)
CHAD: Charlotte Anomaly Dataset CHAD is high-resolution, multi-camera dataset for surveillance video anomaly detection.
7 papers · 1 benchmark
UBI-Fights (Abnormal Event Detection Dataset)
UBI-Fights - Concerning a specific anomaly detection and still providing a wide diversity in fighting scenarios, the UBI-Fights dataset is a unique new large-scale dataset of 80 hours of video fully annotated at the frame level.
7 papers · 2 benchmarks
The 3DSeg-8 is a collection of several publicly available 3D segmentation datasets from different medical imaging modalities, e.g.
6 papers · 0 benchmarks
AnoShift (AnoShift: A Distribution Shift Benchmark for Unsupervised Anomaly Detection)
AnoShift is a large-scale anomaly detection benchmark, which focuses on splitting the test data based on its temporal distance to the training set, introducing three testing splits: IID, NEAR, and FAR.
6 papers · 1 benchmark
UTSD (Unified Time Series Dataset)
Unified Time Series Dataset (UTSD) includes 7 domains with up to 1 billion time points with hierarchical capacities to facilitate research of large models in the field of time series.
6 papers · 0 benchmarks
An abnormal activity data-set for research use that contains 4,83,566 annotated frames.
5 papers · 2 benchmarks
InsPLAD (Inspection Power Line Asset Dataset)
InsPLAD is a Dataset for Power Line Asset Inspection containing 10,607 high-resolution Unmanned Aerial Vehicles colour images.
5 papers · 1 benchmark
This is the data set used for The Third International Knowledge Discovery and Data Mining Tools Competition, which was held in conjunction with KDD-99 The Fifth International Conference on Knowledge Discovery and Data Mining.
5 papers · 1 benchmark
OoDIS (Anomaly Instance Segmentation Benchmark)
OoDIS is a benchmark dataset for anomaly instance segmentation, crucial for autonomous vehicle safety.
5 papers · 2 benchmarks
Thyroid (Thyroid Disease)
Thyroid is a dataset for detection of thyroid diseases, in which patients diagnosed with hypothyroid or subnormal are anomalies against normal patients.
5 papers · 1 benchmark
ToyADMOS2 is a dataset of miniature-machine operating sounds for anomalous sound detection under domain shift conditions.
5 papers · 0 benchmarks
voraus-AD contains machine data of a collaborative robot, which moves a can by performing an industrial pick-and-place task.
5 papers · 1 benchmark
This dataset contains five notable histological artifacts: blur, blood (hemorrhage), air bubbles, folded tissue, and damaged tissue.
4 papers · 1 benchmark
ImageNet-1k vs NINCO (No ImageNet Class Objects)
The NINCO (No ImageNet Class Objects) dataset is introduced in the ICML 2023 paper In or Out?
4 papers · 1 benchmark
MUAD (Multiple Uncertainties for Autonomous Driving)
The MUAD dataset (Multiple Uncertainties for Autonomous Driving), consisting of 10,413 realistic synthetic images with diverse adverse weather conditions (night, fog, rain, snow), out-of-distribution objects, and annotations for semantic…
4 papers · 0 benchmarks
The Musk dataset describes a set of molecules, and the objective is to detect musks from non-musks.
4 papers · 2 benchmarks
SWAT A7 (Secure Water Treatment (SWaT))
11 days of continuous operation: 7 under normal operation and 4 days with attack scenarios: + Collected network traffic & all the values obtained from all the 51 sensors and actuators + Data labelled according to normal and abnormal…
4 papers · 0 benchmarks
AnoVox is a large-scale benchmark for ANOmaly detection in autonomous driving.
3 papers · 0 benchmarks
Hazards&Robots (Hazards&Robots: A Dataset for Visual Anomaly Detection in Robotics)
We consider the problem of detecting, in the visual sensing data stream of an autonomous mobile robot, semantic patterns that are unusual (i.e., anomalous) with respect to the robot’s previous experience in similar environments.
3 papers · 0 benchmarks
SKAB (Skoltech Anomaly Benchmark)
SKAB is designed for evaluating algorithms for anomaly detection.
3 papers · 2 benchmarks
SensumSODF (Sensum Solid Oral Dosage Forms)
Given the unavailability of real-world pharmaceutical inspection-domain datasets, we have created the Sensum Solid Oral Dosage Forms (SensumSODF) dataset intended for research and evaluation purposes.
3 papers · 0 benchmarks
A new dataset for streaming classification consisting of temporally correlated images from 51 distinct object categories and additional evaluation classes outside of the training distribution to test novelty recognition.
3 papers · 0 benchmarks
SupplyGraph (SupplyGraph: A Benchmark Dataset for Supply Chain Planning using Graph Neural Networks)
Graph Neural Networks (GNNs) have gained traction across different domains such as transportation, bio-informatics, language processing, and computer vision.
3 papers · 0 benchmarks
WFDD (Woven Fabric Defect Detection)
WFDD is a dataset for benchmarking anomaly detection methods with a focus on textile inspection.
3 papers · 1 benchmark
Yahoo S5 (Yahoo S5 - A Labeled Anomaly Detection Dataset)
Automatic anomaly detection is critical in today's world where the sheer volume of data makes it impossible to tag outliers manually.
3 papers · 0 benchmarks
COMPASS-XP is a dataset of matched photographic and X-ray images of single objects, made available for use in Machine Learning & Computer Vision research, in particular in the context of transport security.
2 papers · 0 benchmarks
This failure dataset contains information on the events collected in the OpenStack cloud computing platform during three different campaigns of fault-injection experiments performed with three different workloads.
2 papers · 0 benchmarks
HERA RFI Detection (Hydrogen Epoch of Reionization Array (HERA))
This dataset contains simulated and expert-labelled spectrograms from two radio telescopes: the Hydrogen Epoch of Reionization Array (HERA) in South Africa and the Low-Frequency Array (LOFAR) in the Netherlands.
2 papers · 1 benchmark
Bearing acceleration data from three run-to-failure experiments on a loaded shaft.
2 papers · 0 benchmarks
Icons-50 is a dataset for studying surface variation robustness.
2 papers · 0 benchmarks
The Insider Threat Test Dataset is a collection of synthetic insider threat test datasets that provide both background and malicious actor synthetic data.
2 papers · 1 benchmark
LOFAR RFI Detection (Low-Frequency Array (LOFAR) Radio Frequency Interference Detection)
This dataset contains simulated and expert-labelled spectrograms from two radio telescopes: the Hydrogen Epoch of Reionization Array (HERA) in South Africa and the Low-Frequency Array (LOFAR) in the Netherlands.
2 papers · 1 benchmark
Large-scale Anomaly Detection (LAD) is a database to benchmark anomaly detection in video sequences, which is featured in two aspects.
2 papers · 0 benchmarks
MIAD contains more than 100K high-resolution color images in various outdoor industrial scenarios, designed for unsupervised anomaly detection.
2 papers · 0 benchmarks
Scene-focused, multi-modal, episodic data of the images and symbolic world-states seen by an agent completing a pogo-stick assembly task within a video game world.
2 papers · 0 benchmarks
PAD Dataset (Pose-agnostic/Multi-pose Anomaly Detection Dataset)
Multi-pose Anomaly Detection (MAD) dataset, which represents the first attempt to evaluate the performance of pose-agnostic anomaly detection.
2 papers · 1 benchmark
TEP (Tennessee Eastman Process)
The original paper presented a model of the industrial chemical process named Tennessee Eastman Process and a model-based TEP simulator for data generation.
2 papers · 1 benchmark
The TII-SSRC-23 dataset offers a comprehensive collection of network traffic patterns, meticulously compiled to support the development and research of Intrusion Detection Systems (IDS).
2 papers · 3 benchmarks
The code to create the dataset is available here.
2 papers · 2 benchmarks
edeniss2020 (EDEN ISS 2020 Telemetry Dataset)
Overview The edeniss2020 dataset is a time series dataset.
2 papers · 0 benchmarks
5GAD-2022 (5G attack detection dataset)
This dataset contains two types of intercepted network packets: "normal" network traffic packets (i.e.
1 paper · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.