Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 31 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 1441–1488 of 12,172
For the details of the work, the readers are refer to the paper "Feature Pyramid and Hierarchical Boosting Network for Pavement Crack Detection" (FPHB), T-ITS 2019.
34 papers · 0 benchmarks
CSL is a synthetic dataset introduced in Murphy et al.
34 papers · 2 benchmarks
The Ciao dataset contains rating information of users given to items, and also contain item category information.
34 papers · 1 benchmark
We introduce an object detection dataset in challenging adverse weather conditions covering 12000 samples in real-world driving scenes and 1500 samples in controlled weather conditions within a fog chamber.
34 papers · 2 benchmarks
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
The DIHARD II development and evaluation sets draw from a diverse set of sources exhibiting wide variation in recording equipment, recording environment, ambient noise, number of speakers, and speaker demographics.
34 papers · 1 benchmark
DR(eye)VE is a large dataset of driving scenes for which eye-tracking annotations are available.
34 papers · 0 benchmarks
This dataset focus on two blur types: camera motion blur and defocus blur.
34 papers · 0 benchmarks
A benchmark dataset that contains 500K document pages with fine-grained token-level annotations for document layout analysis.
34 papers · 0 benchmarks
ECSSD (Extended Complex Scene Saliency Dataset)
The Extended Complex Scene Saliency Dataset (ECSSD) is comprised of complex scenes, presenting textures and structures common to real-world images.
34 papers · 5 benchmarks
The EgoHands dataset contains 48 Google Glass videos of complex, first-person interactions between two people.
34 papers · 0 benchmarks
GrailQA (Strongly Generalizable Question Answering)
GrailQA is a new large-scale, high-quality dataset for question answering on knowledge bases (KBQA) on Freebase with 64,331 questions annotated with both answers and corresponding logical forms in different syntax (i.e., SPARQL,…
34 papers · 4 benchmarks
IBims-1 (Independent benchmark images and matched scans v1)
iBims-1 (independent Benchmark images and matched scans - version 1) is a new high-quality RGB-D dataset, especially designed for testing single-image depth estimation (SIDE) methods.
34 papers · 2 benchmarks
The LM (Linemod) dataset is a valuable resource introduced by Stefan Hinterstoisser and colleagues in their research on model-based training, detection, and pose estimation of texture-less 3D objects in heavily cluttered scenes¹.
34 papers · 5 benchmarks
MOT20 is a dataset for multiple object tracking.
34 papers · 2 benchmarks
OGB-LSC (OGB Large-Scale Challenge)
OGB Large-Scale Challenge (OGB-LSC) is a collection of three real-world datasets for advancing the state-of-the-art in large-scale graph ML.
34 papers · 3 benchmarks
OxUva is a dataset and benchmark for evaluating single-object tracking algorithms.
34 papers · 0 benchmarks
PEMS-BAY is a dataset for traffic prediction.
34 papers · 2 benchmarks
PolyU Dataset is a large dataset of real-world noisy images with reasonably obtained corresponding “ground truth” images.
34 papers · 1 benchmark
The ProPara dataset is designed to train and test comprehension of simple paragraphs describing processes (e.g., photosynthesis), designed for the task of predicting, tracking, and answering questions about how entities change during the…
34 papers · 0 benchmarks
Re-DocRED (Revisiting Document Level Relation Extraction)
The Re-DocRED Dataset resolved the following problems of DocRED: 1.
34 papers · 3 benchmarks
A dataset consisting of 180,662 triplets of dual-pol synthetic aperture radar (SAR) image patches, multi-spectral Sentinel-2 image patches, and MODIS land cover maps.
34 papers · 0 benchmarks
SUIM (Segmentation of Underwater IMagery)
The Segmentation of Underwater IMagery (SUIM) dataset contains over 1500 images with pixel annotations for eight object categories: fish (vertebrates), reefs (invertebrates), aquatic plants, wrecks/ruins, human divers, robots, and…
34 papers · 2 benchmarks
THCHS-30 is a free Chinese speech database THCHS-30 that can be used to build a full-fledged Chinese speech recognition system.
34 papers · 0 benchmarks
The TUD-L (TUD Light) dataset is part of the Benchmark for 6D Object Pose Estimation (BOP).
34 papers · 0 benchmarks
A new multimodal retrieval dataset.
34 papers · 2 benchmarks
ToxicChat is a novel benchmark dataset constructed based on real user queries from an open-source chatbot.
34 papers · 0 benchmarks
UMDFaces is a face dataset divided into two parts: Still Images - 367,888 face annotations for 8,277 subjects.
34 papers · 0 benchmarks
minesweeper is a synthetic graph emulating the eponymous game.
34 papers · 1 benchmark
xP3 is a multilingual dataset for multitask prompted finetuning.
34 papers · 0 benchmarks
2000 HUB5 English Evaluation Transcripts was developed by the Linguistic Data Consortium (LDC) and consists of transcripts of 40 English telephone conversations used in the 2000 HUB5 evaluation sponsored by NIST (National Institute of…
33 papers · 2 benchmarks
3dshapes is a dataset of 3D shapes procedurally generated from 6 ground truth independent latent factors.
33 papers · 0 benchmarks
AMIGOS (AMIGOS: A Dataset for Affect, Personality and Mood Research on Individuals and Groups)
We present a database for research on affect, personality traits and mood by means of neuro-physiological signals.
33 papers · 1 benchmark
CCAligned consists of parallel or comparable web-document pairs in 137 languages aligned with English.
33 papers · 0 benchmarks
CIHP (Crowd Instance-level Human Parsing)
The Crowd Instance-level Human Parsing (CIHP) dataset has 38,280 diverse human images.
33 papers · 1 benchmark
COCO-WholeBody is an extension of COCO dataset with whole-body annotations.
33 papers · 5 benchmarks
Contains 68,536 activity instances in 68.8 hours of first and third-person video, making it one of the largest and most diverse egocentric datasets available.
33 papers · 1 benchmark
ClevrTex is a new benchmark designed as the next challenge to compare, evaluate and analyze algorithms for unsupervised multi-object segmentation.
33 papers · 1 benchmark
Cosal2015 is a large-scale dataset for co-saliency detection which consists of 2,015 images of 50 categories, and each group suffers from various challenging factors such as complex environments, occlusion issues, target appearance…
33 papers · 1 benchmark
The Completion3D benchmark is a dataset for evaluating state-of-the-art 3D Object Point Cloud Completion methods.
33 papers · 1 benchmark
Bio-decagon is a dataset for polypharmacy side effect identification problem framed as a multirelational link prediction problem in a two-layer multimodal graph/network of two node types: drugs and proteins.
33 papers · 1 benchmark
The Dialog State Tracking Challenges 2 & 3 (DSTC2&3) were research challenge focused on improving the state of the art in tracking the state of spoken dialog systems.
33 papers · 5 benchmarks
EmailEU is a directed temporal network constructed from email exchanges in a large European research institution for a 803-day period.
33 papers · 0 benchmarks
GLUCOSE is a large-scale dataset of implicit commonsense causal knowledge, encoded as causal mini-theories about the world, each grounded in a narrative context.
33 papers · 0 benchmarks
JRDB (JackRabbot Dataset and Benchmark)
A novel egocentric dataset collected from social mobile manipulator JackRabbot.
33 papers · 1 benchmark
The JSB chorales are a set of short, four-voice pieces of music well-noted for their stylistic homogeneity.
33 papers · 1 benchmark
A new multitask action quality assessment (AQA) dataset, the largest to date, comprising of more than 1600 diving samples; contains detailed annotations for fine-grained action recognition, commentary generation, and estimating the AQA…
33 papers · 2 benchmarks
MVP (Multi-View Partial point cloud)
MVP is a multi-view partial point cloud dataset (MVP) containing over 100,000 high-quality scans, which renders partial 3D shapes from 26 uniformly distributed camera poses for each 3D CAD model.
33 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.