Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 28 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 1297–1344 of 12,172
The Florence 3D faces dataset consists of: High-resolution 3D scans of human faces from many subjects.
40 papers · 1 benchmark
LegalBench is a fascinating project that revolves around legal reasoning and evaluation.
40 papers · 0 benchmarks
Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection (MIMII) is a sound dataset of industrial machine sounds.
40 papers · 0 benchmarks
ManiSkill2 is the next generation of the SAPIEN ManiSkill benchmark, to address critical pain points often encountered by researchers when using benchmarks for generalizable manipulation skills.
40 papers · 0 benchmarks
RedCaps is a large-scale dataset of 12M image-text pairs collected from Reddit.
40 papers · 0 benchmarks
Samanantar is the largest publicly available parallel corpora collection for Indic languages: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu.
40 papers · 0 benchmarks
SciCite is a dataset of citation intents that addresses multiple scientific domains and is more than five times larger than ACL-ARC.
40 papers · 3 benchmarks
We propose the first question-answering dataset driven by STEM theorems.
40 papers · 1 benchmark
Veri-Wild is the largest vehicle re-identification dataset (as of CVPR 2019).
40 papers · 3 benchmarks
Visual Wake Words represents a common microcontroller vision use-case of identifying whether a person is present in the image or not, and provides a realistic benchmark for tiny vision models.
40 papers · 1 benchmark
ALCE (Automatic LLMs' Citation Evaluation)
ALCE is a benchmark for Automatic LLMs' Citation Evaluation.
39 papers · 0 benchmarks
AP-10K is the first large-scale benchmark for general animal pose estimation, to facilitate the research in animal pose estimation.
39 papers · 2 benchmarks
The Argoverse 2 Motion Forecasting Dataset is a curated collection of 250,000 scenarios for training and validation.
39 papers · 0 benchmarks
Break is a question understanding dataset, aimed at training models to reason over complex questions.
39 papers · 0 benchmarks
BUFF (Bodies Under Flowing Fashion)
BUFF consists of 5 subjects, 3 male and 2 female wearing 2 clothing styles: a) t-shirt and long pants and b) a soccer outfit.
39 papers · 1 benchmark
BookSum is a collection of datasets for long-form narrative summarization.
39 papers · 2 benchmarks
The Campus and Shelf datasets were presented in the paper 3D Pictorial Structures for Multiple Human Pose Estimation.
39 papers · 2 benchmarks
We present CulturaX, a substantial multilingual dataset with 6.3 trillion tokens in 167 languages, tailored for large language model (LLM) development.
39 papers · 0 benchmarks
DIS5K (Dichotomous Image Segmentation (DIS) Dataset)
To build the highly accurate Dichotomous Image Segmentation dataset (DIS5K), we first manually collected over 12,000 images from Flickr1 based on our pre-designed keywords.
39 papers · 5 benchmarks
A dataset of financial agreements made public through U.S.
39 papers · 0 benchmarks
The largest and cleanest face recognition dataset Glint360K, which contains 17,091,657 images of 360,232 individuals, baseline models trained on Glint360K can easily achieve state-of-the-art performance.
39 papers · 0 benchmarks
H3D (Honda Research Institute 3D)
The H3D is a large scale full-surround 3D multi-object detection and tracking dataset.
39 papers · 0 benchmarks
The MTG-Jamendo dataset is an open dataset for music auto-tagging.
39 papers · 0 benchmarks
The Memetracker corpus contains articles from mainstream media and blogs from August 1 to October 31, 2008 with about 1 million documents per day.
39 papers · 1 benchmark
RoadTracer is a dataset for extraction of road networks from aerial images.
39 papers · 0 benchmarks
The SIXray dataset is constructed by the Pattern Recognition and Intelligent System Development Laboratory, University of Chinese Academy of Sciences.
39 papers · 1 benchmark
TDIUC (Task Directed Image Understanding Challenge)
Task Directed Image Understanding Challenge (TDIUC) dataset is a Visual Question Answering dataset which consists of 1.6M questions and 170K images sourced from MS COCO and the Visual Genome Dataset.
39 papers · 1 benchmark
Univesity is an indoor localization dataset proposed in "Camera Relocalization by Computing Pairwise Relative Poses Using Convolutional Neural Network".
39 papers · 0 benchmarks
VerSe (Large Scale Vertebrae Segmentation Challenge)
Spine or vertebral segmentation is a crucial step in all applications regarding automated quantification of spinal morphology and pathology.
39 papers · 0 benchmarks
WikiMovies is a dataset for question answering for movies content.
39 papers · 0 benchmarks
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus which could be used to train multi-speaker Text-to-Speech (TTS) systems.
38 papers · 0 benchmarks
BIOSSES (Biomedical Semantic Similarity Estimation System)
The BIOSSES data set comprises total 100 sentence pairs all of which were selected from the "TAC2 Biomedical Summarization Track Training Data Set" .
38 papers · 2 benchmarks
BigCodeBench is an easy-to-use benchmark for code generation with practical and challenging programming tasks¹.
38 papers · 2 benchmarks
A new publicly available dataset for verification of climate change-related claims.
38 papers · 1 benchmark
CUAD (Contract Understanding Atticus Dataset)
Contract Understanding Atticus Dataset (CUAD) is a dataset for legal contract review.
38 papers · 0 benchmarks
CVEfixes is a comprehensive vulnerability dataset that is automatically collected and curated from Common Vulnerabilities and Exposures (CVE) records in the public U.S.
38 papers · 0 benchmarks
ChID (Chinese IDiom dataset)
ChID is a large-scale Chinese IDiom dataset for cloze test.
38 papers · 0 benchmarks
ConvFinQA (Conversational Finance Question Answering)
ConvFinQA is a dataset designed to study the chain of numerical reasoning in conversational question answering.
38 papers · 2 benchmarks
DAIR-V2X is a large-scale, multi-modality, multi-view dataset from real scenarios for VICAD.
38 papers · 2 benchmarks
DeepFix consists of a program repair dataset (fix compiler errors in C programs).
38 papers · 1 benchmark
The ECE dataset (Gui et al., 2016a) is collected from SINA city news and contains 2105 instances.
38 papers · 1 benchmark
The EMOTIC dataset, named after EMOTions In Context, is a database of images with people in real environments, annotated with their apparent emotions.
38 papers · 2 benchmarks
A more challenging task to investigate two aspects of few-shot relation classification models: (1) Can they adapt to a new domain with only a handful of instances?
38 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
38 papers · 0 benchmarks
GazeFollow is a large-scale dataset annotated with the location of where people in images are looking.
38 papers · 1 benchmark
HDD (Honda Research Institute Driving Dataset)
Honda Research Institute Driving Dataset (HDD) is a dataset to enable research on learning driver behavior in real-life environments.
38 papers · 0 benchmarks
Imagenette is a subset of 10 easily classified classes from Imagenet (bench, English springer, cassette player, chain saw, church, French horn, garbage truck, gas pump, golf ball, parachute).
38 papers · 1 benchmark
InsuranceQA is a question answering dataset for the insurance domain, the data stemming from the website Insurance Library.
38 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.