Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 36 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 1681–1728 of 12,172
WikiCoref is an English corpus annotated for anaphoric relations, where all documents are from the English version of Wikipedia.
28 papers · 1 benchmark
The WoZ 2.0 dataset is a newer dialogue state tracking dataset whose evaluation is detached from the noisy output of speech recognition systems.
28 papers · 1 benchmark
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
decaNLP (Natural Language Decathlon Benchmark)
Natural Language Decathlon Benchmark (decaNLP) is a challenge that spans ten tasks: question answering, machine translation, summarization, natural language inference, sentiment analysis, semantic role labeling, zero-shot relation…
28 papers · 0 benchmarks
Consists of millions of entries in which the MT element of the training triplets has been obtained by translating the source side of publicly-available parallel corpora, and using the target side as an artificial human post-edit.
28 papers · 0 benchmarks
The 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across 20 different newsgroups.
27 papers · 5 benchmarks
2D-3D Match Dataset is a new dataset of 2D-3D correspondences by leveraging the availability of several 3D datasets from RGB-D scans.
27 papers · 0 benchmarks
AFAD (Asian Face Age Dataset)
The Asian Face Age Dataset (AFAD) is a new dataset proposed for evaluating the performance of age estimation, which contains more than 160K facial images and the corresponding age and gender labels.
27 papers · 1 benchmark
AitW (Android in the Wild)
Android in the Wild (AitW) is a dataset for device-control research which is orders of magnitude larger than current datasets.
27 papers · 0 benchmarks
AndroidWorld is an environment for building and benchmarking autonomous computer control agents.
27 papers · 0 benchmarks
This dataset details the energy consumption of appliances in a low-energy building over 4.5 months.
27 papers · 0 benchmarks
Dataset of 64x64 images of a robot pushing objects on a table top.
27 papers · 2 benchmarks
300 news articles annotated with 1,727 bias spans and find evidence that informational bias appears in news articles more frequently than lexical bias.
27 papers · 0 benchmarks
CONAN (COunter NArratives through Nichesourcing)
COunter NArratives through Nichesourcing (CONAN) is a dataset that consists of 4,078 pairs over the 3 languages.
27 papers · 0 benchmarks
CaseHOLD (Case Holdings On Legal Decisions)
CaseHOLD (Case Holdings On Legal Decisions) is a law dataset comprised of over 53,000+ multiple choice questions to identify the relevant holding of a cited case.
27 papers · 2 benchmarks
Chaos NLI is a Natural Language Inference (NLI) dataset with 100 annotations per example (for a total of 464,500 annotations) for some existing data points in the development sets of SNLI, MNLI, and Abductive NLI.
27 papers · 0 benchmarks
In this work, we make the first attempt to evaluate LLMs in a more challenging code generation scenario, i.e.
27 papers · 0 benchmarks
DeeperForensics-1.0 represents the largest face forgery detection dataset by far, with 60,000 videos constituted by a total of 17.6 million frames, 10 times larger than existing datasets of the same kind.
27 papers · 0 benchmarks
Evidence Inference is a corpus for this task comprising 10,000+ prompts coupled with full-text articles describing RCTs.
27 papers · 0 benchmarks
GID (Gaofen Image Dataset)
Gaofen Image Dataset (GID) is a large-scale land-cover dataset constructed with Gaofen-2 (GF-2) satellite images.
27 papers · 0 benchmarks
Pulsar candidates collected during the HTRU survey.
27 papers · 0 benchmarks
Haze4k is a synthesized dataset with 4,000 hazy images, in which each hazy image has the associate ground truths of a latent clean image, a transmission map, and an atmospheric light ma
27 papers · 1 benchmark
The ISIC 2018 dataset was published by the International Skin Imaging Collaboration (ISIC) as a large-scale dataset of dermoscopy images.
27 papers · 1 benchmark
The data set contains 38 patches (of the same size), each consisting of a true orthophoto (TOP) extracted from a larger TOP mosaic.
27 papers · 2 benchmarks
A benchmark dataset for out-of-distribution detection.
27 papers · 1 benchmark
We contribute an IntentQA dataset with diverse intents in daily social activities.
27 papers · 2 benchmarks
IntrA is an open-access 3D intracranial aneurysm dataset that makes the application of points-based and mesh-based classification and segmentation models available.
27 papers · 2 benchmarks
KPTimes is a large-scale dataset of news texts paired with editor-curated keyphrases.
27 papers · 3 benchmarks
KoDF (Korean DeepFake Detection Dataset)
The Korean DeepFake Detection Dataset (KoDF) is a large-scale collection of synthesized and real videos focused on Korean subjects, used for the task of deepfake detection.
27 papers · 0 benchmarks
LDC2017T10 (Abstract Meaning Representation (AMR) Annotation Release 2.0)
Abstract Meaning Representation (AMR) Annotation Release 2.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
27 papers · 1 benchmark
LHQ (Landscapes High-Quality)
A dataset of 90,000 high-resolution nature landscape images, crawled from Unsplash and Flickr and preprocessed with Mask R-CNN and Inception V3.
27 papers · 4 benchmarks
MedMNIST v2 is a large-scale MNIST-like collection of standardized biomedical images, including 12 datasets for 2D and 6 datasets for 3D.
27 papers · 0 benchmarks
Multi-Modal-CelebA-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ.
27 papers · 3 benchmarks
Nighttime Driving is a dataset of road scenes consisting of 35,000 images ranging from daytime to twilight time and to nighttime.
27 papers · 2 benchmarks
The goal of this benchmark is to introduce a standard evaluation metric to measure the accuracy and robustness of 3D face reconstruction methods under variations in viewing angle, lighting, and common occlusions.
27 papers · 1 benchmark
ProofNet is a benchmark for autoformalization and formal proving of undergraduate-level mathematics.
27 papers · 0 benchmarks
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song)
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains 7,356 files (total size: 24.8 GB).
27 papers · 6 benchmarks
ROAD (ROAD: The ROad event Awareness Dataset for Autonomous Driving)
ROAD is designed to test an autonomous vehicle's ability to detect road events, defined as triplets composed by an active agent, the action(s) it performs and the corresponding scene locations.
27 papers · 0 benchmarks
RadarScenes is a real-world radar point cloud dataset for automotive applications.
27 papers · 0 benchmarks
S2Looking is a building change detection dataset that contains large-scale side-looking satellite images captured at varying off-nadir angles.
27 papers · 1 benchmark
Sentiment analysis of codemixed tweets.
27 papers · 0 benchmarks
The Caltech 101 Silhouettes dataset consists of 4,100 training samples, 2,264 validation samples and 2,307 test samples.
27 papers · 0 benchmarks
TG-ReDial is a a topic-guided conversational recommendation dataset for research on conversational/interactive recommender systems.
27 papers · 0 benchmarks
UVO (Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation)
UVO is a new benchmark for open-world class-agnostic object segmentation in videos.
27 papers · 2 benchmarks
VQA-HAT (Human ATtention) is a dataset to evaluate the informative regions of an image depending on the question being asked about it.
27 papers · 0 benchmarks
The Vid4 dataset is generally used for testing video super-resolution.
27 papers · 3 benchmarks
Dataset of high-resolution (4096×2160), high-fps (1000fps) video frames with extreme motion.
27 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.