Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 27 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 1249–1296 of 12,172
RWC (Real World Computing Music Database)
The RWC (Real World Computing) Music Database is a copyright-cleared music database (DB) that is available to researchers as a common foundation for research.
43 papers · 0 benchmarks
ShARC (Shaping Answers with Rules through Conversation)
ShARC is a Conversational Question Answering dataset focussing on question answering from texts containing rules.
43 papers · 0 benchmarks
The Stacked MNIST dataset is derived from the standard MNIST dataset with an increased number of discrete modes.
43 papers · 1 benchmark
UK-DALE is an open-access dataset from the UK recording Domestic Appliance-Level Electricity to conduct research on disaggregation algorithms, with data describing not just the aggregate demand per building but also the ground truth'…
43 papers · 0 benchmarks
WebQA, is a new benchmark for multimodal multihop reasoning in which systems are presented with the same style of data as humans when searching the web: Snippets and Images.
43 papers · 0 benchmarks
gRefCOCO is the first large-scale Generalized Referring Expression Segmentation dataset that contains multi-target, no-target, and single-target expressions.
43 papers · 2 benchmarks
A2D (Actor-Action Dataset)
A2D (Actor-Action Dataset) is a dataset for simultaneously inferring actors and actions in videos.
42 papers · 1 benchmark
AVSpeech is a large-scale audio-visual dataset comprising speech clips with no interfering background signals.
42 papers · 0 benchmarks
BUCC (Building and Using Comparable Corpora)
The BUCC mining task is a shared task on parallel sentence extraction from two monolingual corpora with a subset of them assumed to be parallel, and that has been available since 2016.
42 papers · 4 benchmarks
C-GQA (Compositional GQA)
We propose a split built on top of Stanford GQA dataset originally proposed for VQA and name it Compositional GQA (C-GQA) dataset (see supplementary for the details).
42 papers · 0 benchmarks
CASIA-FASD is a small face anti-spoofing dataset containing 50 subjects.
42 papers · 0 benchmarks
CCMatrix uses ten snapshots of a curated common crawl corpus (Wenzek et al., 2019) totalling 32.7 billion unique sentences.
42 papers · 0 benchmarks
CHiME-5 (CHiME Speech Separation and Recognition Challenge)
The CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning.
42 papers · 0 benchmarks
DeformingThings4D is a synthetic dataset containing 1,972 animation sequences spanning 31 categories of humanoids and animals.
42 papers · 0 benchmarks
DuoRC contains 186,089 unique question-answer pairs created from a collection of 7680 pairs of movie plots where each pair in the collection reflects two versions of the same movie.
42 papers · 1 benchmark
We release E-commerce Dialogue Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
42 papers · 1 benchmark
The EPIC-KITCHENS-55 dataset comprises a set of 432 egocentric videos recorded by 32 participants in their kitchens at 60fps with a head mounted camera.
42 papers · 3 benchmarks
EmoContext consists of three-turn English Tweets.
42 papers · 1 benchmark
English Web Treebank is a dataset containing 254,830 word-level tokens and 16,624 sentence-level tokens of webtext in 1174 files annotated for sentence- and word-level tokenization, part-of-speech, and syntactic structure.
42 papers · 0 benchmarks
IconQA (Icon Question Answering)
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images in the daily-life context.
42 papers · 1 benchmark
KuaiRand is an unbiased sequential recommendation dataset collected from the recommendation logs of the video-sharing mobile app, Kuaishou (快手).
42 papers · 1 benchmark
Objaverse-XL is an extensive dataset containing over 10 million 3D objects.
42 papers · 0 benchmarks
The Oxford RobotCar Dataset contains over 100 repetitions of a consistent route through Oxford, UK, captured over a period of over a year.
42 papers · 3 benchmarks
PMLB (Penn Machine Learning Benchmarks)
The Penn Machine Learning Benchmarks (PMLB) is a large, curated set of benchmark datasets used to evaluate and compare supervised machine learning algorithms.
42 papers · 0 benchmarks
PearRead is a dataset of scientific peer reviews.
42 papers · 0 benchmarks
SCROLLS (Standardized CompaRison Over Long Language Sequences)
SCROLLS (Standardized CompaRison Over Long Language Sequences) is an NLP benchmark consisting of a suite of tasks that require reasoning over long texts.
42 papers · 1 benchmark
The UCR Time Series Archive - introduced in 2002, has become an important resource in the time series data mining community, with at least one thousand published papers making use of at least one data set from the archive.
42 papers · 2 benchmarks
VitaminC (Fact Verification with Contrastive Evidence)
The VitaminC dataset contains more than 450,000 claim-evidence pairs for fact verification and factual consistent generation.
42 papers · 0 benchmarks
Dataset containing RGB-D data of 4 large scenes, comprising a total of 12 rooms, for the purpose of RGB and RGB-D camera relocalization.
41 papers · 0 benchmarks
ACID (Aerial Coastline Imagery Dataset)
ACID consists of thousands of aerial drone videos of different coastline and nature scenes on YouTube.
41 papers · 1 benchmark
This dataset contains product reviews and metadata from Amazon, including 142.8 million reviews spanning May 1996 - July 2014.
41 papers · 5 benchmarks
CoAID include diverse COVID-19 healthcare misinformation, including fake news on websites and social platforms, along with users' social engagement about such news.
41 papers · 0 benchmarks
A creative writing task where the input is 4 random sentences and the output should be a coherent passage with 4 paragraphs that end in the 4 input sentences respectively.
41 papers · 0 benchmarks
DDD17 (DAVIS Driving Dataset 2017)
DDD17 has over 12 h of a 346x260 pixel DAVIS sensor recording highway and city driving in daytime, evening, night, dry and wet weather conditions, along with vehicle speed, GPS position, driver steering, throttle, and brake captured from…
41 papers · 1 benchmark
ExpW (Expression in-the-Wild)
The Expression in-the-Wild (ExpW) dataset is for facial expression recognition and contains 91,793 faces manually labeled with expressions.
41 papers · 1 benchmark
The I-Haze dataset contains 25 indoor hazy images (size 2833×4657 pixels) training.
41 papers · 1 benchmark
JFT-3B is an internal Google dataset and a larger version of the JFT-300M dataset.
41 papers · 0 benchmarks
KITTI Road is road and lane estimation benchmark that consists of 289 training and 290 test images.
41 papers · 0 benchmarks
Million-AID is a large-scale benchmark dataset containing a million instances for RS scene classification.
41 papers · 0 benchmarks
NT-VOT211 consists of 211 diverse videos, offering 211,000 well-annotated frames with 8 attributes including camera motion, deformation, fast motion, motion blur, tiny target, distractors, occlusion and out-of-view.
41 papers · 1 benchmark
This dataset consists of 33698 images from 221 identities.
41 papers · 2 benchmarks
Partial REID is a specially designed partial person reidentification dataset that includes 600 images from 60 people, with 5 full-body images and 5 occluded images per person.
41 papers · 1 benchmark
The Query-based Video Highlights (QVHighlights) dataset is a dataset for detecting customized moments and highlights from videos given natural language (NL).
41 papers · 4 benchmarks
SMM4H (Social Media Mining for Health Shared Task)
Social Media Mining for Health (SMM4H) Shared Task is a massive data source for biomedical and public health applications.
41 papers · 0 benchmarks
A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research.
41 papers · 0 benchmarks
AID (Aerial Image Dataset)
AID is a new large-scale aerial image dataset, by collecting sample images from Google Earth imagery.
40 papers · 2 benchmarks
BLURB (Biomedical Language Understanding and Reasoning Benchmark)
BLURB is a collection of resources for biomedical natural language processing.
40 papers · 2 benchmarks
Contains around 200K dialogs with a total of 1.6M turns.
40 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.