Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 58 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2737–2784 of 12,172
Update on 3DIdent, where we introduce six additional object classes (Hare, Dragon, Cow, Armadillo, Horse, and Head), and impose a causal graph over the latent variables.
12 papers · 1 benchmark
The CelebA-Dialog dataset has the following properties: 1) Facial images are annotated with rich fine-grained labels, which classify one attribute into multiple degrees according to its semantic meaning; 2) Accompanied with each image,…
12 papers · 0 benchmarks
The CELL benchmark is made of fluorescence microscopy images of cell.
12 papers · 3 benchmarks
Datasets for multi-view crowd counting in wide-area scenes.
12 papers · 1 benchmark
CommitPack is is a 4TB dataset of commits scraped from GitHub repositories that are permissively licensed.
12 papers · 0 benchmarks
ConditionalQA is a Question Answering (QA) dataset that contains complex questions with conditional answers, i.e.
12 papers · 1 benchmark
The latest CosmoFlow dataset includes around 10,000 cosmological N-body dark matter simulations.
12 papers · 0 benchmarks
A popular dataset for node classification on heterogeneous graphs.
12 papers · 1 benchmark
The DSTC7 Task 1 dataset is a dataset and task for goal-oriented dialogue.
12 papers · 1 benchmark
DWD (Diverse Weather Dataset)
Urban-scene detection dataset that consists of five different weather conditions: daytime-sunny, night-sunny, dusk-rainy, daytime-foggy, and night-rainy.
12 papers · 1 benchmark
The DramaQA focuses on two perspectives: 1) Hierarchical QAs as an evaluation metric based on the cognitive developmental stages of human intelligence.
12 papers · 1 benchmark
ECTSum is a dataset with transcripts of earnings calls (ECTs), hosted by public companies, as documents, and short experts-written telegram-style bullet point summaries derived from corresponding Reuters articles.
12 papers · 0 benchmarks
EPIC-SOUNDS is a large scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos from EPIC-KITCHENS-100.
12 papers · 2 benchmarks
Corpus containing 25206 sentences labelled with lexical instances of 717 idiomatic expressions.
12 papers · 0 benchmarks
Earnings-21, a 39-hour corpus of earnings calls containing entity-dense speech from nine different financial sectors.
12 papers · 0 benchmarks
EgoExoLearn is a fascinating dataset designed to bridge the gap between egocentric and exocentric views of procedural activities.
12 papers · 3 benchmarks
FLAME (Fire Luminosity Airborne-based Machine learning Evaluation)
FLAME is a fire image dataset collected by drones during a prescribed burning piled detritus in an Arizona pine forest.
12 papers · 1 benchmark
FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
FoolMeTwice (FM2 for short) is a large dataset of challenging entailment pairs collected through a fun multi-player game.
12 papers · 0 benchmarks
FSDKaggle2018 is an audio dataset containing 11,073 audio files annotated with 41 labels of the AudioSet Ontology.
12 papers · 1 benchmark
Chinese Few-shot Learning Evaluation Benchmark (FewCLUE) is a comprehensive small sample evaluation benchmark in Chinese.
12 papers · 5 benchmarks
GLGE (General Language Generation Evaluation)
GLGE is a general language generation evaluation benchmark which is composed of 8 language generation tasks, including Abstractive Text Summarization (CNN/DailyMail, Gigaword, XSUM, MSNews), Answer-aware Question Generation (SQuAD 1.1,…
12 papers · 0 benchmarks
GMOT-40 (Generic Multiple Object Tracking (GMOT))
GMOT-40 is the first public dense dataset for Generic Multiple Object Tracking (GMOT).
12 papers · 2 benchmarks
81 videos of 31 classes of ground terrain such as grass, gravel, asphalt and sand.
12 papers · 0 benchmarks
The Hands in action dataset (HIC) dataset has RGB-D sequences of hands interacting with objects.
12 papers · 0 benchmarks
HRSC2016 (High resolution ship collections 2016)
High-resolution ship collections 2016 (HRSC2016) is a data set used for scientific research.
12 papers · 1 benchmark
HRWSI (High-Resolution Web Stereo Image)
The HRWSI dataset consists of about 21K diverse high-resolution RGB-D image pairs derived from the Web stereo images.
12 papers · 0 benchmarks
A platform for research in embodied artificial intelligence (AI).
12 papers · 0 benchmarks
An in-the-wild stereo image dataset, comprising 49,368 image pairs contributed by users of the Holopix mobile social platform.
12 papers · 0 benchmarks
The Hutter Prize Wikipedia dataset, also known as enwiki8, is a byte-level dataset consisting of the first 100 million bytes of a Wikipedia XML dump.
12 papers · 1 benchmark
HyperKvasir dataset contains 110,079 images and 374 videos where it captures anatomical landmarks and pathological and normal findings.
12 papers · 2 benchmarks
An Independent components (IC) dataset containing spatiotemporal measures for over 200,000 ICs from more than 6,000 EEG recordings.
12 papers · 0 benchmarks
A new large dataset for illumination estimation.
12 papers · 0 benchmarks
IRS (Indoor Robotics Stereo)
IRS is an open dataset for indoor robotics vision tasks, especially disparity and surface normal estimation.
12 papers · 0 benchmarks
This split was introduced in TEMI (BMVC 2023) Adaloglou, Nikolas, Felix Michels, Hamza Kalisch, and Markus Kollmann.
12 papers · 4 benchmarks
100 tasks from LIBERO-100 suite.
12 papers · 1 benchmark
LLaMEA (algorithms and experiments from the paper)
3500+ Generated evolutionary algorithms by the LLaMEA framework.
12 papers · 0 benchmarks
LANI is a 3D navigation environment and corpus, where an agent navigates between landmarks.
12 papers · 0 benchmarks
Lila is a unified mathematical reasoning benchmark consisting of 23 diverse tasks along four dimensions: (i) mathematical abilities e.g., arithmetic, calculus (ii) language format e.g., question-answering, fill-in-the-blanks (iii) language…
12 papers · 0 benchmarks
Math-Vision (Math-V) dataset is a meticulously curated collection of 3,040 high-quality mathematical problems with visual contexts sourced from real math competitions.
12 papers · 1 benchmark
MM-COVID (Multilingual and Multidimensional COVID-19 Fake News Data Repository)
MM-COVID is a dataset for fake news detection related to COVID-19.
12 papers · 0 benchmarks
MMNeedle (Multimodal Needle in a Haystack)
We introduce the MultiModal Needle-in-a-haystack (MMNeedle) benchmark, specifically designed to assess the long-context capabilities of MLLMs.
12 papers · 1 benchmark
MOW (3D dataset of humans Manipulating Objects in-the-Wild)
3D dataset of humans Manipulating Objects in-the-Wild (MOW).
12 papers · 1 benchmark
MP-100 (Mulit-category Pose Dataset)
The first large-scale pose dataset containing objects of multiple super-categories, termed Multi-category Pose (MP-100).
12 papers · 1 benchmark
MUSES (MUlti-Shot EventS)
MUSES is a large-scale dataset for temporal event (action) localization.
12 papers · 1 benchmark
The Maximum Unbiased Validation (MUV) dataset is a benchmark dataset selected from PubChem BioAssay.
12 papers · 3 benchmarks
The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.
12 papers · 0 benchmarks
MasakhaNEWS is a benchmark dataset for news topic classification covering 16 languages widely spoken in Africa.
12 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.