Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 49 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2305–2352 of 12,172
VQA-E is a dataset for Visual Question Answering with Explanation, where the models are required to generate and explanation with the predicted answer.
17 papers · 0 benchmarks
ViViD++ (Vision for Visibility Dataset)
A dataset capturing diverse visual data formats that target varying luminance conditions, and was recorded from alternative vision sensors, by handheld or mounted on a car, repeatedly in the same space but in different conditions.
17 papers · 0 benchmarks
A temporal counterfactual dataset composing of 1000 short and natural video-caption pairs.
17 papers · 1 benchmark
We manually edited an aerial and a satellite imagery dataset of building samples and named it a WHU building dataset.
17 papers · 3 benchmarks
The Weakly Occluded Scene Text (WOST) dataset is a public dataset for scene text segmentation.
17 papers · 1 benchmark
A publicly available dataset with 242k labeled sections in English and German from two distinct domains: diseases and cities.
17 papers · 0 benchmarks
The York Urban Line Segment Database is a compilation of 102 images (45 indoor, 57 outdoor) of urban environments consisting mostly of scenes from the campus of York University and downtown Toronto, Canada.
17 papers · 2 benchmarks
iSarcasm is a dataset of tweets, each labelled as either sarcastic or nonsarcastic.
17 papers · 1 benchmark
3D Hand Pose is a multi-view hand pose dataset consisting of color images of hands and different kind of annotations for each: the bounding box and the 2D and 3D location on the joints in the hand.
16 papers · 0 benchmarks
AgeDB contains 16, 488 images of various famous people, such as actors/actresses, writers, scientists, politicians, etc.
16 papers · 3 benchmarks
The BASHI dataset is a corpus consisting of 50 Wall Street Journal (WSJ) articles.
16 papers · 0 benchmarks
BG-20k (Background Dataset - 20k)
BG-20k contains 20,000 high-resolution background images excluded salient objects, which can be used to help generate high quality synthetic data.
16 papers · 0 benchmarks
The CLEVR-Hans data set is a novel confounded visual scene data set, which captures complex compositions of different objects.
16 papers · 0 benchmarks
CMU-MOSI (Multimodal Corpus of Sentiment Intensity)
The Multimodal Corpus of Sentiment Intensity (CMU-MOSI) dataset is a collection of 2199 opinion video clips.
16 papers · 2 benchmarks
COLDataset is a dataset to facilitate Chinese offensive language detection and model evaluation.
16 papers · 0 benchmarks
CURE-TSR (CURE Traffic Sign Recognition)
Includes more than two million traffic sign images that are based on real-world and simulator data.
16 papers · 0 benchmarks
Casual Conversations dataset is designed to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of age, genders, apparent skin tones and ambient lighting conditions.
16 papers · 0 benchmarks
ChemProt consists of 1,820 PubMed abstracts with chemical-protein interactions annotated by domain experts and was used in the BioCreative VI text mining chemical-protein interactions shared task.
16 papers · 1 benchmark
CholecT45 is a subset of CholecT50 consisting of 45 videos from the Cholec80 dataset.
16 papers · 1 benchmark
Node classification on Cornell with the fixed 48%/32%/20% splits provided by Geom-GCN.
16 papers · 2 benchmarks
Node classification on Cornell with 60%/20%/20% random splits for training/validation/test.
16 papers · 2 benchmarks
CryoNuSeg is a fully annotated FS-derived cryosectioned and H&E-stained nuclei instance segmentation dataset.
16 papers · 0 benchmarks
DADA-2000 is a large-scale benchmark with 2000 video sequences (named as DADA-2000) is contributed with laborious annotation for driver attention (fixation, saccade, focusing time), accident objects/intervals, as well as the accident…
16 papers · 0 benchmarks
DiPCo (DiPCo -- Dinner Party Corpus)
We present a speech data corpus that simulates a "dinner party" scenario taking place in an everyday home environment.
16 papers · 0 benchmarks
DynaSent is an English-language benchmark task for ternary (positive/negative/neutral) sentiment analysis.
16 papers · 1 benchmark
Expi (Extreme Pose Interaction)
Extreme Pose Interaction (ExPI) Dataset is a new person interaction dataset of Lindy Hop dancing actions.
16 papers · 3 benchmarks
See paper: Caldas, Sebastian, et al.
16 papers · 2 benchmarks
The FIVR-200K dataset has been collected to simulate the problem of Fine-grained Incident Video Retrieval (FIVR).
16 papers · 1 benchmark
FUSS (Free Universal Sound Separation)
The Free Universal Sound Separation (FUSS) dataset is a database of arbitrary sound mixtures and source-level references, for use in experiments on arbitrary sound separation.
16 papers · 0 benchmarks
FaithDial is a new benchmark for hallucination-free dialogues, by editing hallucinated responses in the Wizard of Wikipedia (WoW) benchmark.
16 papers · 0 benchmarks
This dataset, based on Flickr30K, is introduced in Learning with Noisy Correspondence for Cross-modal Matching.
16 papers · 1 benchmark
GRIT (General Robust Image Task Benchmark)
The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources.
16 papers · 5 benchmarks
GSV-Cities is a large-scale dataset for training deep neural network for the task of Visual Place Recognition.
16 papers · 0 benchmarks
GeneCIS benchmark is designed for measuring models’ ability to adapt to a range of similarity conditions, which is zero-shot evaluation only.
16 papers · 1 benchmark
GitTables is a corpus of currently 1M relational tables extracted from CSV files in GitHub covering 96 topics.
16 papers · 0 benchmarks
The Groove MIDI Dataset (GMD) is composed of 13.6 hours of aligned MIDI and (synthesized) audio of human-performed, tempo-aligned expressive drumming.
16 papers · 2 benchmarks
HVU (Holistic Video Understanding)
HVU is organized hierarchically in a semantic taxonomy that focuses on multi-label and multi-task video understanding as a comprehensive problem that encompasses the recognition of multiple semantic aspects in the dynamic scene.
16 papers · 0 benchmarks
HumanEval-X is a benchmark for evaluating the multilingual ability of code generative models.
16 papers · 0 benchmarks
ICFG-PEDES (Identity-Centric and Fine-Grained Person Description Dataset)
One large-scale database for Text-to-Image Person Re-identification, i.e., Text-based Person Retrieval.
16 papers · 3 benchmarks
IDRiD (Indian Diabetic Retinopathy Image Dataset)
Indian Diabetic Retinopathy Image Dataset (IDRiD) dataset consists of typical diabetic retinopathy lesions and normal retinal structures annotated at a pixel level.
16 papers · 3 benchmarks
ILDC (Indian Legal Documents Corpus)
The ILDC dataset (Indian Legal Documents Corpus) is a large corpus of 35k Indian Supreme Court cases annotated with original court decisions.
16 papers · 0 benchmarks
IndicGLUE (Indic General Language Understanding Evaluation Benchmark)
We now introduce IndicGLUE, the Indic General Language Understanding Evaluation Benchmark, which is a collection of various NLP tasks as de- scribed below.
16 papers · 4 benchmarks
Jester Gesture Recognition dataset includes 148,092 labeled video clips of humans performing basic, pre-defined hand gestures in front of a laptop camera or webcam.
16 papers · 6 benchmarks
JuICe is a corpus of 1.5 million examples with a curated test set of 3.7K instances based on online programming assignments.
16 papers · 0 benchmarks
LMDrive Dataset consists of 64K instruction-sensor-control data clips collected in the CARLA simulator, where each clip includes one navigation instruction, several notice instructions, a sequence of multi-modal multi-view sensor data, and…
16 papers · 0 benchmarks
Aims to extract events and their arguments from multimedia documents.
16 papers · 0 benchmarks
MATRES (Multi-Axis Temporal RElations for Start-points)
This is the Multi-Axis Temporal RElations for Start-points (i.e., MATRES) dataset
16 papers · 2 benchmarks
MFR (Ongoing version of ICCV-2021 Masked Face Recognition Challenge & Workshop(MFR))
During the COVID-19 coronavirus epidemic, almost everyone wears a facial mask, which poses a huge challenge to face recognition.
16 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.