Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 47 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2209–2256 of 12,172
A novel dataset facilitating multimodal and Synergetic sociAL Scene Analysis.
18 papers · 1 benchmark
SCAND (Socially CompliAnt Navigation Dataset)
Have you wondered how autonomous mobile robots should share space with humans in public spaces?
18 papers · 0 benchmarks
SCICAP is a large-scale image captioning dataset that contains real-world scientific figures and captions.
18 papers · 1 benchmark
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically.
18 papers · 0 benchmarks
SeaDronesSee (SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water)
SeaDronesSee is a large-scale data set aimed at helping develop systems for Search and Rescue (SAR) using Unmanned Aerial Vehicles (UAVs) in maritime scenarios.
18 papers · 3 benchmarks
TUM-VIE (TUM Stereo Visual-Inertial Event Dataset)
TUM-VIE is an event camera dataset for developing 3D perception and navigation algorithms.
18 papers · 0 benchmarks
TuringBench is a benchmark environment that contains : - Benchmark tasks- Turing Test (i.e., human vs.
18 papers · 2 benchmarks
TempEval-3 (TempEval-3: events, times, and temporal relations)
Within the SemEval-2013 evaluation exercise, the TempEval-3 shared task aims to advance research on temporal information processing.
18 papers · 2 benchmarks
The first parallel corpus composed from United Nations documents published by the original data creator.
18 papers · 0 benchmarks
VAST (VAried Stance Topics)
VAST consists of a large range of topics covering broad themes, such as politics (e.g., ‘a Palestinian state’), education (e.g., ‘charter schools’), and public health (e.g., ‘childhood vaccination’).
18 papers · 1 benchmark
VidSitu is a dataset for the task of semantic role labeling in videos (VidSRL).
18 papers · 0 benchmarks
Violin (VIdeO-and-Language INference)
Video-and-Language Inference is the task of joint multimodal understanding of video and text.
18 papers · 0 benchmarks
WTW (Wired Table in the Wild)
WTW (Wired Table in the Wild) is a large-scale dataset which includes well-annotated structure parsing of multiple style tables in several scenes like the photo, scanning files, web pages.
18 papers · 1 benchmark
Weibo21 is a benchmark of fake news dataset for multi-domain fake news detection (MFND) with domain label annotated, which consists of 4,488 fake news and 4,640 real news from 9 different domains.
18 papers · 0 benchmarks
Many existing datasets for lidar place recognition are solely representative of structured urban environments, and have recently been saturated in performance by deep learning based approaches.
18 papers · 1 benchmark
Node classification on Wisconsin with 60%/20%/20% random splits for training/validation/test.
18 papers · 1 benchmark
X-CSQA is a multilingual dataset for Commonsense reasoning research, based on CSQA.
18 papers · 0 benchmarks
Yeast dataset consists of a protein-protein interaction network.
18 papers · 0 benchmarks
ZInd (Zillow Indoor Dataset)
The Zillow Indoor Dataset (ZInD) provides extensive visual data that covers a real world distribution of unfurnished residential homes.
18 papers · 1 benchmark
bFFHQ (Gender-biased FFHQ dataset)
Gender-biased FFHQ dataset (bFFHQ) has age as a target label and gender as a correlated bias, and the images are from the FFHQ dataset.
18 papers · 1 benchmark
Regression dataset for molecular docking scores (predicted molecule-protein binding affinity).
18 papers · 4 benchmarks
xSID (Cross-lingual Slot and Intent Detection)
xSID, a new evaluation benchmark for cross-lingual (X) Slot and Intent Detection in 13 languages from 6 language families, including a very low-resource dialect, covering Arabic (ar), Chinese (zh), Danish (da), Dutch (nl), English (en),…
18 papers · 0 benchmarks
AVeriTeC (AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the Web)
AVeriTeC (Automated Verification of Textual Claims) is a dataset of 4568 real-world claims covering fact-checks by 50 different organizations.
17 papers · 1 benchmark
ApolloCar3DT is a dataset that contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints.
17 papers · 14 benchmarks
Extracted from the Tashkeela Corpus, the dataset consists of 55K lines containing about 2.3M words.
17 papers · 1 benchmark
BABE (Bias Annotations By Experts)
BABE is an expertly annotated dataset aimed at facilitating media bias research.
17 papers · 0 benchmarks
Composed by 2.7 billion tokens, and has been annotated with tagging and parsing information.
17 papers · 0 benchmarks
BUG is a large-scale gender bias dataset of 108K diverse real-world English sentences, sampled semiautomatically from large corpora using lexical syntactic pattern matching
17 papers · 0 benchmarks
A renovation of Labeled Faces in the Wild (LFW), the de facto standard testbed for unconstraint face verification.
17 papers · 4 benchmarks
CASIA-HWDB is a dataset for handwritten Chinese character recognition.
17 papers · 0 benchmarks
CC152K (Conceptual Captions 152K)
CC152K is a subset of Conceptual Captions.
17 papers · 1 benchmark
CDTB (Color-and-Depth Tracking)
Source: https://www.vicos.si/Projects/CDTB 4.2 State-of-the-art Comparison A TH CTB (color-and-depth visual object tracking) dataset is recorded by several passive and active RGB-D setups and contains indoor as well as outdoor sequences…
17 papers · 0 benchmarks
CLEVR-Ref+ is a synthetic diagnostic dataset for referring expression comprehension.
17 papers · 1 benchmark
COCO-Noisy (Microsoft Common Objects in Context with 20% of Noisy Correspondence and 1K test data)
This dataset is based on MS COCO that have 20% of data randomly shuffled to simulate noisy correspondence.
17 papers · 1 benchmark
A renovation of Labeled Faces in the Wild (LFW), the de facto standard testbed for unconstraint face verification.
17 papers · 4 benchmarks
The CUTE80 dataset is a lightweight collection of images specifically designed for text detection in natural scene images.
17 papers · 1 benchmark
Node classification on Chameleon with 60%/20%/20% random splits for training/validation/test.
17 papers · 2 benchmarks
ConvQuestions is the first realistic benchmark for conversational question answering over knowledge graphs.
17 papers · 0 benchmarks
The D-HAZY dataset is generated from NYU depth indoor image collection.
17 papers · 0 benchmarks
DAiSEE is a multi-label video classification dataset comprising of 9,068 video snippets captured from 112 users for recognizing the user affective states of boredom, confusion, engagement, and frustration "in the wild".
17 papers · 1 benchmark
DESED (Domestic environment sound event detection)
The DESED dataset is a dataset designed to recognize sound event classes in domestic environments.
17 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
17 papers · 3 benchmarks
DialoGLUE is a natural language understanding benchmark for task-oriented dialogue designed to encourage dialogue research in representation-based transfer, domain adaptation, and sample-efficient task learning.
17 papers · 2 benchmarks
EgoTask QA benchmark contains 40K balanced question-answer pairs selected from 368K programmatically generated questions generated over 2K egocentric videos.
17 papers · 1 benchmark
FMD (Fluorescence Microscopy Denoising)
The Fluorescence Microscopy Denoising (FMD) dataset is dedicated to Poisson-Gaussian denoising.
17 papers · 2 benchmarks
FreebaseQA is a data set for open-domain QA over the Freebase knowledge graph.
17 papers · 0 benchmarks
By perturbing the widely used GSM8K dataset, an adversarial dataset for grade-school math called GSM-Plus is created.
17 papers · 1 benchmark
HPO-B is a benchmark for assessing the performance of HPO (Hyperparameter optimization) algorithms.
17 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.