Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 76 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3601–3648 of 12,172
A configurable visual question and answer dataset (COG) to parallel experiments in humans and animals.
7 papers · 0 benchmarks
Random sampled instances of the Capacitated Vehicle Routing Problem with Time Windows (CVRPTW) for 20, 50 and 100 customer nodes.
7 papers · 0 benchmarks
ChangeIt dataset with more than 2600 hours of video with state-changing actions published at CVPR 2022.
7 papers · 0 benchmarks
ChangeSim is a dataset aimed at online scene change detection (SCD) and more.
7 papers · 2 benchmarks
The CheXmask Database presents a comprehensive, uniformly annotated collection of chest radiographs, constructed from five public databases: ChestX-ray8, Chexpert, MIMIC-CXR-JPG, Padchest and VinDr-CXR.
7 papers · 0 benchmarks
CochlScene is a dataset for acoustic scene classification.
7 papers · 1 benchmark
The CocoChorales Dataset CocoChorales is a dataset consisting of over 1400 hours of audio mixtures containing four-part chorales performed by 13 instruments, all synthesized with realistic-sounding generative models.
7 papers · 0 benchmarks
Five curated datasets of one-liner commits from open-source projects.
7 papers · 0 benchmarks
ConfAIde is a benchmark that evaluates the inference-time privacy implications of Language Models (LLMs) in interactive settings.
7 papers · 0 benchmarks
The dataset published here is the largest, most diverse and consistent crack segmentation dataset constructed so far.
7 papers · 0 benchmarks
The Cross-dataset Testbed is a Decaf7 based cross-dataset image classification dataset, which contains 40 categories of images from 3 domains: 3,847 images in Caltech256, 4,000 images in ImageNet, and 2,626 images for SUN.
7 papers · 0 benchmarks
DBLP Temporal is a dataset for temporal entity resolution, based on author profiles extracted from the Digital Bibliography and Library Project (DBLP).
7 papers · 1 benchmark
In this work we create a question answering dataset over the DBLP scholarly knowledge graph (KG).
7 papers · 0 benchmarks
DDD20 (DAVIS Driving Dataset 2020)
The dataset was captured with a DAVIS camera that concurrently streams both dynamic vision sensor (DVS) brightness change events and active pixel sensor (APS) intensity frames.
7 papers · 0 benchmarks
Danish Fungi 2020 (DF20) is a fine-grained dataset and benchmark.
7 papers · 1 benchmark
The Diabetic Foot Ulcers dataset (DFUC2021) is a dataset for analysis of pathology, focusing on infection and ischaemia.
7 papers · 0 benchmarks
DFW (Disguised Faces in the Wild)
Contains over 11000 images of 1000 identities with different types of disguise accessories.
7 papers · 3 benchmarks
DIBCO 2011 is the International Document Image Binarization Contest organized in the context of ICDAR 2011 conference.
7 papers · 0 benchmarks
DOLPHINS (Dataset for Collaborative Perception enabled Harmonious and Interconnected Self-driving)
Vehicle-to-Everything (V2X) network has enabled collaborative perception in autonomous driving, which is a promising solution to the fundamental defect of stand-alone intelligence including blind zones and long-range perception.
7 papers · 0 benchmarks
DPM (Don’t Patronize Me!)
Don’t Patronize Me!
7 papers · 4 benchmarks
DaNetQA (Yes/no Question Answering Dataset for the Russian)
DaNetQA is a question answering dataset for yes/no questions.
7 papers · 1 benchmark
Evaluate a natural language code generation model on real data science pedagogical notebooks!
7 papers · 0 benchmarks
Composes sentence pairs (i.e., twin sentences).
7 papers · 0 benchmarks
Demetr is a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological…
7 papers · 0 benchmarks
Trajectories of 3 dynamical systems: - Pendulum - Lotka-Voltera - 3-body system Code to re-create the datasets is provided in the repo on the folder datageneration
7 papers · 0 benchmarks
Includes egocentric videos containing hands in the wild.
7 papers · 0 benchmarks
EasyCall is a new dysarthric speech command dataset in Italian.
7 papers · 0 benchmarks
Contains annotated egocentric and top-view videos.
7 papers · 1 benchmark
EvoEval is a holistic benchmark suite created by evolving HumanEval problems¹.
7 papers · 0 benchmarks
EvoGym is a large-scale benchmark for co-optimizing the design and control of soft robots.
7 papers · 0 benchmarks
FLAT, a synthetic dataset of 2000 ToF measurements that capture all of these nonidealities, and can be used to simulate different hardware
7 papers · 0 benchmarks
FOR-instance (FOR-instance: a UAV laser scanning benchmark dataset for semantic and instance segmentation of individual trees)
The challenge of accurately segmenting individual trees from laser scanning data hinders the assessment of crucial tree parameters necessary for effective forest management, impacting many downstream applications.
7 papers · 0 benchmarks
FPL (First-Person Locomotion)
Supports new task that predicts future locations of people observed in first-person videos.
7 papers · 0 benchmarks
FPv1 (prior name FAUST-partial) is a 3D registration benchmark dataset created to address the lack of data variability in the existing 3D registration benchmarks such as: 3DMatch, ETH, KITTI.
7 papers · 1 benchmark
FRLL-Morphs is a dataset of morphed faces based on images selected from the publicly available Face Research London Lab dataset [1].
7 papers · 0 benchmarks
FathomNet is an open-source image database that can be used to train, test, and validate state-of-the-art artificial intelligence algorithms to help us understand our ocean and its inhabitants.
7 papers · 0 benchmarks
Former Flickr30k-CN translates the training and validation sets of Flickr30k using machine translation and manually translates the test set.
7 papers · 1 benchmark
FocusPath is a dataset compiled from diverse Whole Slide Image (WSI) scans in different focus (z-) levels.
7 papers · 0 benchmarks
A challenging multi-agent seasonal dataset collected by a fleet of Ford autonomous vehicles at different days and times during 2017-18.
7 papers · 0 benchmarks
ForecastQA is a question-answering dataset consisting of 10,392 event forecasting questions, which have been collected and verified via crowdsourcing efforts.
7 papers · 0 benchmarks
Freiburg Groceries is a groceries classification dataset consisting of 5000 images of size 256x256, divided into 25 categories.
7 papers · 0 benchmarks
GHOSTS is the first natural-language dataset made and curated by working researchers in mathematics that (1) aims to cover graduate-level mathematics and (2) provides a holistic overview of the mathematical capabilities of language models.
7 papers · 0 benchmarks
GINC (Generative IN-Context learning Dataset)
GINC (Generative In-Context learning Dataset) is a small-scale synthetic dataset for studying in-context learning.
7 papers · 0 benchmarks
GeoDE is a geographically diverse dataset with 61,940 images from 40 classes and 6 world regions, and no personally identifiable information, collected through crowd-sourcing.
7 papers · 0 benchmarks
This is a dataset for visual grasp affordance prediction that promotes more robust and heterogeneous robotic grasping methods.
7 papers · 0 benchmarks
Grocery Store is a dataset of natural images of grocery items.
7 papers · 0 benchmarks
A large-scale corpus of Gulf Arabic consisting of 110 million words from 1,200 forum novels.
7 papers · 0 benchmarks
H-DIBCO 2016 is the international Handwritten Document Image Binarization Contest organized in the context of ICFHR 2016 conference
7 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.