Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 18 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 817–864 of 12,172
Multilingual LibriSpeech is a large multilingual corpus suitable for speech research.
77 papers · 2 benchmarks
OC20 (Open Catalyst 2020)
Open Catalyst 2020 is a dataset for catalysis in chemical engineering.
77 papers · 1 benchmark
PRW (Person Re-identification in the Wild)
PRW is a large-scale dataset for end-to-end pedestrian detection and person recognition in raw video frames.
77 papers · 1 benchmark
ASVspoof 2019 (Automatic Speaker Verification Spoofing And Countermeasures Challenge)
76 papers · 0 benchmarks
ECB+ (extension to the EventCorefBank)
The ECB+ corpus is an extension to the EventCorefBank (ECB, Bejan and Harabagiu, 2010).
76 papers · 0 benchmarks
FineGym is an action recognition dataset build on top of gymnasium videos.
76 papers · 0 benchmarks
The Long Range Graph Benchmark (LRGB) is a collection of 5 graph learning datasets that arguably require long-range reasoning to achieve strong performance in a given task.
76 papers · 4 benchmarks
QM9 provides quantum chemical properties (at DFT level) for a relevant, consistent, and comprehensive chemical space of small organic molecules.
76 papers · 9 benchmarks
TAT-QA (Tabular And Textual dataset for Question Answering) is a large-scale QA dataset, aiming to stimulate progress of QA research over more complex and realistic tabular and textual data, especially those requiring numerical reasoning.
76 papers · 1 benchmark
The TuSimple dataset consists of 6,408 road images on US highways.
76 papers · 1 benchmark
A2D2 (Audi Autonomous Driving Dataset)
Audi Autonomous Driving Dataset (A2D2) consists of simultaneously recorded images and 3D point clouds, together with 3D bounding boxes, semantic segmentation, instance segmentation, and data extracted from the automotive bus.
75 papers · 0 benchmarks
AISHELL-2 contains 1000 hours of clean read-speech data from iOS is free for academic usage.
75 papers · 4 benchmarks
ARKitScenes is an RGB-D dataset captured with the widely available Apple LiDAR scanner.
75 papers · 2 benchmarks
HACS (Human Action Clips and Segments)
HACS is a dataset for human action recognition.
75 papers · 2 benchmarks
MetFaces is an image dataset of human faces extracted from works of art.
75 papers · 2 benchmarks
OVIS (Occluded Video Instance Segmentation)
OVIS is a new large scale benchmark dataset for video instance segmentation task.
75 papers · 1 benchmark
Rendered Hand Pose (RHD) is a dataset for hand pose estimation.
75 papers · 0 benchmarks
The Semantic Scholar corpus (S2) is composed of titles from scientific papers published in machine learning conferences and journals from 1985 to 2017, split by year (33 timesteps).
75 papers · 0 benchmarks
ApolloScape is a large dataset consisting of over 140,000 video frames (73 street scene videos) from various locations in China under varying weather conditions.
74 papers · 4 benchmarks
PETA (Pedestrian Attribute)
The PEdesTrian Attribute dataset (PETA) is a dataset fore recognizing pedestrian attributes, such as gender and clothing style, at a far distance.
74 papers · 1 benchmark
DDAD (Dense Depth for Autonomous Driving)
DDAD is a new autonomous driving benchmark from TRI (Toyota Research Institute) for long range (up to 250m) and dense depth estimation in challenging and diverse urban conditions.
73 papers · 1 benchmark
To the best of our knowledge this is the largest publicly available dataset of face images with gender and age labels for training.
73 papers · 0 benchmarks
Continuous speech separation (CSS) is an approach to handling overlapped speech in conversational audio signals.
73 papers · 2 benchmarks
An open-source simulated benchmark for meta-reinforcement learning and multi-task learning consisting of 50 distinct robotic manipulation tasks.
73 papers · 0 benchmarks
RFW (Racial Faces in-the-Wild)
To validate the racial bias of four commercial APIs and four state-of-the-art (SOTA) algorithms.
73 papers · 0 benchmarks
misc @inproceedings{RITAC18, author = {Radenovi\'{c}, F.
73 papers · 3 benchmarks
TREC-COVID is a community evaluation designed to build a test collection that captures the information needs of biomedical researchers using the scientific literature during a pandemic.
73 papers · 1 benchmark
TrecQA (Text Retrieval Conference Question Answering)
Text Retrieval Conference Question Answering (TrecQA) is a dataset created from the TREC-8 (1999) to TREC-13 (2004) Question Answering tracks.
73 papers · 3 benchmarks
VisDrone is a large-scale benchmark with carefully annotated ground-truth for various important computer vision tasks, to make vision meet drones.
73 papers · 1 benchmark
BABEL is a large dataset with language labels describing the actions being performed in mocap sequences.
72 papers · 1 benchmark
Birdsnap is a large bird dataset consisting of 49,829 images from 500 bird species with 47,386 images used for training and 2,443 images used for testing.
72 papers · 2 benchmarks
DS-1000 is a code generation benchmark with a thousand data science questions spanning seven Python libraries that (1) reflects diverse, realistic, and practical use cases, (2) has a reliable metric, (3) defends against memorization by…
72 papers · 0 benchmarks
GRAB is a dataset of full-body motions interacting and grasping 3D objects.
72 papers · 1 benchmark
MASSIVE is a parallel dataset of > 1M utterances across 51 languages with annotations for the Natural Language Understanding tasks of intent prediction and slot annotation.
72 papers · 3 benchmarks
The PKU-MMD dataset is a large skeleton-based action detection dataset.
72 papers · 4 benchmarks
Scan2CAD is an alignment dataset based on 1506 ScanNet scans with 97607 annotated keypoints pairs between 14225 (3049 unique) CAD models from ShapeNet and their counterpart objects in the scans.
72 papers · 1 benchmark
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
CARPK (car parking lot dataset)
The Car Parking Lot Dataset (CARPK) contains nearly 90,000 cars from 4 different parking lots collected by means of drone (PHANTOM 3 PROFESSIONAL).
71 papers · 1 benchmark
SPair-71k contains 70,958 image pairs with diverse variations in viewpoint and scale.
71 papers · 2 benchmarks
The SemanticPOSS dataset for 3D semantic segmentation contains 2988 various and complicated LiDAR scans with large quantity of dynamic instances.
71 papers · 1 benchmark
The shared task of CoNLL-2002 concerns language-independent named entity recognition.
70 papers · 3 benchmarks
The Comprehensive Cars (CompCars) dataset contains data from two scenarios, including images from web-nature and surveillance-nature.
70 papers · 1 benchmark
DiffusionDB is a large-scale text-to-image prompt dataset.
70 papers · 1 benchmark
A new large-scale question-answering dataset that requires reasoning on heterogeneous information.
70 papers · 1 benchmark
MSP-IMPROV (MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception)
We present the MSP-IMPROV corpus, a multimodal emotional database, where the goal is to have control over lexical content and emotion while also promoting naturalness in the recordings.
70 papers · 1 benchmark
MuPoTS-3D (Multiperson Pose Test Set in 3DMulti-person Pose estimation Test Set in 3D)
MuPoTs-3D (Multi-person Pose estimation Test Set in 3D) is a dataset for pose estimation composed of more than 8,000 frames from 20 real-world scenes with up to three subjects.
70 papers · 3 benchmarks
OmniObject3D is a large vocabulary 3D object dataset with massive high-quality real-scanned 3D objects.
70 papers · 0 benchmarks
RSICD (Remote Sensing Image Captioning Dataset)
70 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.