Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 50 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2353–2400 of 12,172
MathBench (MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark)
MathBench is an All in One math dataset for language model evaluation, with: A Sophisticated Five-Stage Difficulty Mechanism: Unlike the usual mathematical datasets that can only evaluate a single difficulty level or have a mix of unclear…
16 papers · 0 benchmarks
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S.
16 papers · 1 benchmark
MinneApple is a benchmark dataset for apple detection and segmentation.
16 papers · 0 benchmarks
The NYU Hand pose dataset contains 8252 test-set and 72757 training-set frames of captured RGBD data with ground-truth hand-pose information.
16 papers · 1 benchmark
NewsCLIPpings is a dataset for detecting mismatched images and captions.
16 papers · 0 benchmarks
OntoNotes Release 4.0 contains the content of earlier releases -- OntoNotes Release 1.0 LDC2007T21, OntoNotes Release 2.0 LDC2008T04 and OntoNotes Release 3.0 LDC2009T24 -- and adds newswire, broadcast news, broadcast conversation and web…
16 papers · 1 benchmark
PIPA (People in Photo Album)
The PIPA database is collected from Flickr photo albums for the task of person recognition.
16 papers · 1 benchmark
PasticineLab is a differentiable physics benchmark, which includes a diverse collection of soft body manipulation tasks.
16 papers · 0 benchmarks
A database with 2,000 videos captured by surveillance cameras in real-world scenes.
16 papers · 1 benchmark
SKM-TEA (Stanford Knee MRI with Multi-Task Evaluation)
The SKM-TEA dataset pairs raw quantitative knee MRI (qMRI) data, image data, and dense labels of tissues and pathology for end-to-end exploration and evaluation of the MR imaging pipeline.
16 papers · 0 benchmarks
SMHD (Self-reported Mental Health Diagnoses)
A novel large dataset of social media posts from users with one or multiple mental health conditions along with matched control users.
16 papers · 0 benchmarks
SONAR, a new multilingual and multimodal fixed-size sentence embedding space, with a full suite of speech and text encoders and decoders.
16 papers · 0 benchmarks
SPGISpeech (pronounced “speegie-speech”) is a large-scale transcription dataset, freely available for academic research.
16 papers · 1 benchmark
SPoC (Pseudocode-to-Code)
Pseudocode-to-Code (SPoC) is a program synthesis dataset, containing 18,356 programs with human-authored pseudocode and test cases.
16 papers · 2 benchmarks
Raw sensor dataset where each sequence captures 7 (few contain 6) images (RAW and JPG) taken by different focal lengths.
16 papers · 0 benchmarks
SemArt is a multi-modal dataset for semantic art understanding.
16 papers · 0 benchmarks
ShEMO (Sharif Emotional Speech Database)
The database includes 3000 semi-natural utterances, equivalent to 3 hours and 25 minutes of speech data extracted from online radio plays.
16 papers · 1 benchmark
SketchGraphs is a dataset of 15 million sketches extracted from real-world CAD models intended to facilitate research in both ML-aided design and geometric program induction.
16 papers · 0 benchmarks
SketchyCOCO dataset consists of two parts: Object-level data Object-level data contains 20198(train18869+val1329) triplets of {foreground sketch, foreground image, foreground edge map} examples covering 14 classes,…
16 papers · 1 benchmark
We introduce Stanford-ORB, a new real-world 3D Object inverse Rendering Benchmark.
16 papers · 3 benchmarks
TSU (Toyota Smarthome Untrimmed)
Toyota Smarthome Untrimmed (TSU) is a dataset for activity detection in long untrimmed videos.
16 papers · 1 benchmark
TV show Caption is a large-scale multimodal captioning dataset, containing 261,490 caption descriptions paired with 108,965 short video moments.
16 papers · 1 benchmark
To address the need for a standard open domain table benchmark dataset, the author propose a novel weak supervision approach to automatically create the TableBank, which is orders of magnitude larger than existing human labeled datasets…
16 papers · 0 benchmarks
Node classification on Texas with 60%/20%/20% random splits for training/validation/test.
16 papers · 1 benchmark
TextComplexityDE is a dataset consisting of 1000 sentences in German language taken from 23 Wikipedia articles in 3 different article-genres to be used for developing text-complexity predictor models and automatic text simplification in…
16 papers · 1 benchmark
TextSeg is a large-scale fine-annotated and multi-purpose text detection and segmentation dataset, collecting scene and design text with six types of annotations: word- and character-wise bounding polygons, masks and transcriptions.
16 papers · 1 benchmark
ToolQA is a question answering benchmark for Large Language Models (LLMs) which is designed to faithfully evaluate LLMs' ability to use external tools for question answering.
16 papers · 0 benchmarks
Traffic (Traffic Flow Forecasting Data Set)
Abstract: The task for this dataset is to forecast the spatio-temporal traffic volume based on the historical traffic volume and other features in neighboring locations.
16 papers · 2 benchmarks
TripClick is a large-scale dataset of click logs in the health domain, obtained from user interactions of the Trip Database health web search engine.
16 papers · 0 benchmarks
The TweepFake dataset consists of 25,572 social media messages posted either by bots or humans on Twitter.
16 papers · 1 benchmark
This task aims to probe stereotyping biases in the QA model/masked LM via underspecified examples, such as the following: Paragraph: An Asian woman was taking classes with a Caucasian woman.
16 papers · 0 benchmarks
V-D4RL provides pixel-based analogues of the popular D4RL benchmarking tasks, derived from the dmcontrol suite, along with natural extensions of two state-of-the-art online pixel-based continuous control algorithms, DrQ-v2 and DreamerV2,…
16 papers · 0 benchmarks
VehicleX is a large-scale synthetic dataset.
16 papers · 0 benchmarks
The ViGGO corpus is a set of 6,900 meaning representation to natural language utterance pairs in the video game domain.
16 papers · 1 benchmark
VideoLQ consists of videos downloaded from various video hosting sites such as Flickr and YouTube, with a Creative Common license.
16 papers · 1 benchmark
The View-of-Delft (VoD) dataset is a novel automotive dataset containing 8600 frames of synchronized and calibrated 64-layer LiDAR-, (stereo) camera-, and 3+1D radar-data acquired in complex, urban traffic.
16 papers · 1 benchmark
VoiceBank+DEMAND is a noisy speech database for training speech enhancement algorithms and TTS models.
16 papers · 1 benchmark
This dataset is a Wikipedia dump, split by relations to perform Few-Shot Knowledge Graph Completion.
16 papers · 0 benchmarks
Wukong is a large-scale Chinese cross-modal dataset for benchmarking different multi-modal pre-training methods to facilitate the Vision-Language Pre-training (VLP).
16 papers · 0 benchmarks
X-FACT is a large publicly available multilingual dataset for factual verification of naturally existing real-world claims.
16 papers · 0 benchmarks
XQLFW (Cross-Quality Labeled Faces in the Wild)
An evaluation protocol for face verification focusing on a large intra-pair image quality difference.
16 papers · 1 benchmark
e-SNLI-VE is a large VL (vision-language) dataset with NLEs (natural language explanations) with over 430k instances for which the explanations rely on the image content.
16 papers · 2 benchmarks
A large-scale stance detection dataset from comments written by candidates of elections in Switzerland.
16 papers · 0 benchmarks
Over 4 million frames of motion capture data for 100 different styles of locomotion.
15 papers · 1 benchmark
The 100PoisonMpts dataset is a significant initiative in the realm of large language model governance.
15 papers · 0 benchmarks
3D AffordanceNet is a dataset of 23k shapes for visual affordance.
15 papers · 1 benchmark
4Seasons is adataset covering seasonal and challenging perceptual conditions for autonomous driving.
15 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.