Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 66 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3121–3168 of 12,172
SpaceNet 2: Building Detection v2 - is a dataset for building footprint detection in geographically diverse settings from very high resolution satellite images.
10 papers · 1 benchmark
SpeakingFaces is a publicly-available large-scale dataset developed to support multimodal machine learning research in contexts that utilize a combination of thermal, visual, and audio data streams; examples include human-computer…
10 papers · 0 benchmarks
StableToolBench is a new benchmark for tool learning that aims to provide a well-balanced combination of stability and reality, building upon its predecessor, ToolBench.
10 papers · 0 benchmarks
A large-scale human image dataset with over 230K samples capturing diverse poses and textures.
10 papers · 0 benchmarks
First large-scale symphony generation dataset.
10 papers · 1 benchmark
We include five substructure counting tasks: 3-stars, triangles, tailed triangles, chordal cycles and attributed triangles.
10 papers · 0 benchmarks
TAPOS is a new dataset developed on sport videos with manual annotations of sub-actions, and conduct a study on temporal action parsing on top.
10 papers · 1 benchmark
TMED (Tufts Medical Echocardiogram Dataset)
TMED is a clinically-motivated benchmark dataset for computer vision and machine learning from limited labeled data.
10 papers · 0 benchmarks
Five datasets used in NeurTraL-AD paper: \textit{RacketSports (RS).} Accelerometer and gyroscope recording of players playing four different racket sports.
10 papers · 1 benchmark
UGIF is a multi-lingual, multi-modal UI grounded dataset for step-by-step task completion on the smartphone.
10 papers · 0 benchmarks
UI-PRMD (University of Idaho – Physical Rehabilitation Movement Dataset)
UI-PRMD is a data set of movements related to common exercises performed by patients in physical therapy and rehabilitation programs.
10 papers · 2 benchmarks
USF (Human ID Gait Challenge Dataset)
The USF Human ID Gait Challenge Dataset is a dataset of videos for gait recognition.
10 papers · 0 benchmarks
Unite The People is a dataset for 3D body estimation.
10 papers · 0 benchmarks
The Vehicular Reference Misbehavior (VeReMi) dataset, is a dataset for the evaluation of misbehavior detection mechanisms for VANETs (vehicular networks).
10 papers · 0 benchmarks
ViP-Bench (Making Large Multimodal Models Understand Arbitrary Visual Prompts)
ViP-Bench is a comprehensive benchmark designed to assess the capability of multimodal models in understanding visual prompts across multiple dimensions.
10 papers · 1 benchmark
DataViSal.rar (including the ground truth data) is our new collected dataset for the following paper.
10 papers · 1 benchmark
We propose a new, scalable video-mining pipeline which transfers captioning supervision from image datasets to video and audio.
10 papers · 0 benchmarks
Composed of 10,000 videos annotated with memorability scores.
10 papers · 0 benchmarks
WikiAtomicEdits is a corpus of 43 million atomic edits across 8 languages.
10 papers · 0 benchmarks
A corpus that encompasses the complete history of conversations between contributors to Wikipedia, one of the largest online collaborative communities.
10 papers · 0 benchmarks
The WinoWhy dataset is a resource that provides human-annotated reasons for answering Winograd Schema Challenge (WSC) questions.
10 papers · 0 benchmarks
XFORMAL is a multilingual formal style transfer benchmark of multiple formal reformulations of informal text in Brazilian Portuguese, French, and Italian.
10 papers · 0 benchmarks
Arxiv ASTRO-PH (Astro Physics) collaboration network is from the e-print arXiv and covers scientific collaborations between authors papers submitted to Astro Physics category.
10 papers · 0 benchmarks
e-ViL is a benchmark for explainable vision-language tasks.
10 papers · 0 benchmarks
iGibson 2.0 is an open-source simulation environment that supports the simulation of a more diverse set of household tasks through three key innovations.
10 papers · 0 benchmarks
Annotated using images taken by a drone in 501 separate flights, totalling in over 62 hours of trajectory data.
10 papers · 0 benchmarks
The selfie dataset contains 46,836 selfie images annotated with 36 different attributes.
10 papers · 1 benchmark
xR-EgoPose is an egocentric synthetic dataset for egocentric 3D human pose estimation.
10 papers · 0 benchmarks
The Sixth Informatics for Integrating Biology and the Bedside (i2b2) Natural Language Processing Challenge for Clinical Records focused on the temporal relations in clinical narratives.
9 papers · 2 benchmarks
Our dataset which consists of multiple indoor and outdoor experiments for up to 30 m gNB-UE link.
9 papers · 0 benchmarks
The AIRS (Aerial Imagery for Roof Segmentation) dataset provides a wide coverage of aerial imagery with 7.5 cm resolution and contains over 220,000 buildings.
9 papers · 1 benchmark
AMZ Computers is a co-purchase graph extracted from Amazon, where nodes represent products, edges represent the co-purchased relations of products, and features are bag-of-words vectors extracted from product reviews.
9 papers · 1 benchmark
The question-answer (QA) pairs are automatically generated using state-of-the-art question generation methods based on paintings and comments provided in an existing art understanding dataset.
9 papers · 0 benchmarks
ART consists of over 20k commonsense narrative contexts and 200k explanations.
9 papers · 0 benchmarks
ATEPP (Automatically Transcribed Expressive Piano Performances)
ATEPP is a dataset of expressive piano performances by virtuoso pianists.
9 papers · 0 benchmarks
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods.
9 papers · 0 benchmarks
Is a collection of action videos from many different countries.
9 papers · 1 benchmark
Is an acronym disambiguation (AD) dataset for scientific domain with 62,441 samples which is significantly larger than the previous scientific AD dataset.
9 papers · 0 benchmarks
Adaptiope is a domain adaptation dataset with 123 classes in the three domains synthetic, product and real life.
9 papers · 0 benchmarks
To systematically evaluate the effectiveness of our approach at accomplishing this, we designed a new benchmark, AdvBench, based on two distinct settings.
9 papers · 0 benchmarks
AmsterTime (AmsterTime: A Visual Place Recognition Benchmark Dataset for Severe Domain Shift)
AmsterTime dataset offers a collection of 2,500 well-curated images matching the same scene from a street view matched to historical archival image data from Amsterdam city.
9 papers · 3 benchmarks
Accurately estimating the 3D pose and shape is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation.
9 papers · 1 benchmark
This dataset contains 8.9M commonsense assertions extracted by the Ascent pipeline developed at the Max Planck Institute for Informatics.
9 papers · 0 benchmarks
Atari-HEAD is a dataset of human actions and eye movements recorded while playing Atari videos games.
9 papers · 0 benchmarks
B-Pref is a benchmark specially designed for preference-based RL.
9 papers · 0 benchmarks
A high-resolution semantic segmentation dataset with 50 validation and 100 test objects.
9 papers · 1 benchmark
BIMCV-COVID19+ dataset is a large dataset with chest X-ray images CXR (CR, DX) and computed tomography (CT) imaging of COVID-19 patients along with their radiographic findings, pathologies, polymerase chain reaction (PCR), immunoglobulin G…
9 papers · 0 benchmarks
This dataset contains 1200 images (1000 WLI images and 200 FICE images) with fine-grained segmentation annotations.
9 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.