Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 52 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2449–2496 of 12,172
LIVE-FB LSVQ (LIVE-FB Large-Scale Social Video Quality (LSVQ) Database)
No-reference (NR) perceptual video quality assessment (VQA) is a complex, unsolved, and important problem to social and streaming media applications.
15 papers · 1 benchmark
MALF (Multi-Attribute Labelled Faces)
The MALF dataset is a large dataset with 5,250 images annotated with multiple facial attributes and it is specifically constructed for fine grained evaluation.
15 papers · 0 benchmarks
Extension test cases of MBPP, as well as generated code.
15 papers · 1 benchmark
MIAP (More Inclusive Annotations for People)
MIAP is a dataset created by obtaining a new set of annotations on a subset of the Open Images dataset, containing bounding boxes and attributes for all of the people visible in those images, as the original Open Images dataset annotations…
15 papers · 0 benchmarks
Existing hate speech datasets contain only textual data.
15 papers · 0 benchmarks
The MSK dataset is a dataset for lesion recognition from the Memorial Sloan-Kettering Cancer Center.
15 papers · 0 benchmarks
MSRA Hands is a dataset for hand tracking.
15 papers · 1 benchmark
MUSDB18-HQ is a high-quality version of the MUSDB18 music tracks dataset.
15 papers · 1 benchmark
MaSS (Multilingual corpus of Sentence-aligned Spoken utterances) is an extension of the CMU Wilderness Multilingual Speech Dataset, a speech dataset based on recorded readings of the New Testament.
15 papers · 3 benchmarks
MedVidQA (Medical Video Question Answering)
The MedVidQA dataset contains the collection of 3, 010 manually created health-related questions and timestamps as visual answers to those questions from trusted video sources, such as accredited medical schools with an established…
15 papers · 0 benchmarks
The Montreal Archive of Sleep Studies (MASS) is an open-access and collaborative database of laboratory-based polysomnography (PSG) recordings O’Reilly, C., et al.
15 papers · 4 benchmarks
MuCGEC (Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction)
MuCGEC is a multi-reference multi-source evaluation dataset for Chinese Grammatical Error Correction (CGEC), consisting of 7,063 sentences collected from three different Chinese-as-a-Second-Language (CSL) learner sources.
15 papers · 1 benchmark
N-ImageNet (Large-Scale Dataset for Event-Based Object Recognition)
The N-ImageNet dataset is an event-camera counterpart for the ImageNet dataset.
15 papers · 2 benchmarks
Preprocessed version of NYT11.
15 papers · 1 benchmark
OCTID (Optical Coherence Tomography Image Retinal Database)
An open-source Optical Coherence Tomography Image Database containing different retinal OCT images with various pathological conditions.
15 papers · 0 benchmarks
Large-scale American Sign Language (ASL) - English dataset collected from online video sites (e.g., YouTube).
15 papers · 0 benchmarks
Opusparcus is a paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish.
15 papers · 0 benchmarks
PANDA is the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis.
15 papers · 0 benchmarks
The Paris-Lille-3D is a Benchmark on Point Cloud Classification.
15 papers · 1 benchmark
Pile of Law is a ∼256GB (and growing) dataset of legal and administrative data which can be used for assessing norms on data sanitization across legal and administrative settings.
15 papers · 0 benchmarks
PointOdyssey is a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms.
15 papers · 1 benchmark
Project CodeNet is a large-scale dataset with approximately 14 million code samples, each of which is an intended solution to one of 4000 coding problems.
15 papers · 0 benchmarks
Node classification on PubMed with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
The goal of PubTables-1M is to create a large, detailed, high-quality dataset for training and evaluating a wide variety of models for the tasks of table detection, table structure recognition, and functional analysis.
15 papers · 0 benchmarks
QED is a linguistically principled framework for explanations in question answering.
15 papers · 1 benchmark
A large-scale dataset for retrieval and event localisation in video.
15 papers · 1 benchmark
REDS (REalistic and Diverse Scenes dataset
realistic and dynamic scenes)
The realistic and dynamic scenes (REDS) dataset was proposed in the NTIRE19 Challenge.
15 papers · 1 benchmark
ROSTD (Real Out-of-Domain Sentences From Task-oriented Dialog)
A dataset of 4K out-of-domain (OOD) examples for the publicly available dataset from (Schuster et al.
15 papers · 0 benchmarks
RTMV is a large-scale synthetic dataset for novel view synthesis consisting of ∼300k images rendered from nearly 2000 complex scenes using high-quality ray tracing at high resolution (1600 × 1600 pixels).
15 papers · 1 benchmark
RepCount (Repetitive Action Counting Dataset)
Counting repetitive actions are widely seen in human activities such as physical exercise.
15 papers · 1 benchmark
RoadAnomaly21 is a dataset for anomaly segmentation, the task of identify the image regions containing objects that have never been seen during training.
15 papers · 0 benchmarks
SVIRO (Synthetic Vehicle Interior Rear Seat Occupancy Dataset)
Contains bounding boxes for object detection, instance segmentation masks, keypoints for pose estimation and depth images for each synthetic scenery as well as images for each individual seat for classification.
15 papers · 0 benchmarks
SYSU-30k contains 30k categories of persons, which is about 20 times larger than CUHK03 (1.3k categories) and Market1501 (1.5k categories), and 30 times larger than ImageNet (1k categories).
15 papers · 2 benchmarks
Salinas Scene is a hyperspectral dataset collected by the 224-band AVIRIS sensor over Salinas Valley, California, and is characterized by high spatial resolution (3.7-meter pixels).
15 papers · 3 benchmarks
Recent advances in language-image pre-training has witnessed the emerging field of building transferable systems that can effortlessly adapt to a wide range of computer vision & multimodal tasks in the wild.
15 papers · 1 benchmark
A Benchmark for Robust Multi-Hop Spatial Reasoning in Texts
15 papers · 1 benchmark
Syn2Real, a synthetic-to-real visual domain adaptation benchmark meant to encourage further development of robust domain transfer methods.
15 papers · 1 benchmark
TITAN consists of 700 labeled video-clips (with odometry) captured from a moving vehicle on highly interactive urban traffic scenes in Tokyo.
15 papers · 0 benchmarks
ToxCast is an initiative by the U.S.
15 papers · 4 benchmarks
ToyADMOS dataset is a machine operating sounds dataset of approximately 540 hours of normal machine operating sounds and over 12,000 samples of anomalous sounds collected with four microphones at a 48kHz sampling rate, prepared by Yuma…
15 papers · 0 benchmarks
UBody is a large-scale Upper-Body dataset with the following annotations: 2D whole-body keypoints 3D SMPLX annotations Frame validity label Person bounding box (bbox) Hand bounding box (bbox)
15 papers · 1 benchmark
This dataset includes 4,500 fully annotated images (over 30,000 license plate characters) from 150 vehicles in real-world scenarios where both the vehicle and the camera (inside another vehicle) are moving.
15 papers · 1 benchmark
A new dataset for the low-resource language as Vietnamese to evaluate MRC models.
15 papers · 0 benchmarks
VGMIDI is a dataset of piano arrangements of video game soundtracks.
15 papers · 0 benchmarks
Washington RGB-D is a widely used testbed in the robotic community, consisting of 41,877 RGB-D images organized into 300 instances divided in 51 classes of common indoor objects (e.g.
15 papers · 0 benchmarks
Node classification on Wisconsin with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 2 benchmarks
A large multilingual benchmark, XL-WiC, featuring gold standards in 12 new languages from varied language families and with different degrees of resource availability, opening room for evaluation scenarios such as zero-shot cross-lingual…
15 papers · 0 benchmarks
xCodeEval is one of the largest executable multilingual multitask benchmarks consisting of 17 programming languages with execution-level parallelism.
15 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.