Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 48 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 2257–2304 of 12,172
JESC (Japanese-English Subtitle Corpus)
Japanese-English Subtitle Corpus is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue.
17 papers · 0 benchmarks
Kleister NDA is a dataset for Key Information Extraction (KIE).
17 papers · 1 benchmark
The Kumar dataset contains 30 1,000×1,000 image tiles from seven organs (6 breast, 6 liver, 6 kidney, 6 prostate, 2 bladder, 2 colon and 2 stomach) of The Cancer Genome Atlas (TCGA) database acquired at 40× magnification.
17 papers · 1 benchmark
LABR (Large-Scale Arabic Book Reviews)
LABR is a large sentiment analysis dataset to-date for the Arabic language.
17 papers · 1 benchmark
LCCC (Large-scale Cleaned Chinese Conversation corpus)
Contains a base version (6.8million dialogues) and a large version (12.0 million dialogues).
17 papers · 0 benchmarks
This dataset contains 2100+ high resolution indoor panoramas, captured using a Canon 5D Mark III and a robotic panoramic tripod head.
17 papers · 0 benchmarks
MEVA (Multiview Extended Video with Activities)
Large-scale dataset for human activity recognition.
17 papers · 0 benchmarks
MLRSNet is a a multi-label high spatial resolution remote sensing dataset for semantic scene understanding.
17 papers · 2 benchmarks
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
17 papers · 1 benchmark
MMDialog is a large-scale multi-turn dialogue dataset containing multi-modal open-domain conversations derived from real human-human chat content in social media.
17 papers · 1 benchmark
MalNet is a large public graph database, representing a large-scale ontology of software function call graphs.
17 papers · 2 benchmarks
Memorability dataset with 10000 3-second videos.
17 papers · 0 benchmarks
MoCA-Mask (Moving Camouflaged Animals (MoCA)-Mask)
The original Moving Camouflaged Animals (MoCA) Dataset includes 37K frames from 141 YouTube Video sequences with resolution and sampling rate of 720 × 1280 and 24fps in the majority of cases.
17 papers · 1 benchmark
The dataset for this challenge was obtained by carefully annotating tissue images of several patients with tumors of different organs and who were diagnosed at multiple hospitals.
17 papers · 2 benchmarks
Molecule3D is a new benchmark that includes a dataset with precise ground-state geometries of approximately 4 million molecules derived from density functional theory (DFT).
17 papers · 2 benchmarks
NINCO (No ImageNet Class Objects)
The NINCO (No ImageNet Class Objects) dataset is introduced in the ICML 2023 paper In or Out?
17 papers · 0 benchmarks
The New College Data is a freely available dataset collected from a robot completing several loops outdoors around the New College campus in Oxford.
17 papers · 0 benchmarks
NoReC (Norwegian Review Corpus)
The Norwegian Review Corpus (NoReC) was created for the purpose of training and evaluating models for document-level sentiment analysis.
17 papers · 0 benchmarks
The One-Minute Gradual-Emotional Behavior dataset (OMG-Emotion) dataset is composed of Youtube videos which are around a minute in length and are annotated taking into consideration a continuous emotional behavior.
17 papers · 0 benchmarks
ORConvQA (Open-Retrieval Conversational Question Answering)
Enhances QuAC by adapting it to an open-retrieval setting.
17 papers · 0 benchmarks
OpenXAI is the first general-purpose lightweight library that provides a comprehensive list of functions to systematically evaluate the quality of explanations generated by attribute-based explanation methods.
17 papers · 0 benchmarks
PCQM4Mv2 is a quantum chemistry dataset originally curated under the PubChemQC project.
17 papers · 1 benchmark
A large-scale English paraphrase dataset that surpasses prior work in both quantity and quality.
17 papers · 0 benchmarks
PartialSpoof is a dataset of partially-spoofed data to evaluate detection of partially-spoofed speech data.
17 papers · 0 benchmarks
PeMS08 is a traffic forecasting dataset.
17 papers · 1 benchmark
QAMR (Question-Answer Meaning Representation Dataset)
Question-Answer Meaning Representation (QAMR) represents a predicate-argument structure of a sentence with a set of question-answer pairs, so that annotations can be easily provided by non-experts.
17 papers · 0 benchmarks
Recent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online.
17 papers · 0 benchmarks
Quasimodo is commonsense knowledge base that focuses on salient properties of objects.
17 papers · 0 benchmarks
RPC (Retail Product Checkout)
RPC is a large-scale retail product checkout dataset and collects 200 retail SKUs.
17 papers · 0 benchmarks
RepoBench is a benchmark designed for evaluating repository-level code auto-completion systems, focusing on more complex, real-world programming scenarios involving multiple files.
17 papers · 0 benchmarks
SECOND (SEmantic Change detectiON Dataset)
SECOND is a well-annotated semantic change detection dataset.
17 papers · 1 benchmark
SMDD (Synthetic Face Morphing Attack Detection Development Dataset)
The Synthetic Morphing Attack Detection Development (SMDD) dataset is a synthetic-based MAD dataset, consisting of 25k bona-fide images generated using the StyleGAN2-ADA framework and 15k morphing attacks created from the bonafide samples…
17 papers · 0 benchmarks
SODA is a high-quality social dialogue dataset.
17 papers · 0 benchmarks
SOSD (Searching on Sorted Data)
SOSD is a collection of dataset to benchmark the lookup performance of learned indexes.
17 papers · 0 benchmarks
SPICE (Small-Molecule/Protein Interaction Chemical Energies)
SPICE is a collection of quantum mechanical data for training potential functions.
17 papers · 0 benchmarks
How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence.
17 papers · 2 benchmarks
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
Sachs (Sachs Protein Dataset)
Sachs dataset measures the expression level of different proteins and phospholipids in human cells.
17 papers · 0 benchmarks
SceneNet is a dataset of labelled synthetic indoor scenes.
17 papers · 0 benchmarks
Automated source code generation is currently a popular machine learning-based task.
17 papers · 0 benchmarks
The shiny folder contains 8 scenes with challenging view-dependent effects used in our paper.
17 papers · 0 benchmarks
- This dataset was originally introduced by [1] for soccer ball and player tracking from six synchronized videos.
17 papers · 1 benchmark
The Switchboard-1 Telephone Speech Corpus (LDC97S62) consists of approximately 260 hours of speech and was originally collected by Texas Instruments in 1990-1, under DARPA sponsorship.
17 papers · 1 benchmark
TID2013 is a dataset for image quality assessment that contains 25 reference images and 3000 distorted images (25 reference images x 24 types of distortions x 5 levels of distortions).
17 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
17 papers · 3 benchmarks
VIPL-HR database is a database for remote heart rate (HR) estimation from face videos under less-constrained situations.
17 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.