Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 84 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 3985–4032 of 12,172
The MEDIA French corpus is dedicated to semantic extraction from speech in a context of human/machine dialogues.
6 papers · 0 benchmarks
MFQE v2 (Multi-Frame Quality Enhancement v2 Dataset)
A dataset for compressed video quality enhancement.
6 papers · 1 benchmark
Contains video clips shot with modern high-resolution mobile cameras, with strong projective distortions and with low lighting conditions.
6 papers · 0 benchmarks
MIT Traffic is a dataset for research on activity analysis and crowded scenes.
6 papers · 0 benchmarks
The Masked LFW (MLFW), based on Cross-Age LFW (CALFW) database, is built using a simple but effective tool that generates masked faces from unmasked faces automatically.
6 papers · 1 benchmark
MLe2 is a dataset for the evaluation of scene text end-to-end reading systems and all intermediate stages such as text detection, script identification and text recognition.
6 papers · 0 benchmarks
MMPTRACK (Multi-camera Multiple People Tracking Dataset)
Multi-camera Multiple People Tracking (MMPTRACK) dataset has about 9.6 hours of videos, with over half a million frame-wise annotations.
6 papers · 1 benchmark
A minimalist, low-memory, and low-compute alternative to classic deep learning benchmarks.
6 papers · 0 benchmarks
MOLD (Marathi Offensive Language Dataset)
MOLD is a Marathi dataset for offensive language identification
6 papers · 0 benchmarks
A large-scale video dataset for MOR in aerial videos.
6 papers · 0 benchmarks
The dataset contains a total of 27,558 cell images with equal instances of parasitized and uninfected cells.
6 papers · 2 benchmarks
The Malimg Dataset contains 9,339 malware byteplot images from 25 different families.
6 papers · 1 benchmark
The official dataset contains a training set (137 images), a validation set (4 images), and a testing set (10 images)
6 papers · 1 benchmark
MatterportLayout extends the Matterport3D dataset with general Manhattan layout annotations.
6 papers · 0 benchmarks
MedShapeNet contains over 100,000 medical shapes, including bones, organs, vessels, muscles, etc., as well as surgical instruments.
6 papers · 0 benchmarks
The MegaVeridicality Dataset is a collection of ordinal veridicality judgments as well as ordinal acceptability judgments for 773 clause-embedding verbs of English.
6 papers · 0 benchmarks
Introduces a new dataset, MELINDA, for Multimodal biomEdicaL experImeNt methoD clAssification.
6 papers · 0 benchmarks
Microsoft Research Social Media Conversation Corpus consists of 127M context-message-response triples from the Twitter FireHose, covering the 3-month period June 2012 through August 2012.
6 papers · 0 benchmarks
Mid-Air, The Montefiore Institute Dataset of Aerial Images and Records, is a multi-purpose synthetic dataset for low altitude drone flights.
6 papers · 2 benchmarks
The MidiCaps dataset [1] is a large-scale dataset of 168,385 midi music files with descriptive text captions, and a set of extracted musical features.
6 papers · 0 benchmarks
Generate high-quality 3D ground-truth shapes for reconstruction evaluation is extremely challenging because even 3D scanners can only generate pseudo ground-truth shapes with artefacts.
6 papers · 0 benchmarks
Modern Office-31 is a refurbished version of the commonly used Office-31 dataset.
6 papers · 0 benchmarks
MoocRadar is a fine-grained and multiaspect knowledge repository that consists of 2,513 exercises, 5,600 concepts, and 14,224 students’ 12,715,126 behavioral records for improving cognitive student modeling in MOOCs.
6 papers · 0 benchmarks
A quantitative benchmark for developing and understanding video of fill-in-the-blank question-answering dataset with over 300,000 examples, based on descriptive video annotations for the visually impaired.
6 papers · 0 benchmarks
MovieShots is a dataset to facilitate the shot type analysis in videos.
6 papers · 0 benchmarks
Moviescope is a large-scale dataset of 5,000 movies with corresponding video trailers, posters, plots and metadata.
6 papers · 0 benchmarks
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio.
6 papers · 0 benchmarks
MuSeRC (Russian Multi-Sentence Reading Comprehension)
We present a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences.
6 papers · 1 benchmark
We collect a large-scale synthetic dataset for robotic hands with Differentiable Force Closure(DFC).
6 papers · 0 benchmarks
The MusicBench dataset is a music audio-text pair dataset that was designed for text-to-music generation purpose and released along with Mustango text-to-music model.
6 papers · 1 benchmark
The generation of data-driven prognostics models requires the availability of datasets with run-to-failure trajectories.
6 papers · 1 benchmark
NCD (Natural-Color Dataset)
The Natural-Color Dataset (NCD) is an image colorization dataset where images are true to their colors.
6 papers · 0 benchmarks
NLU++ (NLLU++ : A Multi-Label, Slot-Rich, Generalisable Dataset for Natural Language Understanding in Task-Oriented Dialogue)
nlu++ is a dataset for natural language understanding (NLU) in task-oriented dialogue (ToD) systems, with the aim to provide a much more challenging evaluation environment for dialogue NLU models, up to date with the current application…
6 papers · 0 benchmarks
There are two versions of the NLmaps corpus.
6 papers · 0 benchmarks
Naamapadam is a Named Entity Recognition (NER) dataset for the 11 major Indian languages from two language families.
6 papers · 0 benchmarks
NorQuAD is the first Norwegian question answering dataset specifically designed for machine reading comprehension.
6 papers · 0 benchmarks
OG-MARL (Off-the-Grid MARL Datasets)
Diverse datasets for offline multi-agent reinforcement learning research.
6 papers · 0 benchmarks
OMICS (Open Mind Indoor Common Sense)
OMICS is an extensive collection of knowledge for indoor service robots gathered from internet users.
6 papers · 0 benchmarks
OOD-CV (Out Of Distribution Generalization in Computer Vision)
Enhancing the robustness of vision algorithms in real-world scenarios is challenging.
6 papers · 1 benchmark
OTTers is a dataset of human one-turn topic transitions.
6 papers · 0 benchmarks
The ObjectFolder Real dataset contains multisensory data collected from 100 real-world household objects.
6 papers · 0 benchmarks
OmniFlow is a synthetic omnidirectional human optical flow dataset.
6 papers · 0 benchmarks
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
1.0 Version of OpenEA benchmark datasets.
6 papers · 4 benchmarks
PGR (Phenotype-Gene Relations)
Phenotype-Gene Relations (PGR) is a corpus that consists of 1712 abstracts, 5676 human phenotype annotations, 13835 gene annotations, and 4283 relations.
6 papers · 2 benchmarks
PHASE (PHysically-grounded Abstract Social Events)
PHASE is a dataset of physically-grounded abstract social events, that resemble a wide range of real-life social interactions by including social concepts such as helping another agent.
6 papers · 0 benchmarks
PIE (Pedestrian Intention Estimation)
PIE is a new dataset for studying pedestrian behavior in traffic.
6 papers · 1 benchmark
PLABA (Plain Language Adaptation of Biomedical Abstracts)
Plain Language Adaptation of Biomedical Abstracts (PLABA) is a dataset designed for automatic adaptation that is both document- and sentence-aligned.
6 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.