Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 102 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 4849–4896 of 12,172
This dataset is a benchmark for complex reasoning abilities in large language models, drawing on United Kingdom Linguistics Olympiad problems which cover a wide range of languages.
4 papers · 1 benchmark
The lipophilicity database refers to a collection of information related to the lipophilic properties of various molecules.
4 papers · 1 benchmark
This is a set of small programs with logic bombs.
4 papers · 0 benchmarks
The Lorenz dataset contains 100000 time-series with length 24.
4 papers · 1 benchmark
M³-VOS (M³-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation)
💡 Description A new benchmark, Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation (M³-VOS), to verify the ability of models to understand object phases, which consists of 479 high-resolution videos spanning over 10…
4 papers · 1 benchmark
The M5Product dataset is a large-scale multi-modal pre-training dataset with coarse and fine-grained annotations for E-products.
4 papers · 0 benchmarks
MA-52 (Micro-Action 52 dataset)
The MA-52 dataset provides the whole-body perspective including gestures, upper- and lower-limb movements, attempting to reveal comprehensive micro-action cues.
4 papers · 1 benchmark
MACSum a human-annotated summarization dataset for controlling mixed attributes.
4 papers · 0 benchmarks
MAS3K (MAS3K: An Open Dataset for Marine Animal Segmentation)
MAS3K contains a total of 3,103 images, where 1,588 are for camouflaged cases, 1,322 are for common cases, and 193 are underwater images in absence of marine animals.
4 papers · 1 benchmark
MCubeS (P) (Multimodal Material Segmentation Dataset)
Multimodal material segmentation (MCubeS) dataset contains 500 sets of images from 42 street scenes.
4 papers · 1 benchmark
Description The consists of 92 groups of video clips with 113, 918 high resolution frames taken by two drones and 63 groups of video clips with 145, 875 high resolution frames taken by three drones.
4 papers · 0 benchmarks
MENSA (Movie Scene Saliency Dataset)
MENSA: Movie Scene Saliency Dataset Dataset Summary The dataset, MENSA (Movie Scene Saliency Dataset) is from the paper "Select and Summarize: Scene Saliency for Movie Script Summarization", and consists of movie scripts and their…
4 papers · 1 benchmark
MERL-RAV (MERL Reannotation of AFLW with Visibility)
The MERL-RAV (MERL Reannotation of AFLW with Visibility) Dataset contains over 19,000 face images in a full range of head poses.
4 papers · 2 benchmarks
MESA (Multi-Ethnic Study of Atherosclerosis)
Multi-Ethnic Study of Atherosclerosis (MESA) is an NHLBI-sponsored 6-center collaborative longitudinal investigation of factors associated with the development of subclinical cardiovascular disease and the progression of subclinical to…
4 papers · 1 benchmark
The METU Trademark Dataset is a large dataset (the largest publicly available logo dataset as of 2014, and the largest one not requiring any preprocessing as of 2017), which is composed of more than 900K real logos belonging to real…
4 papers · 0 benchmarks
MEVID (Multi-view Extended Videos with Identities Dataset)
Multi-view Extended Videos with Identities dataset (MEVID) is a dataset for large-scale, video person re-identification (ReID) in the wild.
4 papers · 0 benchmarks
During the MILAN research project (MachIne Learning for AstroNomy), the research team uses Stellina observation stations to collect raw images of deep sky objects.
4 papers · 0 benchmarks
MIMIC-IV-ED is a large, freely available database of emergency department (ED) admissions at the Beth Israel Deaconess Medical Center between 2011 and 2019.
4 papers · 0 benchmarks
MJU-Waste is an RGBD waste object segmentation dataset that is made public to facilitate future research in this area.
4 papers · 1 benchmark
MLFP (Multispectral Latex Mask based Video Face Presentation Attack)
The MLFP dataset consists of face presentation attacks captured with seven 3D latex masks and three 2D print attacks.
4 papers · 1 benchmark
A new multilingual multi-aspect hate speech analysis dataset and use it to test the current state-of-the-art multilingual multitask learning approaches.
4 papers · 0 benchmarks
MMToM-QA (Multimodal Theory of Mind Question Answering)
MMToM-QA is the first multimodal benchmark to evaluate machine Theory of Mind (ToM), the ability to understand people's minds.
4 papers · 0 benchmarks
MN-DS (Multilabeled News Dataset)
Multilabeled News Dataset (MN-DS) is a dataset for news classification.
4 papers · 0 benchmarks
The MNIST Large Scale dataset is based on the classic MNIST dataset, but contains large scale variations up to a factor of 16.
4 papers · 1 benchmark
MOPRD, a multidisciplinary open peer review dataset consists of paper metadata, multiple version manuscripts, review comments, meta-reviews, author's rebuttal letters, and editorial decisions from 6578 papers.
4 papers · 0 benchmarks
MPV (Multi-Pose Virtual try on)
Consists of 37,723/14,360 person/clothes images, with the resolution of 256x192.
4 papers · 1 benchmark
Message Queuing Telemetry Transport (MQTT) protocol is one of the most used standards used in Internet of Things (IoT) machine to machine communication.
4 papers · 0 benchmarks
MSRVTT-CTN Dataset This dataset contains CTN annotations for the MSRVTT-CTN benchmark dataset in JSON format.
4 papers · 1 benchmark
This is a dataset for video deinterlacing problem.
4 papers · 1 benchmark
MSVD-CTN (MSVD Causal-Temporal Narrative)
MSVD-CTN Dataset This dataset contains CTN annotations for the MSVD-CTN benchmark dataset in JSON format.
4 papers · 1 benchmark
MS^2 (Multi-Document Summarization of Medical Studies)
MS^2 (Multi-Document Summarization of Medical Studies) is a dataset of over 470k documents and 20k summaries derived from the scientific literature.
4 papers · 1 benchmark
MTASS is an open-source dataset in which mixtures contain three types of audio signals.
4 papers · 0 benchmarks
MUAD (Multiple Uncertainties for Autonomous Driving)
The MUAD dataset (Multiple Uncertainties for Autonomous Driving), consisting of 10,413 realistic synthetic images with diverse adverse weather conditions (night, fog, rain, snow), out-of-distribution objects, and annotations for semantic…
4 papers · 0 benchmarks
MUSES offers 2500 multi-modal scenes, evenly distributed across various combinations of weather conditions (clear, fog, rain, and snow) and types of illumination (daytime, nighttime).
4 papers · 5 benchmarks
MassSpecGym (MassSpecGym: A benchmark for the discovery and identification of molecules)
MassSpecGym provides three challenges for benchmarking the discovery and identification of new molecules from MS/MS spectra: - 💥 De novo molecule generation (MS/MS spectrum → molecular structure) - ✨ Bonus chemical formulae challenge…
4 papers · 6 benchmarks
Med-EASi (Medical dataset for Elaborative and Abstractive Simplification), a uniquely crowdsourced and finely annotated dataset for supervised simplification of short medical texts.
4 papers · 0 benchmarks
The Medical Abstracts dataset contains 14,438 medical abstracts describing 5 different classes of patient conditions, with all of the dataset being annotated.
4 papers · 1 benchmark
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection This is MetaHate: a meta-collection of 36 hate speech datasets from social media comments.
4 papers · 0 benchmarks
Collected by leveraging background knowledge from a larger, more highly represented dialogue source.
4 papers · 0 benchmarks
To construct the MICROSOFT RESEARCH MULTIMODAL ALIGNED RECIPE CORPUS the authors first extract a large number of text and video recipes from the web.
4 papers · 0 benchmarks
The Middlebury 2001 is a stereo dataset of indoor scenes with multiple handcrafted layouts.
4 papers · 0 benchmarks
We generate epistemic reasoning problems using modal logic to target theory of mind (tom) in natural language processing models.
4 papers · 0 benchmarks
MIXSET comprises a total of 3.6k mixtext instances.
4 papers · 1 benchmark
The MoCapAct dataset contains training data and models for humanoid locomotion research.
4 papers · 0 benchmarks
Time Series Forecasting Repository containing datasets of related time series for global forecasting.
4 papers · 0 benchmarks
MovingFashion is a dataset for video-to-shop, the task of retrieving clothes which are worn in social media videos.
4 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.