Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 190 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9073–9120 of 12,172
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis.
1 paper · 1 benchmark
M-Phasis (A Feature-Based Corpus of Hate Online)
A corpus of 9k German and French user comments collected from migration-related news articles.
1 paper · 0 benchmarks
M3AV (Multimodal, Multigenre, and Multipurpose Audio-Visual)
The M3AV (Multimodal, Multigenre, and Multipurpose Audio-Visual) is a novel dataset proposed for academic lectures¹.
1 paper · 0 benchmarks
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
In this project, we tried to make malaria detection easily possible at a low cost.
1 paper · 1 benchmark
MADDPG AND P2P-VFRL FOR MINIMIZING AOI IN NTN NETWORK UNDER CSI UNCERTAINTY
1 paper · 0 benchmarks
Characterising multimedia content with relevant, reliable and discriminating tags is vital for multimedia information retrieval.
1 paper · 0 benchmarks
MAI (Multi-scene Aerial Image)
MAI is a dataset for multi-scene recognition in single aerial images.
1 paper · 0 benchmarks
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
An annotated dataset of 4869 transient and 71207 non-transient object lightcurves built from the Catalina Real Time Transient Survey.
1 paper · 0 benchmarks
The rising interest in leveraging higher-order interactions present in complex systems has led to a surge in more expressive models exploiting high-order structures in the data, especially in topological deep learning (TDL), which designs…
1 paper · 0 benchmarks
MAOMaps is a dataset for evaluation of Visual SLAM, RGB-D SLAM and Map Merging algorithms.
1 paper · 0 benchmarks
The MAPLE benchmark constructed by us contains 20 datasets across 19 fields for scientific literature tagging.
1 paper · 0 benchmarks
MAQA (Multi-Answer Question Answering dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MARIO (Monitoring Age-related Macular Degeneration Progression In Optical Coherence Tomography)
MICCAI Challenge 2024
1 paper · 0 benchmarks
Analogical reasoning is fundamental to human cognition and holds an important place in various fields.
1 paper · 1 benchmark
MARS Map is a set of three dataset collected to evaluate the performance of mapping algorithms within a room and between rooms.
1 paper · 0 benchmarks
MARS dataset processed with our re-Detect and Link (DL) module.
1 paper · 0 benchmarks
MASC (Manually Annotated Sub-Corpus)
The Manually Annotated Sub-Corpus (MASC) consists of approximately 500,000 words of contemporary American English written and spoken data drawn from the Open American National Corpus (OANC).
1 paper · 0 benchmarks
MAST (Multi-Attributed Structured Text-to-face Dataset)
A new data consolidation called Multi-Attributed and Structured Text-to-face (MAST) dataset.
1 paper · 0 benchmarks
The MATHWELL Human Annotation Dataset contains 5,084 synthetic word problems and answers generated by MATHWELL, a reference-free educational grade school math word problem generator released in MATHWELL: Generating Educational Math Word…
1 paper · 0 benchmarks
The dataset contains 3 million attribute-value annotations across 1257 unique categories created from 2.2 million cleaned Amazon product profiles.
1 paper · 0 benchmarks
MAVEN-Arg is an advanced event argument extraction dataset, which offers three main advantages: A comprehensive schema covering 162 event types and 612 argument roles, all with expert-written definitions and examples.
1 paper · 0 benchmarks
Manually vAlidated Vq2a Examples fRom Image/Caption datasetS (MAVERICS) is a suite of test-only visual question answering datasets.
1 paper · 0 benchmarks
MAX-60K (Masked Autoencoder for X-ray Fluorescence 60K Dataset)
The dataset for masked autoencoder for X-ray fluorescence (XRF) is a following development after the dataset (Chao et al., 2022).
1 paper · 0 benchmarks
a dataset of reading pointer meter
1 paper · 0 benchmarks
A large-scale machine comprehension dataset (based on the COCO images and captions).
1 paper · 0 benchmarks
Here we release the dataset (MultiChannelGrid, abbreviated as MCGrid) used in our paper LIMUSE: LIGHTWEIGHT MULTI-MODAL SPEAKER EXTRACTION](https://arxiv.org/abs/2111.04063)).
1 paper · 0 benchmarks
MCiteBench is a benchmark to evaluate multimodal citation text generation in Multimodal Large Language Models (MLLMs).
1 paper · 0 benchmarks
A small-scale training set, which only contains 4K images.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
MEIS (M-mode Echocardiograms for Instance Segmentation)
MEIS comprises a total of 2,639 images in the size of 1024 × 768 toward two recording views (Aortic Valve (AV) and Left Ventricle (LV)) with 1,521 (747 in AV + 774 in LV) images for training and 1,118 (559 in AV + 559 in LV) for testing,…
1 paper · 1 benchmark
MENYO-20k is the first multi-domain parallel corpus with a special focus on clean orthography for Yorùbá--English with standardized train-test splits for benchmarking.
1 paper · 0 benchmarks
This dataset contains pre-processed versions of datasets introduced in prior works.
1 paper · 0 benchmarks
METABRIC (MOLECULAR TAXONOMY OF BREAST CANCER INTERNATIONAL CONSORTIUM)
https://ega-archive.org/studies/EGAS00000000083
1 paper · 0 benchmarks
METAR (Meteorological Terminal Aviation Routine Dataset)
Weather reports of 57 stations in the east coast.
1 paper · 0 benchmarks
METEOR is a complex traffic dataset which captures traffic patterns in unstructured scenarios in India.
1 paper · 0 benchmarks
METU-ALET is an image dataset for the detection of the tools in the wild.
1 paper · 0 benchmarks
METU-VIREF is a video referring expression dataset comprising of videos from VIRAT Ground and ILSVRC2015 VID datasets.
1 paper · 0 benchmarks
MF (Mathematical Formulas)
Mathematical dataset containing formulas based on the AMPS Khan dataset and the ARQMath dataset V1.3.
1 paper · 0 benchmarks
MF3QA (Medical Free Form Farsi Question Answering dataset)
real-world doctor-patient question- answering dataset cleaned manually and automatically
1 paper · 0 benchmarks
MF3QA_uncleaned (Medical Free Form Farsi Question Answering dataset (uncleaned))
real-world doctor-patient question- answering dataset
1 paper · 0 benchmarks
MFA (Many Faces of Anger)
The MFA (Many Faces of Anger) dataset includes 200 in-the-wild videos from North American and Persian cultures with fine-grained labels of: 'annoyed', 'anger', 'disgust', 'hatred' and 'furious' and 13 related emojis.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.