Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 192 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9169–9216 of 12,172
The MLSum-it dataset is the translated version (Helsinki-NLP/opus-mt-es-it) of the spanish portion of MLSum, containing news articles taken from BBC/mundo.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MM-Eval (Modern Mongolian Evaluation)
Large language models (LLMs) excel in high-resource languages but face notable challenges in low-resource languages like Mongolian.
1 paper · 0 benchmarks
MM-Locate-News is a dataset for location estimation of news.
1 paper · 0 benchmarks
MMCOMPOSITION is a high-quality benchmark specifically designed to comprehensively evaluate the compositionality of pre-trained Vision-Language Models (VLMs) across three main dimensions—VL compositional perception, reasoning, and…
1 paper · 0 benchmarks
MMInstruct-GPT4V (MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity)
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: 1.
1 paper · 0 benchmarks
Mix of Minimal Optimal Sets (MMOS) of dataset has two advantages for two aspects, higher performance and lower construction costs on math reasoning.
1 paper · 0 benchmarks
MMPD Dataset is proposed in ECCV'2024 "When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset".
1 paper · 1 benchmark
MMSQL (Multi-Turn Multi-Type Text-to-SQL test suit)
A dataset for training and testing tin various problem types and multi-turn Q&A scenarios, including a training set, test set, and test scripts.
1 paper · 1 benchmark
MMTB (Multi-Mission Tool Bench)
Our test data has undergone five rounds of manual inspection and correction by five senior algorithm researcher with years of experience in NLP, CV, and LLM, taking about one month in total.
1 paper · 0 benchmarks
MMFace4D is a large-scale multi-modal 4D (3D sequence) face dataset consisting of 431 identities, 35,904 sequences, and 3.9 million frames MMFace4D has three appealing characteristics: 1) highly diversified subjects and corpus, 2)…
1 paper · 0 benchmarks
MNIST Multiview Datasets ======================== MNIST is a publicly available dataset consisting of 70, 000 images of handwritten digits distributed over ten classes.
1 paper · 0 benchmarks
MNIST dataset with included uncertainty.
1 paper · 0 benchmarks
MNIST-MIX is a multi-language handwritten digit recognition dataset.
1 paper · 0 benchmarks
The MNumGLUESub dataset is a multi-task arithmetic reasoning benchmark.
1 paper · 0 benchmarks
MOD20 is an action recognition dataset consisting of videos collected from YouTube and our own drone.
1 paper · 0 benchmarks
MODA dataset (Massive Online Data Annotation Spindle Dataset)
MODA is a large open-source dataset of high quality, human-scored sleep spindles (5342 spindles, from 180 subjects) that was produced by the Massive Online Data Annotation project.
1 paper · 1 benchmark
Structured atmospheric data for AI/ML Long-term, pre-processed, atmospheric datasets for use in Machine Learning/AI based forecasting.
1 paper · 0 benchmarks
MOET a dataset consists of gaze data from participants tracking specific objects, annotated with labels and bounding boxes, in crowded real-world videos, for training and evaluating attention decoding algorithms.
1 paper · 0 benchmarks
There are three MONK's problems.
1 paper · 0 benchmarks
This dataset was used in the paper 'Template-based Abstractive Microblog Opinion Summarisation' (to be published at TACL, 2022).
1 paper · 0 benchmarks
MOSAD (Mobile Sensing Human Activity Data Set)
MOSAD (Mobile Sensing Human Activity Data Set) is a multi-modal, annotated time series (TS) data set that contains 14 recordings of 9 triaxial smartphone sensor measurements (126 TS) from 6 human subjects performing (in part) 3 motion…
1 paper · 0 benchmarks
MOTIF (MOTIF: A Large Malware Reference Dataset with Ground Truth Family Labels)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MOViD-A is a video-based synthesized dataset.
1 paper · 0 benchmarks
Multi-Person 3D HumanPose Dataset (MP-3DHP) is a depth sensor-based dataset, which was constructed to facilitate the development of multi-person 3D pose estimation methods targeting real-world challenges.
1 paper · 0 benchmarks
MP-IDB (MP-IDB: The Malaria Parasite Image Database for Image Processing and Analysis)
MP-IDB comprises four species of Malaria parasites: Falciparum, Malariae, Ovale, Vivax.
1 paper · 4 benchmarks
The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations.
1 paper · 0 benchmarks
MPM-Verse (MPMVerse Physics Simulation Dataset)
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, elasticity, jelly, rigid collisions, and melting.
1 paper · 0 benchmarks
MPOSE2021 (MPOSE2021 Dataset for Short-time Human Action Recognition)
MPOSE2021, a dataset for real-time short-time HAR, suitable for both pose-based and RGB-based methodologies.
1 paper · 0 benchmarks
MPSGaze (Multi-Person Swap Gaze Dataset)
This is a synthetic dataset containing full images (instead of only cropped faces) that provides ground truth 3D gaze directions for multiple people in one image.
1 paper · 1 benchmark
MQUAKE (Multi-hop Question Answering for Knowledge Editing)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A new knowledge editing dataset removing the incorrect annotation mistakes in the prior dataset MQuAKE.
1 paper · 0 benchmarks
A new knowledge editing benchmark including the challenging instances collected from MQuAKE.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
1、 Competition name: The 2nd China Society of Image and Graphics (CSIG) Image and Graphics Technology Challenge: MRSpineSeg Challenge: Automated Multi-class Segmentation of Spinal Structures on Volumetric MR Images.
1 paper · 0 benchmarks
Whole-body, low-level control/manipulation demonstration dataset for ManiSkill-HAB.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is a challenging dataset from real courtrooms to predict the legal judgment in a reasonably encyclopedic manner by leveraging the genuine input of the case -- plaintiff's claims and court debate data, from which the case's facts are…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MSNER (Multilingual Spoken Named Entity Recognition)
This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish.
1 paper · 0 benchmarks
The MSRA-B dataset is a dataset for salient object detection.
1 paper · 0 benchmarks
MSRB (Marine Snow Removal Benchmarking)
MSRB is a benchmarking dataset for marine snow removal of underwater images.
1 paper · 0 benchmarks
MSVD-Indonesian is derived from the MSVD dataset, which is obtained with the help of a machine translation service.
1 paper · 3 benchmarks
WMVeID863 is captured with vehicles in motion with more challenges, such as motion blur, huge background changes, and especially intense flare degradation from car lamps, and sunlight.
1 paper · 0 benchmarks
Mathematical dataset containing mathematical texts, i.e., texts containing LaTeX formulas, based on the AMPS Khan dataset and the ARQMath dataset V1.3.
1 paper · 0 benchmarks
The MT40K dataset for predicting malware threat intelligence is a collection of 40,000 triples generated from 27,354 unique entities and 34 relations.
1 paper · 0 benchmarks
MT560 (MT560 - A Many-to-English Machine Translation Dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.