Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 191 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9121–9168 of 12,172
The MFH dataset is a multi-viewpoint fine-grained hand hygiene dataset.
1 paper · 0 benchmarks
MFSD (Masked Face Segmentation Dataset)
During the covid-19 era wearing face masks posed new challenges to face-related tasks, including facial recognition, face inpainting, expression recognition, and object removal.
1 paper · 1 benchmark
MFW+ is a benchmark dataset for masked face recognition and an extended version of MFW.
1 paper · 1 benchmark
MGPFD (multi-goal path finding dataset)
MGPFD is a dataset for multi-goal path finding problem, including a training dataset and a simulation dataset.
1 paper · 0 benchmarks
MH-FED (Meta Human Facial Expression Dataset)
This dataset provides a collection of 162K images and 70 Videos of Meta-Humans.
1 paper · 0 benchmarks
MHJ (Multi-Turn Human Jailbreaks)
We compile successful jailbreaks into the Multi-Turn Human Jailbreaks (MHJ) dataset, consisting of 2,912 prompts across 537 multi-turn conversations.
1 paper · 0 benchmarks
MHMD (Modern Historical Movies Dataset)
MHMD (Modern Historical Movies Dataset) is a dataset for old image colorization, built from historical movies.
1 paper · 0 benchmarks
Multi-Person Interaction Motion (MI-Motion) Dataset includes skeleton sequences of multiple individuals collected by motion capture systems and refined and synthesized using a game engine.
1 paper · 0 benchmarks
Source: Radar reflectivity data from the HydroMeteorological Service of Arpae (Emilia-Romagna, Italy).
1 paper · 0 benchmarks
The dataset from the study "A Fully Generative Motivational Interviewing Counsellor Chatbot for Moving Smokers Towards the Decision to Quit".
1 paper · 0 benchmarks
The dataset contains 11,913 frame pairs of urban driving footage with and without moving objects, synthetically generated with the CARLA simulator.
1 paper · 0 benchmarks
MICCAI'2015 Gland Segmentation Challenge Contest Dataset Welcome to the challenge on gland segmentation in histology images.
1 paper · 0 benchmarks
To assess a model’s ability to create microcontroller-driven electronic devices, we developed a benchmark, MICRO25, that includes 25 tasks intended for the common ARDUINO microcontroller ecosystem..
1 paper · 0 benchmarks
Intrinsic component extension of MIT Multi-Illumination Dataset proposed in the paper "Intrinsic Image Decomposition via Ordinal Shading", Chris Careaga and Yağız Aksoy, ACM Transactions on Graphics, 2023 Project Page | Paper | Video |…
1 paper · 0 benchmarks
Consists of manually annotated dangerous and non-dangerous Kiki challenge videos.
1 paper · 0 benchmarks
This dataset contains 4606 articles from 1996 to 2024 that were presented in MIE (Medical Informatics Europe Conference) conferences.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MILU (Multi-task Indic Language Understanding Benchmark)
Overview MILU (Multi-task Indic Language Understanding Benchmark) is a comprehensive evaluation dataset designed to assess the performance of Large Language Models (LLMs) across 11 Indic languages.
1 paper · 0 benchmarks
MIMI dataset (Multi-aspect Integrated Migration Indicators dataset)
Nowadays, new branches of research are proposing the use of non-traditional data sources for the study of migration trends in order to find an original methodology to answer open questions about cross-border human mobility.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
1083 cases from the MIMIC-CXR dataset.
1 paper · 0 benchmarks
MIMIC-IV ICD-9 contains 209,326 discharge summaries—free-text medical documents—annotated with ICD-9 diagnosis and procedure codes.
1 paper · 1 benchmark
Retrospectively collected medical data has the opportunity to improve patient care through knowledge discovery and algorithm development.
1 paper · 0 benchmarks
The MIMIC-IV-ICD10 dataset, featuring the top 50 most frequently occurring labels.
1 paper · 1 benchmark
The MIMIC-IV-ICD9 dataset, featuring the top 50 most frequently occurring labels.
1 paper · 1 benchmark
MIMIC-IV-Note (MIMIC-IV-Note: Deidentified free-text clinical notes)
The advent of large, open access text databases has driven advances in state-of-the-art model performance in natural language processing (NLP).
1 paper · 0 benchmarks
Brazilian Sign Language (Libras) data set with 20 signs for sign language and gesture recognition benchmark: - Acontecer (To happen) - Aluno (Student) - Amarelo (Yellow) - América (America) - Aproveitar (To enjoy) - Bala (Candy) - Banco…
1 paper · 1 benchmark
MINT (a Multi-modal Image and Narrative Text Dubbing Dataset)
Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape.
1 paper · 0 benchmarks
MIPD (Manipulation and Intention In a Novel Corpus of Polish Disinformation)
A novel corpus of 15,356 Polish web articles, including articles identified as containing disinformation.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MIR-ST500 Good for the following task: Singing transcription (singing pitch to music note conversion) Used in several papers published by Roger Jang's Lab
1 paper · 0 benchmarks
MIRACL-VC1 is a lip-reading dataset including both depth and color images.
1 paper · 1 benchmark
Minor Irrigation Structures Check-Dam Dataset is a public dataset annotated by domain experts using images from Google static map for instance segmentation and object detection tasks.
1 paper · 0 benchmarks
This database includes 22 half-hour ECG recordings of subjects who experienced episodes of sustained ventricular tachycardia, ventricular flutter, and ventricular fibrillation.
1 paper · 0 benchmarks
The data generated from this study are grouped into 3 main types: (1) participant demographic and clinical data, (2) sensor data from the different devices, as well as clinical scores and metadata related to the tasks performed, and (3)…
1 paper · 0 benchmarks
An example dataset of 110,000 question/query pairs across four WikiData domains.
1 paper · 0 benchmarks
This dataset is a supplement to the github repositiry (https://github.com/pfilonenko/MLforTwoSampleTesting) and paper addressed to solve the two-sample problem under right-censored observations using Machine Learning.
1 paper · 0 benchmarks
Logic synthesis is a challenging and widely-researched combinatorial optimization problem during integrated circuit (IC) design.
1 paper · 0 benchmarks
ML-CB (ML-CB: Machine Learning Canvas Block)
In this paper, we develop a new privacy enhancing tool: ML-CB—a means of using distinguishable pictorial information combined with underlying website source code to produce accurate and robust machine learning classifiers able to discern…
1 paper · 0 benchmarks
ML25m dataset which is pre-processed in TGN Style.
1 paper · 0 benchmarks
MLB (Mouse Lockbox Dataset)
This dataset provides high-resolution videos recorded from three perspectives with more than 110 hours of total playtime showing mice solving complex tasks.
1 paper · 0 benchmarks
MLDS (Machine Learning Datasets)
MLDS is a collection of thousands of trained neural networks labelled with the data used to train them.
1 paper · 0 benchmarks
MLFF-distill data (Hessian Labels and Specialized Subdatasets for Distilling ML Force Fields)
Contains the datasets and distillation labels which were used in our paper Amin, I., Raja, S., Krishnapriyan, A.S.
1 paper · 0 benchmarks
MlGesture is a dataset for hand gesture recognition tasks, recorded in a car with 5 different sensor types at two different viewpoints.
1 paper · 0 benchmarks
The Mauna Loa Seeing Study was performed by the EOL/Integrated Surface Flux System team, capturing surface meteorology and flux products at the Mauna Loa Observatory in Hawaii.
1 paper · 2 benchmarks
MLQuestions is a domain-adaptation dataset for the machine learning domain containing 50K unaligned passages and 35K unaligned questions, and 3K aligned passage and question pairs.
1 paper · 0 benchmarks
MLRegTest (A Benchmark for the Machine Learning of Regular Languages)
MLRegTest is a benchmark for sequence classification, containing training, development, and test sets from 1,800 regular languages.
1 paper · 0 benchmarks
MLS (Multiple Light Source)
The Multiple Light Source dataset (MLS) is a collection of 24 multiple object scenes each recorded under 18 multiple light source illumination scenarios.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.