Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 245 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 11713–11760 of 12,172
MEFB (benchmark of multi-exposure image fusion)
The MEFB consists of a test set of 100 image pairs.
0 papers · 0 benchmarks
The dataset includes recordings from 10 different users teaching the robot different common kitchen objects, that consists of synchronized recordings from three cameras and a microphone mounted on the robot: An RGB-d camera covers the user…
0 papers · 0 benchmarks
The MICCAI iSEG dataset was described in "https://iseg2017.web.unc.edu/how-to-cite/", that has a total of 10 images, including T1-1 through T1-10, T2-1 through T2-10, and a ground truth for the training set.
0 papers · 0 benchmarks
This dataset contains 4606 articles from 1996 to 2024 that were presented in MIE (Medical Informatics Europe Conference) conferences.
0 papers · 0 benchmarks
The MIM-GOLD-NER dataset is an Icelandic named entity (NE) corpus.
0 papers · 1 benchmark
MIMIC Meme Dataset (Misogyny Identification in Multimodal Internet Content in Hindi-English Code-Mix Language)
This dataset endeavors to fill the research void by presenting a meticulously curated collection of misogynistic memes in a code-mixed language of Hindi and English.
0 papers · 0 benchmarks
MINDS-14 is a dataset designed for the intent detection task with spoken data.
0 papers · 0 benchmarks
This is the standard dataset for solving the TSP problem with uniformly distributed points using machine learning, covering six scales: 50, 100, 200, 500, 1K, and 10K.
0 papers · 0 benchmarks
Machine Learning for Two-Sample Testing under Right-Censored Data: A Simulation Study - Petr PHILONENKO, Ph.D.
0 papers · 0 benchmarks
IQ testing has served as a foundational methodology for evaluating human cognitive capabilities, deliberately decoupling assessment from linguistic background, language proficiency, or domain-specific knowledge to isolate core competencies…
0 papers · 0 benchmarks
We introduce MMKE-Bench, a benchmark designed to evaluate the ability of LMMs to edit visual knowledge in real-world scenarios.
0 papers · 0 benchmarks
MNAD (Moroccan News Articles Dataset)
About the MNAD Dataset The MNAD corpus is a collection of over 1 million Moroccan news articles written in modern Arabic language.
0 papers · 0 benchmarks
MO7 dataset consists of 50,000 images with over 900 unique objects and over 18 classes.
0 papers · 0 benchmarks
The MOBIO database consists of bi-modal (audio and video) data taken from 152 people.
0 papers · 0 benchmarks
The MS-EVS Dataset is the first large-scale event-based dataset for face detection.
0 papers · 0 benchmarks
List of ontologies in the domain of Materials Science and Engineering.
0 papers · 0 benchmarks
Article: A novel hierarchical model based on different emotion induction modalities for EEG emotion recognition
0 papers · 0 benchmarks
MVP-24K (Multi-grained Vehicle Parsing dataset)
Multi-grained Vehicle Parsing (MVP) is a large-scale dataset for semantic analysis of vehicles in the wild, which has several featured properties.
0 papers · 0 benchmarks
The MVTec Industrial 3D Object Detection Dataset (MVTec ITODD), introduced by Bertram Drost, Markus Ulrich, Paul Bergmann, and Carsten Steger from MVTec Software GmbH, is a valuable resource for 3D object detection and pose estimation in…
0 papers · 0 benchmarks
Mac-Morpho is a corpus of Brazilian Portuguese texts annotated with part-of-speech tags.
0 papers · 0 benchmarks
A dataset of female face images assembled for studying the impact of makeup on face recognition.
0 papers · 0 benchmarks
Multi-domain Image Editing Benchmark
0 papers · 0 benchmarks
MapAI: Precision in Building Segmentation Dataset The dataset comprises 7500 training images and 1500 validation images from Denmark.
0 papers · 0 benchmarks
This dataset was developed within an analysis of research data generated and managed within the University of Bologna, with respect to the differences and commonalities between disciplines and potential challenges for institutional data…
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 7000+ original Masks images captured and crowdsourced from over 1200+ urban and rural areas, where each image is manually reviewed and verified by computer vision professionals at DC…
0 papers · 0 benchmarks
MathEval is a benchmark dedicated to a comprehensive evaluation of the mathematical capabilities of large models.
0 papers · 0 benchmarks
The task aims to measure the capability of models to predict the shape of the result of a chain of matrix manipulations, given the inputs' shapes.
0 papers · 0 benchmarks
Media-Text (MediaText: a media industry-based dataset for scene text detetcion)
Media-Text dataset comprising images of banners, posters, covers and another images characterised for media industry.
0 papers · 0 benchmarks
MedleyDB 2.0 is a superset of the MedleyDB – a dataset of annotated, royalty-free multitrack recordings.
0 papers · 0 benchmarks
The MegaIntensionality dataset is a part of the MegaAttitude project.
0 papers · 0 benchmarks
The MegaOrientation Dataset is a linguistic resource that consists of ordinal acceptability judgments for 898 clause-embedding verbs of English with a variety of nonfinite subordinate clause structures.
0 papers · 0 benchmarks
Data Acquisition EEG and NIRS data was collected in an ordinary bright room.
0 papers · 0 benchmarks
This dataset comprises micro-ultrasound scans and human prostate annotations of 75 patients who underwent micro-ultrasound guided prostate biopsy at the University of Florida.
0 papers · 0 benchmarks
The MIVIA audio events data set is composed of a total of 6000 events for surveillance applications, namely glass breaking, gun shots and screams.
0 papers · 0 benchmarks
Mixing Secrets is an instrument recognition dataset containing 258 multi-track recordings sourced from the Mixing Secrets for The Small Studio website.
0 papers · 0 benchmarks
A large ground truth training corpus of top-down fisheye images.
0 papers · 0 benchmarks
MoCap (CMU Graphics Lab Motion Capture Database)
Collection of various motion capture recordings (walking, dancing, sports, and others) performed by over 140 subjects.
0 papers · 0 benchmarks
This dataset is collected by DataCluster Labs, India.
0 papers · 0 benchmarks
Moh (SeyedMohammad Kashani)
We introduce an open-source physical-layer dataset of Bluetooth Low Energy (BLE) IoT sensor devices recorded in an anechoic chamber using USRP x310.
0 papers · 0 benchmarks
The dataset comprises time-series data capturing distinct periodic motions (gaits) of an energetically conservative one-legged hopper.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.