Home › Datasets
Datasets
archive 2025-07-28
12,172 datasets listed, ordered by the archive's paper count. Page 194 of 254: 48 shown of 12,172.
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter a dataset can carry several tags, so counts overlap
Modality 39
Task 500 shown of 3,717, by dataset count
Language 367
All datasets 9265–9312 of 12,172
MatSim (MatSim dataset for materials similarity recognition from images)
MatSim is a synthetic dataset, and natural image benchmark for computer vision-based recognition of similarities and transitions between materials and textures, focusing on identifying any material under any conditions using one or a few…
1 paper · 0 benchmarks
pymatgencodeqa benchmark: qabenchmark/generatedqa/generationresultscode.json, which consists of 34,621 QA pairs.
1 paper · 0 benchmarks
Matador (Matador: A Material Image Dataset)
The Matador dataset is a material image dataset with hierarchical labels.
1 paper · 0 benchmarks
MathEquiv (mathematical statement equivalence)
MathEquiv dataset is accompanied to EquivPruner .
1 paper · 0 benchmarks
Existing arithmetic benchmarks have a limited number of multiple-choice questions.
1 paper · 1 benchmark
Existing arithmetic benchmarks have a limited number of True-or-False questions.
1 paper · 1 benchmark
Mathematical dataset based on 71 famous mathematical identities.
1 paper · 0 benchmarks
MatriVasha the largest dataset of handwritten Bangla compound characters for research on handwritten Bangla compound character recognition.
1 paper · 0 benchmarks
The dataset consists of random electromagnetic scatterers and their associated fields when illuminated by a 1000nm plane-wave illumination.
1 paper · 0 benchmarks
McQueen dataset contains 15k visual conversations and over 80k queries where each one is associated with a fully-specified rewrite version.
1 paper · 0 benchmarks
MeSHup (A Corpus for Full Text Biomedical Document Indexing)
Contains 1,342,667 full text articles in English, together with the associated MeSH labels and metadata, authors, and publication venues that are collected from the MEDLINE database.
1 paper · 0 benchmarks
The Mechanical MNIST – Distribution Shift dataset contains the results of finite element simulation of heterogeneous material subject to large deformation due to equibiaxial extension at a fixed boundary displacement of d = 7.0.
1 paper · 0 benchmarks
This repository contains data for a research project involving graph neural networks (GNNs) applied to mechanical metamaterials and their deformations.
1 paper · 0 benchmarks
Novel benchmark adapted from the MedQA, with confounding statements introduced within the question regarding an irrelevant third party.
1 paper · 0 benchmarks
Novel benchmark adapted from the MedQA, with confounding statements introduced within the question regarding an irrelevant clinical term used in a nonclinical context (e.g., The patient's Zodiac sign is Cancer).
1 paper · 0 benchmarks
The MedLEA package provides morphological and structural features of 471 medicinal plant leaves and 1099 leaf images of 31 species and 29-45 images per species.
1 paper · 0 benchmarks
MedLFQA (Medical Long-form Question Answering)
MedLFQA is reconstructed by reformulating the current four biomedical long-form question-answering benchmark datasets: LiveQA, MedicationQA, HealthsearchQA, and K-QA.
1 paper · 0 benchmarks
A new in-context visual question answering dataset encompassing interleaved image and EHR data derived from MIMIC-IV and MIMIC-CXR-JPG databases.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A comprehensive Turkish dataset for question-answering tasks in medical domain
1 paper · 1 benchmark
The MedVidCL dataset contains a collection of 6, 617 videos annotated into ‘medical instructional’, ‘medical non-instructional' and ‘non-medical’ classes.
1 paper · 0 benchmarks
Mechanistic cardiac electrophysiology models allow for personalized simulations of the electrical activity in the heart and the ensuing electrocardiogram (ECG) on the body surface.
1 paper · 0 benchmarks
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
MediConfusion is a challenging medical Visual Question Answering (VQA) benchmark dataset, that probes the failure modes of medical Multimodal Large Language Models (MLLMs) from a vision perspective.
1 paper · 0 benchmarks
Mediapi-RGB is a bilingual corpus of French Sign Language (LSF) and written French in the form of subtitled videos, accompanied by complementary data (various representations, segmentation, vocabulary, etc.).
1 paper · 1 benchmark
Original dataset for "HIGH PRECISION MEDICINE BOTTLES VISION ONLINE INSPECTION SYSTEM AND CLASSIFICATION BASED ON MULTI-FEATURES AND ENSEMBLE LEARNING VIA INDEPENDENCE TEST"
1 paper · 0 benchmarks
Medical Case Report Corpus is a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library.
1 paper · 0 benchmarks
This dataset contains demographic and personal health information for individuals, along with the corresponding medical insurance charges billed to them.
1 paper · 1 benchmark
A medical Wiki paralell corpus for medical text simplification.
1 paper · 0 benchmarks
A dataset called Medley2K that consists of 2,000 medleys and 7,712 labeled transitions.
1 paper · 0 benchmarks
A large-scale reference dataset for bioacoustics.
1 paper · 1 benchmark
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
The MegaAcceptability dataset is a collection of ordinal acceptability judgments for clause-embedding verbs of English in various surface-syntactic frames and matrix tenses.
1 paper · 0 benchmarks
we introduce a large-scale and diverse symbolic melody dataset called MelodyNet that contains more than 0.4 million melody pieces extracted from approximately 1.6 million songs.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The MemoTrap dataset is a diagnostic dataset designed to assess whether language models (LMs) fall into memorization traps.
1 paper · 0 benchmarks
Collecting data with a HIKVISION USB Camera DS-E11, we build a dataset called MentalHAD with four abnormal actions (climbing walls, hitting windows, climbing, and hitting) and six normal actions (crouching, standing, sitting, hand waving,…
1 paper · 0 benchmarks
MerRec (MerRec Recommendation Dataset)
A large scale, C2C marketplace e-commerce dataset.
1 paper · 0 benchmarks
MeshFLeet (eshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain Specific Generative Modeling Resources)
MeshFleet is a filtered and annotated dataset of High Quality vehicles derived from Objaverse XL.
1 paper · 0 benchmarks
We introduce a new task of rephrasing for amore natural virtual assistant.
1 paper · 0 benchmarks
MessyTable features a large number of scenes with messy tables captured from multiple camera views.
1 paper · 0 benchmarks
Meta Omnium is a dataset-of-datasets spanning multiple vision tasks including recognition, keypoint localization, semantic segmentation and regression.
1 paper · 0 benchmarks
This data adds textual meta-infomation data to two existing corpora for cross language information retrieval: BoostCLIR, and the Large Scale CLIR Dataset (wiki-clir).
1 paper · 0 benchmarks
MetaEval is a collection of 101 NLP tasks.
1 paper · 0 benchmarks
There has been increasing interest in smart factories powered by robotics systems to tackle repetitive, laborious tasks.
1 paper · 0 benchmarks
There has been increasing interest in smart factories powered by robotics systems to tackle repetitive, laborious tasks.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.