Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 61 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2881–2928 of 3,998
MMCOMPOSITION is a high-quality benchmark specifically designed to comprehensively evaluate the compositionality of pre-trained Vision-Language Models (VLMs) across three main dimensions—VL compositional perception, reasoning, and…
1 paper · 0 benchmarks
MMInstruct-GPT4V (MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity)
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: 1.
1 paper · 0 benchmarks
Mix of Minimal Optimal Sets (MMOS) of dataset has two advantages for two aspects, higher performance and lower construction costs on math reasoning.
1 paper · 0 benchmarks
MMPD Dataset is proposed in ECCV'2024 "When Pedestrian Detection Meets Multi-Modal Learning: Generalist Model and Benchmark Dataset".
1 paper · 1 benchmark
MMSQL (Multi-Turn Multi-Type Text-to-SQL test suit)
A dataset for training and testing tin various problem types and multi-turn Q&A scenarios, including a training set, test set, and test scripts.
1 paper · 1 benchmark
MMTB (Multi-Mission Tool Bench)
Our test data has undergone five rounds of manual inspection and correction by five senior algorithm researcher with years of experience in NLP, CV, and LLM, taking about one month in total.
1 paper · 0 benchmarks
MNIST Multiview Datasets ======================== MNIST is a publicly available dataset consisting of 70, 000 images of handwritten digits distributed over ten classes.
1 paper · 0 benchmarks
MODA dataset (Massive Online Data Annotation Spindle Dataset)
MODA is a large open-source dataset of high quality, human-scored sleep spindles (5342 spindles, from 180 subjects) that was produced by the Massive Online Data Annotation project.
1 paper · 1 benchmark
Structured atmospheric data for AI/ML Long-term, pre-processed, atmospheric datasets for use in Machine Learning/AI based forecasting.
1 paper · 0 benchmarks
This dataset was used in the paper 'Template-based Abstractive Microblog Opinion Summarisation' (to be published at TACL, 2022).
1 paper · 0 benchmarks
MOSAD (Mobile Sensing Human Activity Data Set)
MOSAD (Mobile Sensing Human Activity Data Set) is a multi-modal, annotated time series (TS) data set that contains 14 recordings of 9 triaxial smartphone sensor measurements (126 TS) from 6 human subjects performing (in part) 3 motion…
1 paper · 0 benchmarks
MP-IDB (MP-IDB: The Malaria Parasite Image Database for Image Processing and Analysis)
MP-IDB comprises four species of Malaria parasites: Falciparum, Malariae, Ovale, Vivax.
1 paper · 4 benchmarks
The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations.
1 paper · 0 benchmarks
MPM-Verse (MPMVerse Physics Simulation Dataset)
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, elasticity, jelly, rigid collisions, and melting.
1 paper · 0 benchmarks
MPOSE2021 (MPOSE2021 Dataset for Short-time Human Action Recognition)
MPOSE2021, a dataset for real-time short-time HAR, suitable for both pose-based and RGB-based methodologies.
1 paper · 0 benchmarks
1、 Competition name: The 2nd China Society of Image and Graphics (CSIG) Image and Graphics Technology Challenge: MRSpineSeg Challenge: Automated Multi-class Segmentation of Spinal Structures on Volumetric MR Images.
1 paper · 0 benchmarks
This is a challenging dataset from real courtrooms to predict the legal judgment in a reasonably encyclopedic manner by leveraging the genuine input of the case -- plaintiff's claims and court debate data, from which the case's facts are…
1 paper · 0 benchmarks
Mathematical dataset containing mathematical texts, i.e., texts containing LaTeX formulas, based on the AMPS Khan dataset and the ARQMath dataset V1.3.
1 paper · 0 benchmarks
The MT40K dataset for predicting malware threat intelligence is a collection of 40,000 triples generated from 27,354 unique entities and 34 relations.
1 paper · 0 benchmarks
MTA-KDD'19 (Malware Traffic Analysis Knowledge Dataset 2019)
Malware Traffic Analysis Knowledge Dataset 2019 (MTA-KDD'19) is an updated and refined dataset specifically tailored to train and evaluate machine learning based malware traffic analysis algorithms.
1 paper · 0 benchmarks
MVALUE (Multilingual human VALUE dataset)
Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and…
1 paper · 0 benchmarks
MVTec-FS (MVTec few-shot detection and classfication dataset)
The MVTec-FS dataset is a refined version of the MVTec AD dataset, designed for few-shot learning research.
1 paper · 0 benchmarks
If you plan to test your method on our road network, you can find road network files in env\map.
1 paper · 0 benchmarks
This dataset is used to train and evaluate models for the detection of machine-paraphrased text.
1 paper · 0 benchmarks
Dataset introduction There are four dimension in MBTI.
1 paper · 0 benchmarks
MalVis (MalVis: A Large-Scale Android Malware Visualization Dataset and Framework for Improved Classification)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MalayalamMixSentiment is a Sentiment Analysis Dataset for Code-Mixed Malayalam-English.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MapEval contains 700 question-answer pairs.
1 paper · 0 benchmarks
MapEval-Textual contains 300 question-answer pairs.
1 paper · 1 benchmark
MapEval-Textual contains 300 context-question-answer triplets.
1 paper · 1 benchmark
MapEval-Visual contains 400 image-question-answer triplets.
1 paper · 1 benchmark
This dataset was developed within an analysis of research data generated and managed within the University of Bologna, with respect to the differences and commonalities between disciplines and potential challenges for institutional data…
1 paper · 0 benchmarks
This dataset accompanies the ICWSM 2022 paper "Mapping Topics in 100,000 Real-Life Moral Dilemmas".
1 paper · 0 benchmarks
It contains grayscale mono and stereo images (NavCam and LocCam) from laboratory tests performed by a prototype rover on a martian-like testbed.
1 paper · 0 benchmarks
This dataset is a collection of marxist fragments mixed and cut randomly from the Marxist archive (marxists.org).
1 paper · 0 benchmarks
MatSeg (Dataset for Zero-Shot Material States Segmentation)
MatSeg Dataset for Zero-Shot Material States Segmentation: The dataset contains large-scale synthetic images for training data and highly diverse real-world image benchmarks for testing.
1 paper · 0 benchmarks
MatSim (MatSim dataset for materials similarity recognition from images)
MatSim is a synthetic dataset, and natural image benchmark for computer vision-based recognition of similarities and transitions between materials and textures, focusing on identifying any material under any conditions using one or a few…
1 paper · 0 benchmarks
pymatgencodeqa benchmark: qabenchmark/generatedqa/generationresultscode.json, which consists of 34,621 QA pairs.
1 paper · 0 benchmarks
Matador (Matador: A Material Image Dataset)
The Matador dataset is a material image dataset with hierarchical labels.
1 paper · 0 benchmarks
MathEquiv (mathematical statement equivalence)
MathEquiv dataset is accompanied to EquivPruner .
1 paper · 0 benchmarks
Mathematical dataset based on 71 famous mathematical identities.
1 paper · 0 benchmarks
McQueen dataset contains 15k visual conversations and over 80k queries where each one is associated with a fully-specified rewrite version.
1 paper · 0 benchmarks
MeSHup (A Corpus for Full Text Biomedical Document Indexing)
Contains 1,342,667 full text articles in English, together with the associated MeSH labels and metadata, authors, and publication venues that are collected from the MEDLINE database.
1 paper · 0 benchmarks
A new in-context visual question answering dataset encompassing interleaved image and EHR data derived from MIMIC-IV and MIMIC-CXR-JPG databases.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The MedVidCL dataset contains a collection of 6, 617 videos annotated into ‘medical instructional’, ‘medical non-instructional' and ‘non-medical’ classes.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.