Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 62 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 2929–2976 of 3,998

MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
MediConfusion is a challenging medical Visual Question Answering (VQA) benchmark dataset, that probes the failure modes of medical Multimodal Large Language Models (MLLMs) from a vision perspective.
1 paper · 0 benchmarks
This dataset contains demographic and personal health information for individuals, along with the corresponding medical insurance charges billed to them.
1 paper · 1 benchmark
A medical Wiki paralell corpus for medical text simplification.
1 paper · 0 benchmarks
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Collecting data with a HIKVISION USB Camera DS-E11, we build a dataset called MentalHAD with four abnormal actions (climbing walls, hitting windows, climbing, and hitting) and six normal actions (crouching, standing, sitting, hand waving,…
1 paper · 0 benchmarks
MerRec (MerRec Recommendation Dataset)
A large scale, C2C marketplace e-commerce dataset.
1 paper · 0 benchmarks
We introduce a new task of rephrasing for amore natural virtual assistant.
1 paper · 0 benchmarks
This data adds textual meta-infomation data to two existing corpora for cross language information retrieval: BoostCLIR, and the Large Scale CLIR Dataset (wiki-clir).
1 paper · 0 benchmarks
MetaEval is a collection of 101 NLP tasks.
1 paper · 0 benchmarks
Metric-Type of Numerical Tables is a dataset extracted from scientific papers (ACL anthology website) consisting of header tables, captions, and metric-types.
1 paper · 0 benchmarks
MiMIC (Multi-Modal Indian Earnings Calls Dataset)
Predicting stock market prices following corporate earnings calls remains a significant challenge for investors and researchers alike, requiring innovative approaches that can process diverse information sources.
1 paper · 0 benchmarks
MiST (Modals In Scientific Text) is a dataset containing 3737 modal instances in five scientific domains annotated for their semantic, pragmatic, or rhetorical function.
1 paper · 0 benchmarks
Microscopy Image Dataset of Pulmonary Vascular Changes (Microscopy Image Dataset for Deep Learning-Based Quantitative Assessment of Pulmonary Vascular Changes)
Pulmonary hypertension (PH) is a syndrome complex that accompanies a number of diseases of different etiologies, associated with basic mechanisms of structural and functional changes of the pulmonary circulation vessels and revealed…
1 paper · 0 benchmarks
Milling Data Set (UC Berkeley Milling Data Set)
Experiments on a metal milling machine for different speeds, feeds, and depth of cut.
1 paper · 0 benchmarks
MindReader is a novel dataset providing explicit user ratings over a knowledge graph within the movie domain.
1 paper · 0 benchmarks
MineralImage5k (Benchmark for 5k raw mineral species recognition)
We present a comprehensive dataset comprising a vast collection of raw mineral samples for the purpose of mineral recognition.
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
MoNuSAC (MoNuSAC 2020)
Different types of cells play a vital role in the initiation, development, invasion, metastasis and therapeutic response of tumors of various organs.
1 paper · 1 benchmark
MoToMQA (Multi-Order Theory of Mind Question & Answer)
The MoToMQA (Multi-Order Theory of Mind Question & Answer) benchmark is a test suite introduced to examine the extent to which large language models (LLMs) have developed higher-order theory of mind (ToM); the human ability to reason about…
1 paper · 0 benchmarks
https://huggingface.co/datasets/OpenDFM/MobA-MobBench
1 paper · 0 benchmarks
The Modified Swiss Dwellings (MSD) dataset is an ML-ready dataset for floor plan generation and analysis at building-level scale.
1 paper · 0 benchmarks
We present MoleculeCLA: a large-scale dataset consisting of approximately 140,000 small molecules derived from computational ligand-target binding analysis, providing nine properties that cover chemical, physical, and biological aspects.
1 paper · 0 benchmarks
Mon(IoT)r Testbed The Mon(IoT)r Testbed is the traffic capture software developed for the Mon(IoT)r Lab.
1 paper · 0 benchmarks
This dataset of medical misinformation was collected and is published by Kempelen Institute of Intelligent Technologies (KInIT).
1 paper · 0 benchmarks
Dataset can be used by anyone who is interested to perform morphological classification of galaxies.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is a dataset of robot motions based on physics simulations.
1 paper · 0 benchmarks
Movie Reviews (Movie Review Polarity Dataset Enriched with "Annotator Rationales")
This dataset is based on the movie review polarity dataset (v2.0) collected and maintained by Bo Pang and Lillian Lee.
1 paper · 0 benchmarks
MovieCLIP is a movie-centric taxonomy of 179 scene labels derived from movie scripts and auxiliary web-based video datasets designed for visual scene recognition.
1 paper · 0 benchmarks
Mpm-Verse-Large (MPMVerse Physics Simulation Dataset)
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, jelly, and rigid collisions.
1 paper · 0 benchmarks
MuDoCo_QueryRewrite (The MuDoCo dataset with Query Rewrite Annotations)
Given an ongoing dialogue between a user and a dialogue assistant, for the user query, the model is required to predict both coreference links between the query and the dialogue context, and the self-contained rewritten user query that is…
1 paper · 0 benchmarks
MuLMS (Multi-Layer Materials Science)
The Multi-Layer Materials Science corpus (MuLMS) consists of 50 documents (licensed CC BY) from the materials science domain, spanning across the following 7 subareas: "Electrolysis", "Graphene", "Polymer Electrolyte Fuel Cell (PEMFC)",…
1 paper · 0 benchmarks
MuSoHu (Toward human-like social robot navigation: A large-scale, multi-modal, social human navigation dataset)
A large-scale, egocentric, multimodal, and context-aware dataset of human demonstrations of social navigation.
1 paper · 0 benchmarks
Dataset Description The dataset used in this study comprises bug reports extracted from the Visual Studio Code GitHub repository, specifically focusing on those labeled with the english-please tag.
1 paper · 1 benchmark
This dataset was created as part of the Master's thesis titled "Multi-Class Depression Detection Through Tweets Using Artificial Intelligence." It contains tweets labeled for five types of depression (Bipolar, Major, Psychotic, Atypical,…
1 paper · 0 benchmarks
This dataset is a multi-labelled SMILES odor dataset with 138 odor descriptors.
1 paper · 1 benchmark
The Multi-domain Image Characteristic Dataset consists of thousands of images sourced from the internet.
1 paper · 0 benchmarks
Multi-template MRI mouse brain atlas (Multi-template MRI mouse brain atlas for both in vivo and ex vivo analysis)
Mouse Brain MRI atlas (both in-vivo and ex-vivo) (repository relocated from the original webpage) List of atlases - FVBNCrl: Brain MRI atlas of the wild-type FVBNCrl mouse strain (used as the background strain for the rTg4510 which is a…
1 paper · 0 benchmarks
MultiCite is a dataset of 12,653 citation contexts from over 1,200 computational linguistics papers used for Citation context analysis (CCA).
1 paper · 0 benchmarks
MultiRefKGC (multi-reference KGC)
MultiRefKGC is a dataset created from conversations from Reddit designed for Knowledge-Grounded Dialogue Generation tasks.
1 paper · 0 benchmarks
MultiSenseBadminton (MultiSenseBadminton: Wearable Sensor–Based Biomechanical Dataset for Evaluation of Badminton Performance)
The sports industry is witnessing an increasing trend of utilizing multiple synchronized sensors for player data collection, enabling personalized training systems with multi-perspective real-time feedback.
1 paper · 0 benchmarks
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/7845311#.ZK-jty9BxhE
1 paper · 0 benchmarks
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/7845361#.ZK-k7y9BxhE
1 paper · 0 benchmarks
A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/8119042#.ZK-jJC9BxhE
1 paper · 0 benchmarks
This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.