Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 60 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2833–2880 of 3,998
The Linguistic Benchmark (JSON), consisting of 30 questions was developed to be easy for human adults to answer but challenging for LLMs.
1 paper · 0 benchmarks
The LinkedResults dataset contains around 1,600 results capturing performance of machine learning models from tables of 239 papers.
1 paper · 0 benchmarks
This is a dataset of 3 English books which do not contain the letter "e" in them.
1 paper · 1 benchmark
CSV file with a list of all examined OWL reasoners.
1 paper · 0 benchmarks
A listwise multi-response dataset for human preferences alignment.
1 paper · 0 benchmarks
An open-source online generative dictionary that takes a word and context containing the word as input and automatically generates a definition as output.
1 paper · 0 benchmarks
Liver-US (Liver Ultrasound Dataset for Medical Image Classification)
The Liver-US dataset is a comprehensive collection of high-quality ultrasound images of the liver, including both normal and abnormal cases.
1 paper · 1 benchmark
This dataset was collected during a LoRaWAN measurement campaign in a multi-room indoor office environment at the University of Siegen, Germany.
1 paper · 0 benchmarks
LocoVR (LocoVR: Multiuser Indoor Locomotion Dataset in Virtual Reality)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A 160B bilingual long-text dataset with 3 categories: holistic, aggregated and chaotic long texts.
1 paper · 0 benchmarks
LoT-insts contains over 25k classes whose frequencies are naturally long-tail distributed.
1 paper · 2 benchmarks
This dataset contains simulated synthetic particle decays, simulated using the PhaseSpace library.
1 paper · 0 benchmarks
The M-AILABS Speech Dataset is the first large dataset that we are providing free-of-charge, freely usable as training data for speech recognition and speech synthesis.
1 paper · 1 benchmark
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
In this project, we tried to make malaria detection easily possible at a low cost.
1 paper · 1 benchmark
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
The MAPLE benchmark constructed by us contains 20 datasets across 19 fields for scientific literature tagging.
1 paper · 0 benchmarks
MAQA (Multi-Answer Question Answering dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Analogical reasoning is fundamental to human cognition and holds an important place in various fields.
1 paper · 1 benchmark
MASC (Manually Annotated Sub-Corpus)
The Manually Annotated Sub-Corpus (MASC) consists of approximately 500,000 words of contemporary American English written and spoken data drawn from the Open American National Corpus (OANC).
1 paper · 0 benchmarks
MAST (Multi-Attributed Structured Text-to-face Dataset)
A new data consolidation called Multi-Attributed and Structured Text-to-face (MAST) dataset.
1 paper · 0 benchmarks
The MATHWELL Human Annotation Dataset contains 5,084 synthetic word problems and answers generated by MATHWELL, a reference-free educational grade school math word problem generator released in MATHWELL: Generating Educational Math Word…
1 paper · 0 benchmarks
The dataset contains 3 million attribute-value annotations across 1257 unique categories created from 2.2 million cleaned Amazon product profiles.
1 paper · 0 benchmarks
MAVEN-Arg is an advanced event argument extraction dataset, which offers three main advantages: A comprehensive schema covering 162 event types and 612 argument roles, all with expert-written definitions and examples.
1 paper · 0 benchmarks
Here we release the dataset (MultiChannelGrid, abbreviated as MCGrid) used in our paper LIMUSE: LIGHTWEIGHT MULTI-MODAL SPEAKER EXTRACTION](https://arxiv.org/abs/2111.04063)).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 1 benchmark
MENYO-20k is the first multi-domain parallel corpus with a special focus on clean orthography for Yorùbá--English with standardized train-test splits for benchmarking.
1 paper · 0 benchmarks
MF (Mathematical Formulas)
Mathematical dataset containing formulas based on the AMPS Khan dataset and the ARQMath dataset V1.3.
1 paper · 0 benchmarks
MGPFD (multi-goal path finding dataset)
MGPFD is a dataset for multi-goal path finding problem, including a training dataset and a simulation dataset.
1 paper · 0 benchmarks
MH-FED (Meta Human Facial Expression Dataset)
This dataset provides a collection of 162K images and 70 Videos of Meta-Humans.
1 paper · 0 benchmarks
The dataset from the study "A Fully Generative Motivational Interviewing Counsellor Chatbot for Moving Smokers Towards the Decision to Quit".
1 paper · 0 benchmarks
To assess a model’s ability to create microcontroller-driven electronic devices, we developed a benchmark, MICRO25, that includes 25 tasks intended for the common ARDUINO microcontroller ecosystem..
1 paper · 0 benchmarks
This dataset contains 4606 articles from 1996 to 2024 that were presented in MIE (Medical Informatics Europe Conference) conferences.
1 paper · 0 benchmarks
MILU (Multi-task Indic Language Understanding Benchmark)
Overview MILU (Multi-task Indic Language Understanding Benchmark) is a comprehensive evaluation dataset designed to assess the performance of Large Language Models (LLMs) across 11 Indic languages.
1 paper · 0 benchmarks
MIMI dataset (Multi-aspect Integrated Migration Indicators dataset)
Nowadays, new branches of research are proposing the use of non-traditional data sources for the study of migration trends in order to find an original methodology to answer open questions about cross-border human mobility.
1 paper · 0 benchmarks
MIMIC-IV ICD-9 contains 209,326 discharge summaries—free-text medical documents—annotated with ICD-9 diagnosis and procedure codes.
1 paper · 1 benchmark
MIMIC-IV-Note (MIMIC-IV-Note: Deidentified free-text clinical notes)
The advent of large, open access text databases has driven advances in state-of-the-art model performance in natural language processing (NLP).
1 paper · 0 benchmarks
MINT (a Multi-modal Image and Narrative Text Dubbing Dataset)
Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape.
1 paper · 0 benchmarks
MIRACL-VC1 is a lip-reading dataset including both depth and color images.
1 paper · 1 benchmark
Minor Irrigation Structures Check-Dam Dataset is a public dataset annotated by domain experts using images from Google static map for instance segmentation and object detection tasks.
1 paper · 0 benchmarks
This database includes 22 half-hour ECG recordings of subjects who experienced episodes of sustained ventricular tachycardia, ventricular flutter, and ventricular fibrillation.
1 paper · 0 benchmarks
This dataset is a supplement to the github repositiry (https://github.com/pfilonenko/MLforTwoSampleTesting) and paper addressed to solve the two-sample problem under right-censored observations using Machine Learning.
1 paper · 0 benchmarks
MLB (Mouse Lockbox Dataset)
This dataset provides high-resolution videos recorded from three perspectives with more than 110 hours of total playtime showing mice solving complex tasks.
1 paper · 0 benchmarks
The Mauna Loa Seeing Study was performed by the EOL/Integrated Surface Flux System team, capturing surface meteorology and flux products at the Mauna Loa Observatory in Hawaii.
1 paper · 2 benchmarks
MLQuestions is a domain-adaptation dataset for the machine learning domain containing 50K unaligned passages and 35K unaligned questions, and 3K aligned passage and question pairs.
1 paper · 0 benchmarks
MM-Locate-News is a dataset for location estimation of news.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.