Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 41 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1921–1968 of 3,998
Description Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation" (Li et al., 2022).
2 papers · 0 benchmarks
The LeukemiaAttri dataset is a large-scale, multi-domain collection of microscopy images derived from leukemia patient samples, enriched with detailed morphological information.
2 papers · 2 benchmarks
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks.
2 papers · 0 benchmarks
LymphoMNIST is a comprehensive dataset designed for the nuanced classification of lymphocyte images.
2 papers · 0 benchmarks
MECD (Multi-Event Causal Discovery)
Provide: 1,105 lifestyle videos that span diverse scenarios.
2 papers · 1 benchmark
MGSM8KInstruct, the multilingual math reasoning instruction dataset, encompassing ten distinct languages, thus addressing the issue of training data scarcity in multilingual math reasoning.
2 papers · 0 benchmarks
This repository provides a cleaned dataset, which is intended to be used for text classification, language modeling, and AI-generated content detection tasks.
2 papers · 0 benchmarks
MIAD contains more than 100K high-resolution color images in various outdoor industrial scenarios, designed for unsupervised anomaly detection.
2 papers · 0 benchmarks
You need to request access to download and use the dataset.
2 papers · 1 benchmark
The MIMIC PERform Testing dataset contains the following physiological signals recorded from 200 critically-ill patients during routine clinical care: - electrocardiogram (ECG) - photoplethysmogram (PPG) - impedance pneumography (imp),…
2 papers · 2 benchmarks
We provide a dataset called MMAC Captions for sensor-augmented egocentric-video captioning.
2 papers · 0 benchmarks
MMSD2.0 (Towards a Reliable Multi-modal Sarcasm Detection System)
Multi-modal sarcasm detection has attracted much recent attention.
2 papers · 0 benchmarks
MMVR (Millimeter-wave Multi-View Radar (MMVR) Dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
MMVax-Stance includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
MOMAland is an open source Python library for developing and comparing multi-objective multi-agent reinforcement learning algorithms by providing a standard API to communicate between learning algorithms and environments, as well as a…
2 papers · 0 benchmarks
MPHOI-72 (Multi-person Human-object Interaction Dataset 72)
MPHOI-72 is a multi-person human-object interaction dataset that can be used for a wide variety of HOI/activity recognition and pose estimation/object tracking tasks.
2 papers · 0 benchmarks
MSDA (Multi-source domain adaptation dataset for text recognition)
5 domains: synthetic domain, document domain, street view domain, handwritten domain, and car license domain over five million images
2 papers · 2 benchmarks
The MUSE dataset contains bilingual dictionaries for 110 pairs of languages.
2 papers · 2 benchmarks
We introduce the first dataset, MUSIC-AVQA-R, to evaluate the robustness of AVQA models.
2 papers · 0 benchmarks
MaNGA (Mapping Nearby Galaxies at APO)
MaNGA is a component of the Fourth-Generation Sloan Digital Sky Survey whose goal is to map the detailed composition and kinematic structure of nearby galaxies.
2 papers · 0 benchmarks
The E-MASAC Dataset is a collection of code-mixed conversations sourced from an Indian TV series, focusing on Hindi-English interactions.
2 papers · 1 benchmark
MagicBathyNet is a benchmark dataset made up of image patches of Sentinel-2, SPOT-6 and aerial imagery, bathymetry in raster format and seabed classes annotations.
2 papers · 0 benchmarks
Our primary objective in creating this dataset is to support researchers in the advancement of algorithms for keypoints detection and the pretraining of large models on retinal images using a self-supervised approach.
2 papers · 0 benchmarks
MedMNIST-C is an open-source data set collection comprising algorithmically generated corruptions applied to the test sets of the MedMNIST collection following the concept of ImageNet-C.
2 papers · 0 benchmarks
The process by which sections in a document are demarcated and labeled is known as section identification.
2 papers · 2 benchmarks
MentSum (Mental Health Summarization Dataset)
Mental health remains a significant challenge of public health worldwide.
2 papers · 1 benchmark
MAUD is an expert-annotated merger agreement reading comprehension dataset based on the American Bar Association's 2021 Public Target Deal Points study, where lawyers and law students answered 92 questions about 152 merger agreements.
2 papers · 0 benchmarks
MetaVD is a Meta Video Dataset for enhancing human action recognition datasets.
2 papers · 0 benchmarks
This dataset contains the extraction made in 2022 of all the 622 datasets that existed then at the UCI Machine Learning Repository.
2 papers · 0 benchmarks
Mila Simulated Floods Dataset is a 1.5 square km virtual world using the Unity3D game engine including urban, suburban and rural areas.
2 papers · 1 benchmark
MoB (Malicious or Benign Cartoon Videos)
A dataset of cartoon video clips.
2 papers · 1 benchmark
Morph Call is a suite of 46 probing tasks for four Indo-European languages that fall under different morphology: Russian, French, English, and German.
2 papers · 0 benchmarks
A version of the CMU Movie Summary Corpus (http://www.cs.cmu.edu/~ark/personas/), which was originally scraped from plot summaries from Wikipedia, with some cleaning and sentences turned into events & sorted into "genres" (via LDA).
2 papers · 0 benchmarks
MultiOOD (Multimodal Out-of-Distribution Detection Benchmark)
MultiOOD is the first benchmark for Multimodal OOD Detection and covers diverse dataset sizes and modalities.
2 papers · 0 benchmarks
MultiOpEd is a corpus of multi-perspective news editorials.
2 papers · 0 benchmarks
MultiReQA is a cross-domain evaluation for retrieval question answering models.
2 papers · 0 benchmarks
Texture-based studies and designs have been in focus recently.
2 papers · 0 benchmarks
The original dataset was provided by Orange telecom in France, which contains anonymized and aggregated human mobility data.
2 papers · 0 benchmarks
The MultiviewC dataset mainly contributes to multiview cattle action recognition, 3D objection detection and tracking.
2 papers · 0 benchmarks
The MusicBrainz20K dataset for entity resolution and entity clustering is based on real records about songs from the MusicBrainz database.
2 papers · 1 benchmark
NBA: This is extended from a Kaggle dataset containing around 400 NBA basketball players.
2 papers · 1 benchmark
This database offers iris images (with and without contact lenses) of the same eyes captured shortly one after another with illumination coming from two different locations.
2 papers · 0 benchmarks
The dataset consists of titles and abstracts from NLP-related papers.
2 papers · 0 benchmarks
NMED-T (Naturalistic Music EEG Dataset - Tempo)
Losorelli, Steven, Nguyen, Duc T., Dmochowski, Jacek P., and Kaneshiro, Blair This dataset contains cortical (EEG) and behavioral data collected during natural music listening.
2 papers · 0 benchmarks
This collection contains images from 422 non-small cell lung cancer (NSCLC) patients.
2 papers · 0 benchmarks
NVGaze (NVGaze: An Anatomically-Informed Dataset for Low-Latency, Near-Eye Gaze Estimation)
Quality, diversity, and size of training dataset are critical factors for learning-based gaze estimators.
2 papers · 0 benchmarks
A collection of over 2,500 novel English words published in the New York Times between November 2017 and March 2019, manually annotated for their class of novelty (such as lexical derivation, dialectal variation, blending, or compounding).
2 papers · 0 benchmarks
The Nations dataset is a small knowledge graph with 14 entities, 55 relations, and 1992 triples describing countries and their political relationships.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.