Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 81 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3841–3888 of 3,998

The MOBIO database consists of bi-modal (audio and video) data taken from 152 people.
0 papers · 0 benchmarks
MVP-24K (Multi-grained Vehicle Parsing dataset)
Multi-grained Vehicle Parsing (MVP) is a large-scale dataset for semantic analysis of vehicles in the wild, which has several featured properties.
0 papers · 0 benchmarks
MapAI: Precision in Building Segmentation Dataset The dataset comprises 7500 training images and 1500 validation images from Denmark.
0 papers · 0 benchmarks
This dataset was developed within an analysis of research data generated and managed within the University of Bologna, with respect to the differences and commonalities between disciplines and potential challenges for institutional data…
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Media-Text (MediaText: a media industry-based dataset for scene text detetcion)
Media-Text dataset comprising images of banners, posters, covers and another images characterised for media industry.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Moh (SeyedMohammad Kashani)
We introduce an open-source physical-layer dataset of Bluetooth Low Energy (BLE) IoT sensor devices recorded in an anechoic chamber using USRP x310.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Monopedal Gaits (Periodic Trajectories of a Passive One-Legged Hopper)
The dataset comprises time-series data capturing distinct periodic motions (gaits) of an energetically conservative one-legged hopper.
0 papers · 0 benchmarks
MoralChoice is a survey dataset to evaluate the moral beliefs encoded in LLMs.
0 papers · 0 benchmarks
A dataset of all Moroccan money
0 papers · 0 benchmarks
Mudestreda (Mudestreda Multimodal Device State Recognition Dataset)
Mudestreda Multimodal Device State Recognition Dataset obtained from real industrial milling device with Time Series and Image Data for Classification, Regression, Anomaly Detection, Remaining Useful Life (RUL) estimation, Signal Drift…
0 papers · 0 benchmarks
Multi-Spectral Leaf Segmentation (Multi-Spectral Leaf Segmentation For Crop/Weed Identification)
This dataset were acquired with the Airphen (Hyphen, Avignon, France) six-band multi-spectral camera configured using the 450/570/675/710/730/850 nm bands with a 10 nm FWHM.
0 papers · 0 benchmarks
Abstract: We introduce the multi-spectral stereo (MS2) outdoor dataset, including stereo RGB, stereo NIR, stereo thermal, stereo LiDAR data, and GPS/IMU information.
0 papers · 0 benchmarks
Multimodal Humor Dataset (Multimodal Humor Dataset: Predicting Laughter Tracks for Sitcoms)
A great number of situational comedies (sitcoms) are being regularly made and the task of adding laughter tracks to these is a critical task.
0 papers · 0 benchmarks
Multimodal Large Language Models (MLLMs) have shown significant promise in various applications, leading to broad interest from researchers and practitioners alike.
0 papers · 0 benchmarks
NEMO (NEMO: A Database for Emotion Analysis Using Functional Near-Infrared Spectroscopy)
We present a dataset for the analysis of human affective states using functional near-infrared spectroscopy (fNIRS).
0 papers · 0 benchmarks
NHR-Edit (NoHumansRequired Edit Dataset)
NHR-Edit is a training dataset for instruction-based image editing.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
NIAN (Needle in a Needlestack)
The Needle in a Needlestack (NIAN) is a new benchmark designed to measure how well Language Learning Models (LLMs) pay attention to the information in their context window¹.
0 papers · 0 benchmarks
NeoRL-2 includes new task scenarios that better reflect real-world task properties and includes traditional control methods as the data-collecting method.
0 papers · 0 benchmarks
Based on RADDLE and SNIPS , we construct Noise-SF, which includes two different perturbation settings.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
This study’s sample consists of seven corporations (Black Rock, Google, Meta, JP Morgan, Walgreens, Netflix, and Pepsico) analyzed across seven quarters beginning in 2021.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Data in this study come from western Ecuador's Choco tropical forest, including \textit{Fundación para la Conservación de los Andes Tropicales Reserve and adjacent Reserva Ecológica Mache-Chindul park} (FCAT; 00°23'28'' N, 79°41'05'' W),…
0 papers · 0 benchmarks
PANACEA (PANACEA dataset - Heterogeneous COVID-19 Claims)
The peer-reviewed publication for this dataset has been presented in the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), and can be accessed here:…
0 papers · 0 benchmarks
PJM Hourly Energy Consumption Data PJM Interconnection LLC (PJM) is a regional transmission organization (RTO) in the United States.
0 papers · 0 benchmarks
Post-Spraying Image Evaluation This dataset is for the paper Deep Learning for Precision Agriculture: Post-Spraying Evaluation and Deposition Estimation (https://arxiv.org/abs/2409.16213).
0 papers · 0 benchmarks
All existing databases of spoofed speech contain attack data that is spoofed in its entirety.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Citation Request : 1.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Pose Estimation Lunar Robot (Dataset for camera pose estimation research using computer simulated images from rovers on the lunar surface)
Overview The goal: using simulation data to train neural networks to estimate the pose of a rover's camera with respect to a known target object The mission context: A simulated lunar surface, with lunar landers and lunar rovers.
0 papers · 0 benchmarks
0 papers · 0 benchmarks
Cancer genomics and precision oncology: The TCGA Research Network started in 2005 has profiled and analyzed a large number of human tumors to discover molecular aberrations at the DNA, RNA, protein, and epigenetic levels and thereby…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.