Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 63 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 2977–3024 of 3,998

The first annotated corpus for multilingual analysis of potentially unfair clauses in online Terms of Service.
1 paper · 0 benchmarks
Multimedia Goal-oriented Generative Script Learning Dataset This link contains a dataset consisting of multimedia steps for two categories: gardening and crafts.
1 paper · 0 benchmarks
Spectroscopic techniques are essential tools for determining the structure of molecules.
1 paper · 0 benchmarks
Multiview Manipulation Data (Multiview Manipulation Expert Data and Trained Models)
Accompanying expert data and trained models for 2021 IROS paper on Multiview Manipulation.
1 paper · 0 benchmarks
NAIST COVID is a multilingual dataset of social media posts related to COVID-19, consisting of microblogs in English and Japanese from Twitter and those in Chinese from Weibo.
1 paper · 0 benchmarks
This dataset contains names that are exclusively associated with a single gender and that have no ambiguous meanings, therefore being exact with respect to both gender and meaning.
1 paper · 0 benchmarks
This dataset extends NAMEXACT by including words that can be used as names, but may not exclusively be used as names in every context.
1 paper · 0 benchmarks
NAS Dataset for DIP (Neural Architecture Search Dataset for DIP)
Dataset for our CVPR paper: "ISNAS-DIP: Image-Specific Neural Architecture Search for Deep Image Prior".
1 paper · 0 benchmarks
NBA_Box_Scores_Odds (NBA Team-Level Box Score Statistics (2015-2019), Historical Win Percentages (2014-2018) and Betting Odds (2018/2019))
Dataset Description: NBA Team Statistics, Historical Performance & Betting Odds (2015-2019) Overview This dataset contains team-level box score statistics, historical win percentages, and closing betting odds for NBA games from 2015 to…
1 paper · 0 benchmarks
NBMOD (Noisy Background Multi-Object Dataset for grasp detection)
Introduction NBMOD is a dataset created for researching the task of specific object grasp detection by robots in noisy environments.
1 paper · 1 benchmark
NCANDA (National Consortium on Alcohol and Neurodevelopment in Adolescence)
The NCANDA consortium is composed of an Administrative component at the University of California San Diego, a Data Analysis and Informatics component at SRI International, and five research sites (University of California San Diego, SRI…
1 paper · 0 benchmarks
NCSE v2.0 (NCSE v2.0: A Dataset of OCR-Processed 19th Century English Newspapers)
The NCSE v2.0 is a digitized collection of six 19th-century English periodicals The ground truth contains 358 cropped images of text blocks from 31 pages of 19th century newspaper data
1 paper · 0 benchmarks
NEREL-BIO is an annotation scheme and corpus of PubMed abstracts in Russian and English.
1 paper · 0 benchmarks
NES-VMDB is a dataset containing 98,940 gameplay videos from 389 NES games, each paired with its original soundtrack in symbolic format (MIDI).
1 paper · 0 benchmarks
NII-CU MAPD (NII-CU Multispectral Aerial Person Detection Dataset)
The National Institute of Informatics - Chiba University (NII-CU) Multispectral Aerial Person Detection Dataset consists of 5,880 pairs of aligned RGB+FIR (Far infrared) images captured from a drone flying at heights between 20 and 50…
1 paper · 2 benchmarks
We announce the release of a new multilingual speaker dataset called NITK-IISc Multilingual Multi-accent Speaker Profiling(NISP) dataset.
1 paper · 0 benchmarks
The NISQA Corpus includes more than 14,000 speech samples with simulated (e.g.
1 paper · 0 benchmarks
NJH (Not Just Hate)
NJH is a dataset of over 40,000 tweets about immigration from the US and UK, annotated with six labels for different aspects of incivility and intolerance.
1 paper · 0 benchmarks
NL2GQL Dataset (NL2GQL developped for R3-NL2GQL)
A bilingual (English and Chinese natural language queries) dataset which has NL queries annotated with their corresponding GQL queries (i.e.
1 paper · 0 benchmarks
NLI4Wills Corpus can be used to train transformers and sentence-transformer models for the validity evaluation of the legal will statements.
1 paper · 0 benchmarks
This project is a collection of three corpora which can be used for evaluating chatbots or other conversational interfaces.
1 paper · 0 benchmarks
NMF (Named Mathematical Formulas)
Mathematical dataset based on 71 famous mathematical identities.
1 paper · 0 benchmarks
This corpus contains data files that were generated as part of the NOVIC paper (see above).
1 paper · 0 benchmarks
NPO (Negative and Positive Obstacles)
The dataset is recorded with an on-vehicle ZED stereo camera in both urban and rural environments The dataset contains various lighting conditions, such as normal lights, large-area shadows, dim lights, and sun glare.
1 paper · 1 benchmark
Unique radiogenomic dataset from a Non-Small Cell Lung Cancer (NSCLC) cohort of 211 subjects.
1 paper · 0 benchmarks
NTIC Screening Dataset (Niramai Thermal Image for COVID19 Screening)
In the last two years, millions of lives have been lost due to COVID-19.
1 paper · 0 benchmarks
Motion similarity annotations for NTU RGB+D 120 dataset to evaluate motion similarity in the real world.
1 paper · 0 benchmarks
Includes co-referent name string pairs along with their similarities.
1 paper · 0 benchmarks
Narvik Road Dataset (DIT4BEARs Smart Road Dataset)
DIT4BEARs Internship Project (at UiT-The Arctic University of Norway) Dataset The dataset contains data of 5 months including weather conditions, friction coefficient, distance traveled, wind speed, surface temperature, air temperature,…
1 paper · 0 benchmarks
We scraped the Gutenberg Project and a subset of English Wikipedia to obtain the list of sentences that contain any.
1 paper · 0 benchmarks
Introduction to the Needle In A Haystack Test The Needle In A Haystack test, inspired by NeedleInAHaystack, is an evaluation method that randomly inserts key information into long texts to create prompts for large language models (LLMs).
1 paper · 0 benchmarks
Nelson-Plosser (Nelson-Plosser US Macroeconomic Time Series)
US Macroeconomic dataset containing 14 time series of monthly observations.
1 paper · 0 benchmarks
The dataset comprises 1641 questions and answers generated as three separate parts.
1 paper · 0 benchmarks
NeurIT Dataset is open-sourced for public research usage.
1 paper · 0 benchmarks
NewsMTSC is a dataset for target-dependent sentiment classification (TSC) on news articles reporting on policy issues.
1 paper · 0 benchmarks
Introduction The Niko Chord Progression Dataset is used in AccoMontage2.
1 paper · 0 benchmarks
Niramai Oncho Dataset (Niramai Onchocerciasis/RiverBlindness Dataset)
Onchocerciasis is causing blindness in over half a million people in the world today.
1 paper · 0 benchmarks
The AASL-Clear dataset is a collection of RGB images featuring Arabic alphabet sign Language gestures with backgrounds removed.
1 paper · 1 benchmark
NoLiMa (NoLiMa: Long-Context Evaluation Beyond Literal Matching)
A benchmark extending needle-in-a-haystack (NIAH) test with a carefully designed needle set, where questions and needles have minimal lexical overlap, requiring models to infer latent associations to locate the needle within the haystack.
1 paper · 0 benchmarks
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
1 paper · 0 benchmarks
Dataset composed of two main parts 1.
1 paper · 0 benchmarks
This dataset artifact contains the intermediate datasets from pipeline executions necessary to reproduce the results of the paper.
1 paper · 0 benchmarks
Authors of the Dataset: - Pratik Bhowal (B.E., Dept of Electronics and Instrumentation Engineering, Jadavpur University Kolkata, India)…
1 paper · 1 benchmark
Dynamic occupancy grids generated from NuScenes dataset.
1 paper · 0 benchmarks
NurViD (A Large Expert-Level Video Database for Nursing Procedure Activity Understanding)
We propose NurViD, a large video dataset with expert-level annotation for nursing procedure activity understanding.
1 paper · 0 benchmarks
OAGT (Paper Topic Dataset)
OAGL is a paper topic dataset consisting of 6942930 records which comprise various scientific publication attributes like abstracts, titles, keywords, publication years, venues, etc.
1 paper · 0 benchmarks
OCB (Open Circuit Benchmark)
OCB contains two graph datasets, Ckt-Bench-101 and Ckt-Bench-301, for representation learning over analog circuits.
1 paper · 0 benchmarks
OCTScenes contains 5000 tabletop scenes with a total of 15 everyday objects.
3D
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.