Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 15 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 673–720 of 3,998

Node classification on Texas with 60%/20%/20% random splits for training/validation/test.
16 papers · 1 benchmark
ToolQA is a question answering benchmark for Large Language Models (LLMs) which is designed to faithfully evaluate LLMs' ability to use external tools for question answering.
16 papers · 0 benchmarks
Traffic (Traffic Flow Forecasting Data Set)
Abstract: The task for this dataset is to forecast the spatio-temporal traffic volume based on the historical traffic volume and other features in neighboring locations.
16 papers · 2 benchmarks
TripClick is a large-scale dataset of click logs in the health domain, obtained from user interactions of the Trip Database health web search engine.
16 papers · 0 benchmarks
The TweepFake dataset consists of 25,572 social media messages posted either by bots or humans on Twitter.
16 papers · 1 benchmark
V-D4RL provides pixel-based analogues of the popular D4RL benchmarking tasks, derived from the dmcontrol suite, along with natural extensions of two state-of-the-art online pixel-based continuous control algorithms, DrQ-v2 and DreamerV2,…
16 papers · 0 benchmarks
The ViGGO corpus is a set of 6,900 meaning representation to natural language utterance pairs in the video game domain.
16 papers · 1 benchmark
VideoLQ consists of videos downloaded from various video hosting sites such as Flickr and YouTube, with a Creative Common license.
16 papers · 1 benchmark
This dataset is a Wikipedia dump, split by relations to perform Few-Shot Knowledge Graph Completion.
16 papers · 0 benchmarks
XQLFW (Cross-Quality Labeled Faces in the Wild)
An evaluation protocol for face verification focusing on a large intra-pair image quality difference.
16 papers · 1 benchmark
e-SNLI-VE is a large VL (vision-language) dataset with NLEs (natural language explanations) with over 430k instances for which the explanations rely on the image content.
16 papers · 2 benchmarks
BeerAdvocate is a dataset that consists of beer reviews from beeradvocate.
15 papers · 1 benchmark
Data was collected for normal bearings, single-point drive end and fan end defects.
15 papers · 1 benchmark
Node classification on Citeseer with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
Node classification on Cora with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
CrossNER is a cross-domain NER (Named Entity Recognition) dataset, a fully-labeled collection of NER data spanning over five diverse domains (Politics, Natural Science, Music, Literature, and Artificial Intelligence) with specialized…
15 papers · 1 benchmark
CustomHumans is recorded by a multi-view photogrammetry system equipped with 53 RGB (12 Megapixels) and 53 (4 Megapixels) IR cameras.
15 papers · 1 benchmark
The database consists of 150 annotated pages of three different medieval manuscripts with challenging layouts.
15 papers · 2 benchmarks
The DUC2004 dataset is a dataset for document summarization.
15 papers · 4 benchmarks
FaVIQ (Fact Verification from Information-seeking Questions)
FaVIQ (Fact Verification from Information-seeking Questions) is a challenging and realistic fact verification dataset that reflects confusions raised by real users.
15 papers · 0 benchmarks
The GoodsAD dataset contains 6124 images with 6 categories of common supermarket goods.
15 papers · 1 benchmark
Humicroedit is a humorous headline dataset.
15 papers · 0 benchmarks
LIVE-FB LSVQ (LIVE-FB Large-Scale Social Video Quality (LSVQ) Database)
No-reference (NR) perceptual video quality assessment (VQA) is a complex, unsolved, and important problem to social and streaming media applications.
15 papers · 1 benchmark
MMHS150k (Multimodal Hate Speech)
Existing hate speech datasets contain only textual data.
15 papers · 0 benchmarks
MaSS (Multilingual corpus of Sentence-aligned Spoken utterances) is an extension of the CMU Wilderness Multilingual Speech Dataset, a speech dataset based on recorded readings of the New Testament.
15 papers · 3 benchmarks
MedVidQA (Medical Video Question Answering)
The MedVidQA dataset contains the collection of 3, 010 manually created health-related questions and timestamps as visual answers to those questions from trusted video sources, such as accredited medical schools with an established…
15 papers · 0 benchmarks
The Montreal Archive of Sleep Studies (MASS) is an open-access and collaborative database of laboratory-based polysomnography (PSG) recordings O’Reilly, C., et al.
15 papers · 4 benchmarks
N-ImageNet (Large-Scale Dataset for Event-Based Object Recognition)
The N-ImageNet dataset is an event-camera counterpart for the ImageNet dataset.
15 papers · 2 benchmarks
Large-scale American Sign Language (ASL) - English dataset collected from online video sites (e.g., YouTube).
15 papers · 0 benchmarks
Opusparcus is a paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish.
15 papers · 0 benchmarks
Pile of Law is a ∼256GB (and growing) dataset of legal and administrative data which can be used for assessing norms on data sanitization across legal and administrative settings.
15 papers · 0 benchmarks
Project CodeNet is a large-scale dataset with approximately 14 million code samples, each of which is an intended solution to one of 4000 coding problems.
15 papers · 0 benchmarks
Node classification on PubMed with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 1 benchmark
PubTables-1M (PubMed Tables One Million)
The goal of PubTables-1M is to create a large, detailed, high-quality dataset for training and evaluating a wide variety of models for the tasks of table detection, table structure recognition, and functional analysis.
15 papers · 0 benchmarks
A Benchmark for Robust Multi-Hop Spatial Reasoning in Texts
15 papers · 1 benchmark
This dataset includes 4,500 fully annotated images (over 30,000 license plate characters) from 150 vehicles in real-world scenarios where both the vehicle and the camera (inside another vehicle) are moving.
15 papers · 1 benchmark
VGMIDI is a dataset of piano arrangements of video game soundtracks.
15 papers · 0 benchmarks
Node classification on Wisconsin with the fixed 48%/32%/20% splits provided by Geom-GCN.
15 papers · 2 benchmarks
xCodeEval is one of the largest executable multilingual multitask benchmarks consisting of 17 programming languages with execution-level parallelism.
15 papers · 0 benchmarks
ACL Anthology Reference Corpus (ACL ARC) is a collection of 10,920 academic papers from the ACL Anthology.
14 papers · 3 benchmarks
ARAD-1K (Ntire 2022 spectral recovery challenge and data set)
The dataset used for NTIRE 2022 Spectral Recovery Challenge
14 papers · 1 benchmark
AVSD (Audio-Visual Scene-Aware Dialog)
The Audio Visual Scene-Aware Dialog (AVSD) dataset, or DSTC7 Track 3, is a audio-visual dataset for dialogue understanding.
14 papers · 1 benchmark
Development of a benchmark corpus to support the automatic extraction of drug-related adverse effects from medical case reports.
14 papers · 3 benchmarks
BabyLM is a dataset for small scale language modeling, human language acquisition, low-resource NLP, and cognitive modeling.
14 papers · 0 benchmarks
BioLAMA is a benchmark comprised of 49K biomedical factual knowledge triples for probing biomedical Language Models.
14 papers · 0 benchmarks
A SemEval shared task in which participants must extract definitions from free text using a term-definition pair corpus that reflects the complex reality of definitions in natural language.
14 papers · 0 benchmarks
DailyTalk is a high-quality conversational speech dataset designed for Text-to-Speech.
14 papers · 0 benchmarks
ELPV (A dataset of functional and defective solar cells extracted from EL images of solar modules)
The dataset contains 2,624 samples of 300×300 pixels 8-bit grayscale images of functional and defective solar cells with varying degree of degradations extracted from 44 different solar modules.
14 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.