Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 5 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 193–240 of 3,998

Tudataset: A collection of benchmark datasets for learning with graphs
96 papers · 1 benchmark
UAVDT (Unmanned Aerial Vehicle Benchmark Object Detection and Tracking)
UAVDT is a large scale challenging UAV Detection and Tracking benchmark (i.e., about 80, 000 representative frames from 10 hours raw videos) for 3 important fundamental tasks, i.e., object DETection (DET), Single Object Tracking (SOT) and…
96 papers · 2 benchmarks
CBT (Children’s Book Test)
Children’s Book Test (CBT) is designed to measure directly how well language models can exploit wider linguistic context.
92 papers · 1 benchmark
A large-scale multi-object tracking dataset for human tracking in occlusion, frequent crossover, uniform appearance and diverse body gestures.
91 papers · 1 benchmark
PopQA is an open-domain QA dataset with 14k QA pairs with fine-grained Wikidata entity ID, Wikipedia page views, and relationship type information.
91 papers · 1 benchmark
WikiMatrix is a dataset of parallel sentences in the textual content of Wikipedia for all possible language pairs.
91 papers · 0 benchmarks
E2E (End-to-End NLG Challenge)
End-to-End NLG Challenge (E2E) aims to assess whether recent end-to-end NLG systems can generate more complex output by learning from datasets containing higher lexical richness, syntactic complexity and diverse discourse phenomena.
90 papers · 4 benchmarks
ST-VQA (Scene Text Visual Question Answering)
ST-VQA aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the VQA process.
90 papers · 0 benchmarks
Winoground is a dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning.
90 papers · 1 benchmark
The COCO-Text dataset is a dataset for text detection and recognition.
89 papers · 2 benchmarks
Dataset produced for the SAPIEN simulation environment.
88 papers · 0 benchmarks
ASPEC (Asian Scientific Paper Excerpt Corpus)
ASPEC, Asian Scientific Paper Excerpt Corpus, is constructed by the Japan Science and Technology Agency (JST) in collaboration with the National Institute of Information and Communications Technology (NICT).
87 papers · 0 benchmarks
GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and unsupervised…
87 papers · 3 benchmarks
KP20k is a large-scale scholarly articles dataset with 528K articles for training, 20K articles for validation and 20K articles for testing.
87 papers · 3 benchmarks
The MovieQA dataset is a dataset for movie question answering.
86 papers · 1 benchmark
CAMUS (Cardiac Acquisitions for Multi-structure Ultrasound Segmentation)
This project aims to provide all the materials to the community to resolve the problem of echocardiographic image segmentation and volume estimation from 2D ultrasound sequences (both two and four-chamber views).
85 papers · 0 benchmarks
A large-scale and machine-generated dataset of 274,186 toxic and benign statements about 13 minority groups.
85 papers · 0 benchmarks
CodeContests is a competitive programming dataset for machine-learning.
84 papers · 1 benchmark
The How2 dataset contains 13,500 videos, or 300 hours of speech, and is split into 185,187 training, 2022 development (dev), and 2361 test utterances.
84 papers · 2 benchmarks
MiniF2F is a dataset of formal Olympiad-level mathematics problems statements intended to provide a unified cross-system benchmark for neural theorem proving.
84 papers · 2 benchmarks
MusicCaps is a dataset composed of 5.5k music-text pairs, with rich text descriptions provided by human experts.
84 papers · 1 benchmark
TweetEval introduces an evaluation framework consisting of seven heterogeneous Twitter-specific classification tasks.
84 papers · 1 benchmark
WikiBio (Wikipedia Biography Dataset)
This dataset gathers 728,321 biographies from English Wikipedia.
84 papers · 1 benchmark
GovReport is a dataset for long document summarization, with significantly longer documents and summaries.
83 papers · 2 benchmarks
NLVR (Natural Language Visual Reasoningnatural language for visual reasoning)
NLVR contains 92,244 pairs of human-written English sentences grounded in synthetic images.
83 papers · 3 benchmarks
Our task is to localize and provide a pixel-level mask of an object on all video frames given a language referring expression obtained either by looking at the first frame only or the full video.
82 papers · 1 benchmark
LoveDA (Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation)
1.
81 papers · 1 benchmark
VQG (Visual Question Generation)
VQG is a collection of datasets for visual question generation.
80 papers · 1 benchmark
CoNLL-2014 will continue the CoNLL tradition of having a high profile shared task in natural language processing.
79 papers · 0 benchmarks
MOSES (Molecular sets (MOSES))
The set is based on the ZINC Clean Leads collection.
79 papers · 0 benchmarks
WikiTableQuestions is a question answering dataset over semi-structured tables.
79 papers · 2 benchmarks
PHOENIX14T (RWTH-PHOENIX-Weather-2014T)
Over a period of three years (2009 - 2011) the daily news and weather forecast airings of the German public tv-station PHOENIX featuring sign language interpretation have been recorded and the weather forecasts of a subset of 386 editions…
78 papers · 0 benchmarks
CoNaLa (CMU CoNaLa, the Code/Natural Language Challenge)
The CMU CoNaLa, the Code/Natural Language Challenge dataset is a joint project from the Carnegie Mellon University NeuLab and Strudel labs.
77 papers · 1 benchmark
Few-NERD is a large-scale, fine-grained manually annotated named entity recognition dataset, which contains 8 coarse-grained types, 66 fine-grained types, 188,200 sentences, 491,711 entities, and 4,601,223 tokens.
77 papers · 3 benchmarks
Multilingual LibriSpeech is a large multilingual corpus suitable for speech research.
77 papers · 2 benchmarks
ECB+ (extension to the EventCorefBank)
The ECB+ corpus is an extension to the EventCorefBank (ECB, Bejan and Harabagiu, 2010).
76 papers · 0 benchmarks
TAT-QA (Tabular And Textual dataset for Question Answering) is a large-scale QA dataset, aiming to stimulate progress of QA research over more complex and realistic tabular and textual data, especially those requiring numerical reasoning.
76 papers · 1 benchmark
TREC-COVID is a community evaluation designed to build a test collection that captures the information needs of biomedical researchers using the scientific literature during a pandemic.
73 papers · 1 benchmark
WikiANN (PAN-X)
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
The shared task of CoNLL-2002 concerns language-independent named entity recognition.
70 papers · 3 benchmarks
DiffusionDB is a large-scale text-to-image prompt dataset.
70 papers · 1 benchmark
MSP-IMPROV (MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception)
We present the MSP-IMPROV corpus, a multimodal emotional database, where the goal is to have control over lexical content and emotion while also promoting naturalness in the recordings.
70 papers · 1 benchmark
YouTube-UGC (YouTube UGC dataset)
This YouTube dataset is a sampling from thousands of User Generated Content (UGC) as uploaded to YouTube distributed under the Creative Commons license.
70 papers · 1 benchmark
CMRC (Chinese Machine Reading Comprehension)
CMRC is a dataset is annotated by human experts with near 20,000 questions as well as a challenging set which is composed of the questions that need reasoning over multiple clues.
69 papers · 0 benchmarks
Multimodal Opinionlevel Sentiment Intensity (MOSI) contains: (1) multimodal observations including transcribed speech and visual gestures as well as automatic audio and visual features, (2) opinion-level subjectivity segmentation, (3)…
69 papers · 1 benchmark
QMSum is a new human-annotated benchmark for query-based multi-domain meeting summarisation task, which consists of 1,808 query-summary pairs over 232 meetings in multiple domains.
69 papers · 1 benchmark
UBFC-rPPG (Univ. Bourgogne Franche-Comté Remote PhotoPlethysmoGraphy)
We introduce here a new database called UBFC-rPPG (stands for Univ.
69 papers · 1 benchmark
VSR (Visual Spatial Reasoning)
The Visual Spatial Reasoning (VSR) corpus is a collection of caption-image pairs with true/false labels.
69 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.