Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 21 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 961–1008 of 3,998
Emilia Dataset (An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation)
Recent advancements in speech generation models have been significantly driven by the use of large-scale training data.
8 papers · 0 benchmarks
Europarl-ASR (EN) is a 1300-hour English-language speech and text corpus of parliamentary debates for (streaming) Automatic Speech Recognition training and benchmarking, speech data filtering and speech data verbatimization, based on…
8 papers · 2 benchmarks
FeTS2022 (Federated Tumor Segmentation Challenge 2022)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
8 papers · 0 benchmarks
A dataset for fine-grained entity typing of knowledge graph entities built from Freebase.
8 papers · 0 benchmarks
A GQA-based dataset with 1,040,830 multi-modal explanations of visual reasoning processes.
8 papers · 1 benchmark
GTA (A Benchmark for General Tool Agents)
A benchmark to evaluate the tool-use capabilities of LLM-based agents in real-world scenarios.
8 papers · 0 benchmarks
GVFC (Gun Violence Frame Corpus)
This is a new dataset of news headlines and their frames related to the issue of gun violence in the United States.
8 papers · 0 benchmarks
GigaST is a large-scale pseudo speech translation (ST) corpus.
8 papers · 0 benchmarks
HANNA (HANNA, a large annotated dataset of Human-ANnotated NArratives for ASG evaluation.)
HANNA, a large annotated dataset of Human-ANnotated NArratives for Automatic Story Generation (ASG) evaluation, has been designed for the benchmarking of automatic metrics for ASG.
8 papers · 0 benchmarks
The HInt dataset is frequently used as a generalizability benchmark for 3D Hand Reconstruction.
8 papers · 1 benchmark
HumAID (Human-Annotated Disaster Incidents Data)
Social networks are widely used for information consumption and dissemination, especially during time-critical events such as natural disasters.
8 papers · 0 benchmarks
The dataset contains transactions made by credit cards in September 2013 by European cardholders.
8 papers · 2 benchmarks
Kinetics-100 is a dataset split created from the Kinetics dataset to evaluate the performance of few-shot action recognition models.
8 papers · 1 benchmark
Logic2Text is a large-scale dataset with 10,753 descriptions involving common logic types paired with the underlying logical forms.
8 papers · 0 benchmarks
LongForm dataset is created by leveraging English corpus examples with augmented instructions.
8 papers · 0 benchmarks
MISAW (MIcro-Surgical Anastomose Workflow recognition on training sessions)
The MISAW data set is composed of 27 sequences of micro-surgical anastomosis on artificial blood vessels performed by 3 surgeons and 3 engineering students.
8 papers · 1 benchmark
MO-Gymnasium is an open source Python library for developing and comparing multi-objective reinforcement learning algorithms by providing a standard API to communicate between learning algorithms and environments, as well as a standard set…
8 papers · 0 benchmarks
MRDA (ICSI Meeting Recorder Dialog Act Corpus)
The MRDA corpus consists of about 75 hours of speech from 75 naturally-occurring meetings among 53 speakers.
8 papers · 1 benchmark
MathMLben is a benchmark to the evaluate tools for mathematical format conversion (LaTeX ↔ MathML ↔ CAS).
8 papers · 0 benchmarks
The NLC2CMD Competition hosted at NeurIPS 2020 aimed to bring the power of natural language processing to the command line.
8 papers · 1 benchmark
OPRA (Online Product Reviews for Affordances)
The OPRA Dataset was introduced in Demo2Vec: Reasoning Object Affordances From Online Videos (CVPR'18) for reasoning object affordances from online demonstration videos.
8 papers · 2 benchmarks
OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data.
8 papers · 1 benchmark
QA2D (Question to Declarative Sentence (QA2D) Dataset)
The Question to Declarative Sentence (QA2D) Dataset contains 86k question-answer pairs and their manual transformation into declarative sentences.
8 papers · 0 benchmarks
QUASAR-S (QUestion Answering by Search And Reading – Stack Overflow)
QUASAR-S is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
8 papers · 0 benchmarks
RFUND (Revised FUNSD and XFUND)
RFUND is a relabeled version of FUNSD and XFUND datasets, tackling the following issues in their original annotations: 1.
8 papers · 0 benchmarks
This dataset, called RodoSol-ALPR dataset, contains 20,000 images captured by static cameras located at pay tolls owned by the Rodovia do Sol (RodoSol) concessionaire, which operates 67.5 kilometers of a highway (ES-060) in the Brazilian…
8 papers · 0 benchmarks
This dataset aims at evaluating the License Plate Character Segmentation (LPCS) problem.
8 papers · 1 benchmark
SciGraphQA is a large-scale, open-domain dataset focused on generating multi-turn conversational question-answering dialogues centered around understanding and describing scientific graphs and figures.
8 papers · 0 benchmarks
NLPContributionGraph was introduced as Task 11 at SemEval 2021 for the first time.
8 papers · 0 benchmarks
The SmartLights benchmark from Snipstests the capability of controlling lights in different rooms.
8 papers · 1 benchmark
We introduce a new audio dataset called SoundDescs that can be used for tasks such as text to audio retrieval, audio captioning etc.
8 papers · 1 benchmark
Spider 2.0 is a comprehensive code generation agent task that includes 632 examples.
8 papers · 1 benchmark
SubjQA is a question answering dataset that focuses on subjective (as opposed to factual) questions and answers.
8 papers · 0 benchmarks
T³Bench is the first comprehensive text-to-3D benchmark containing diverse text prompts of three increasing complexity levels that are specially designed for 3D generation (300 prompts in total).
8 papers · 1 benchmark
This dataset is collected from various global and local news sources.
8 papers · 0 benchmarks
TalkDown is a labelled dataset for condescension detection in context.
8 papers · 0 benchmarks
Autonomous trucking is a promising technology that can greatly impact modern logistics and the environment.
8 papers · 1 benchmark
Vision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored.
8 papers · 0 benchmarks
The Video-based Multimodal Summarization with Multimodal Output (VMSMO) corpus consists of 184,920 document-summary pairs, with 180,000 training pairs, 2,460 validation and test pairs.
8 papers · 0 benchmarks
VideoXum is an enriched large-scale dataset for cross-modal video summarization.
8 papers · 1 benchmark
WANDS (Wayfair ANnotation Dataset)
The dataset contains: 42,994 candidate products with data comprising product class, title, description, attributes, category hierarchy, average rating, and number of reviews 480 search query strings with predicted product class 233,448…
8 papers · 0 benchmarks
Many e-shops have started to mark-up product data within their HTML pages using the schema.org vocabulary.
8 papers · 4 benchmarks
News translation is a recurring WMT task.
8 papers · 0 benchmarks
Who's Waldo is a dataset of 270K image–caption pairs, depicting interactions of people, that is automatically mined from Wikimedia Commons.
8 papers · 1 benchmark
We manually performed the task of Open Information Extraction on 5 short documents, elaborating tentative guidelines for the task, and resulting in a ground truth reference of 347 tuples.
8 papers · 1 benchmark
Wiki (Web Traffic Time Series Forecasting)
Context There's a story behind every dataset and here's your opportunity to share yours.
8 papers · 3 benchmarks
WildReceipt is a collection of receipts.
8 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.