Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 17 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 769–816 of 3,998
This paper introduces the Broad Twitter Corpus (BTC), which is not only significantly bigger, but sampled across different regions, temporal periods, and types of Twitter users.
12 papers · 2 benchmarks
CDCP (Cornell eRulemaking Corpus)
The Cornell eRulemaking Corpus – CDCP is an argument mining corpus annotated with argumentative structure information capturing the evaluability of arguments.
12 papers · 3 benchmarks
CDR (BioCreative V CDR Task Corpus)
The BioCreative V CDR task corpus is manually annotated for chemicals, diseases and chemical-induced disease (CID) relations.
12 papers · 2 benchmarks
CICERO (Contextualized Commonsense Inference in Dialogues)
CICERO contains 53,000 inferences for five commonsense dimensions -- cause, subsequent event, prerequisite, motivation, and emotional reaction -- collected from 5600 dialogues.
12 papers · 4 benchmarks
CIRCLE is a dataset containing 10 hours of full-body reaching motion from 5 subjects across nine scenes, paired with ego-centric information of the environment represented in various forms, such as RGBD videos.
12 papers · 1 benchmark
The COCO-MLT is created from MS COCO-2017, containing 1,909 images from 80 classes.
12 papers · 2 benchmarks
Along with COVID-19 pandemic we are also fighting an infodemic'.
12 papers · 1 benchmark
ConditionalQA is a Question Answering (QA) dataset that contains complex questions with conditional answers, i.e.
12 papers · 1 benchmark
The DSTC7 Task 1 dataset is a dataset and task for goal-oriented dialogue.
12 papers · 1 benchmark
Earnings-21, a 39-hour corpus of earnings calls containing entity-dense speech from nine different financial sectors.
12 papers · 0 benchmarks
FoolMeTwice (FM2 for short) is a large dataset of challenging entailment pairs collected through a fun multi-player game.
12 papers · 0 benchmarks
GLGE (General Language Generation Evaluation)
GLGE is a general language generation evaluation benchmark which is composed of 8 language generation tasks, including Abstractive Text Summarization (CNN/DailyMail, Gigaword, XSUM, MSNews), Answer-aware Question Generation (SQuAD 1.1,…
12 papers · 0 benchmarks
GMOT-40 (Generic Multiple Object Tracking (GMOT))
GMOT-40 is the first public dense dataset for Generic Multiple Object Tracking (GMOT).
12 papers · 2 benchmarks
The Hutter Prize Wikipedia dataset, also known as enwiki8, is a byte-level dataset consisting of the first 100 million bytes of a Wikipedia XML dump.
12 papers · 1 benchmark
HyperKvasir dataset contains 110,079 images and 374 videos where it captures anatomical landmarks and pathological and normal findings.
12 papers · 2 benchmarks
100 tasks from LIBERO-100 suite.
12 papers · 1 benchmark
LLaMEA (algorithms and experiments from the paper)
3500+ Generated evolutionary algorithms by the LLaMEA framework.
12 papers · 0 benchmarks
Lila is a unified mathematical reasoning benchmark consisting of 23 diverse tasks along four dimensions: (i) mathematical abilities e.g., arithmetic, calculus (ii) language format e.g., question-answering, fill-in-the-blanks (iii) language…
12 papers · 0 benchmarks
Math-Vision (Math-V) dataset is a meticulously curated collection of 3,040 high-quality mathematical problems with visual contexts sourced from real math competitions.
12 papers · 1 benchmark
MM-COVID (Multilingual and Multidimensional COVID-19 Fake News Data Repository)
MM-COVID is a dataset for fake news detection related to COVID-19.
12 papers · 0 benchmarks
MMNeedle (Multimodal Needle in a Haystack)
We introduce the MultiModal Needle-in-a-haystack (MMNeedle) benchmark, specifically designed to assess the long-context capabilities of MLLMs.
12 papers · 1 benchmark
MasakhaNEWS is a benchmark dataset for news topic classification covering 16 languages widely spoken in Africa.
12 papers · 0 benchmarks
Neptune (Neptune Long Video Understanding Benchmark)
Neptune is a dataset consisting of challenging question-answer-decoy (QAD) sets for long videos (up to 15 minutes).
12 papers · 0 benchmarks
The largest real-world night-time semantic segmentation dataset with pixel-level labels.
12 papers · 0 benchmarks
The Overruling dataset is a law dataset corresponding to the task of determining when a sentence is overruling a prior decision.
12 papers · 1 benchmark
PROST (Physical Reasoning about Objects Through Space and Time)
The PROST (Physical Reasoning about Objects Through Space and Time) dataset contains 18,736 multiple-choice questions made from 14 manually curated templates, covering 10 physical reasoning concepts.
12 papers · 0 benchmarks
We present PeerQA, a real-world, scientific, document-level Question Answering (QA) dataset.
12 papers · 3 benchmarks
A dataset of color images corrupted by natural noise due to low-light conditions, together with spatially and intensity-aligned low noise images of the same scenes.
12 papers · 1 benchmark
Co-speech gestures are everywhere.
12 papers · 1 benchmark
The Dataset is part of the KELM corpus This is the Wikipedia text--Wikidata KG aligned corpus used to train the data-to-text generation model.
12 papers · 1 benchmark
We construct the long-tailed version of VOC from its 2012 train-val set.
12 papers · 2 benchmarks
WildScenes is a bi-modal benchmark dataset consisting of multiple large-scale, sequential traversals in natural environments, including semantic annotations in high-resolution 2D images and dense 3D LiDAR point clouds, and accurate 6-DoF…
12 papers · 2 benchmarks
This dataset has 20 classes and each class has about 1000 documents.
11 papers · 1 benchmark
4D-OR includes a total of 6734 scenes, recorded by six calibrated RGB-D Kinect sensors 1 mounted to the ceiling of the OR, with one frame-per-second, providing synchronized RGB and depth images.
11 papers · 3 benchmarks
AMPS (Auxiliary Mathematics Problems and Solutions)
AMPS contains over 100,000 problems pulled from Khan Academy and approximately 5 million problems generated from manually designed Mathematica scripts.
11 papers · 0 benchmarks
ARCH is a computational pathology (CP) multiple instance captioning dataset to facilitate dense supervision of CP tasks.
11 papers · 0 benchmarks
ASAP-AES (Automated Student Assessment Prize)
There are eight essay sets.
11 papers · 1 benchmark
A set of 19 ASC datasets (reviews of 19 products) producing a sequence of 19 tasks.
11 papers · 1 benchmark
COST (COCO Segmentation Text)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
11 papers · 0 benchmarks
This is the home of a collaborative data collection effort by U.
11 papers · 1 benchmark
This is the home of a collaborative data collection effort by U.
11 papers · 1 benchmark
ClariQ is an extension of the Qulac dataset with additional new topics, questions, and answers in the training set.
11 papers · 0 benchmarks
Complementary Commonsense (Com2Sense) is a dataset for benchmarking commonsense reasoning ability of NLP models.
11 papers · 0 benchmarks
The EARS-WHAM dataset mixes speech from the EARS dataset with real noise recordings from the WHAM!
11 papers · 1 benchmark
the YF-E6 emotion dataset using the 6 basic emotion type as keywords on social video-sharing websites including YouTube and Flickr, leading to a total of 3000 videos.
11 papers · 1 benchmark
FLIP (Fitness Landscape Inference for Proteins)
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function).
11 papers · 0 benchmarks
Predicting forest cover type from cartographic variables only (no remotely sensed data).
11 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
11 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.