Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 27 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 1249–1296 of 3,998
PSI (IUPUI-CSRC Pedestrian Situated Intent)
The IUPUI-CSRC Pedestrian Situated Intent (PSI) benchmark dataset has two innovative labels besides comprehensive computer vision annotations.
5 papers · 0 benchmarks
ParaQA is a question answering (QA) dataset with multiple paraphrased responses for single-turn conversation over knowledge graphs (KG).
5 papers · 0 benchmarks
The PodcastFillers dataset consists of 199 full-length podcast episodes in English with manually annotated filler words and automatically generated transcripts.
5 papers · 1 benchmark
PyTorrent contains 218,814 Python package libraries from PyPI and Anaconda environment.
5 papers · 0 benchmarks
QAConv is a new question answering (QA) dataset that uses conversations as a knowledge source.
5 papers · 0 benchmarks
RARE (Randomized AMRs with Rewired Edges)
RARE consists of English AMR pairs with similarity scores that reflect the structural differences between them.
5 papers · 1 benchmark
The evaluation of object detection models is usually performed by optimizing a single metric, e.g.
5 papers · 1 benchmark
RITE (Retinal Images vessel Tree Extraction)
The RITE (Retinal Images vessel Tree Extraction) is a database that enables comparative studies on segmentation or classification of arteries and veins on retinal fundus images, which is established based on the public available DRIVE…
5 papers · 2 benchmarks
RealMAN (A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization)
The Audio Signal and Information Processing Lab at Westlake University, in collaboration with AISHELL, has released the Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset, which provides annotated multi-channel…
5 papers · 2 benchmarks
RefRef (RefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects)
RefRef is a synthetic dataset and benchmark designed for the task of reconstructing scenes with complex refractive and reflective objects.
5 papers · 1 benchmark
Ruddit is a dataset of English language Reddit comments that has fine-grained, real-valued scores for offensive language detection between -1 (maximally supportive) and 1 (maximally offensive).
5 papers · 0 benchmarks
S.MID (SeMantic InDustry)
SeMantic InDustry (S.MID) is a dataset designed to advance the field of LiDAR semantic segmentation, specifically for robotic applications and large-scale industrial scene.
5 papers · 1 benchmark
SAF (Short Answer Feedback Dataset)
This dataset can be found on HuggingFace: https://huggingface.co/datasets/Short-Answer-Feedback/safcommunicationnetworksenglish https://huggingface.co/datasets/Short-Answer-Feedback/safmicrojobgerman
5 papers · 0 benchmarks
SAFIM (Syntax-Aware Fill-In-the-Middle)
Syntax-Aware Fill-in-the-Middle (SAFIM) is a benchmark for evaluating Large Language Models (LLMs) on the code Fill-in-the-Middle (FIM) task.
5 papers · 1 benchmark
SDSD-indoor (Seeing Dynamic Scene in the Dark: High-Quality Video Dataset with Mechatronic Alignment)
The dataset collected by the paper Seeing Dynamic Scene in the Dark: High-Quality Video Dataset with Mechatronic Alignment, ICCV 2021
5 papers · 1 benchmark
SERV-CT (SERV-CT: A disparity dataset from CT for validation of endoscopic 3D reconstruction)
Endoscopic stereo reconstruction for surgical scenes gives rise to specific problems, including the lack of clear corner features, highly specular surface properties, and the presence of blood and smoke.
5 papers · 0 benchmarks
SODA-D is a large-scale dataset tailored for small object detection in driving scenario, which is built on top of MVD dataset and owned data, where the former is a dataset dedicated to pixel-level understanding of street scenes, and the…
5 papers · 1 benchmark
The SOFC-Exp corpus contains 45 scientific publications about solid oxide fuel cells (SOFCs), published between 2013 and 2019 as open-access articles all with a CC-BY license.
5 papers · 0 benchmarks
SSN (Semantic Scholar Network)
SSN (short for Semantic Scholar Network) is a scientific papers summarization dataset which contains 141K research papers in different domains and 661K citation relationships.
5 papers · 0 benchmarks
Sales (Rossmann Store Sales)
Forecast Sales using ARIMA and SARIMA
5 papers · 0 benchmarks
The Sarcasm Corpus contains sarcastic and non-sarcastic utterances of three different types, which are balanced with half of the samples being sarcastic and half non-sarcastic.
5 papers · 0 benchmarks
Sewer-ML is a sewer defect dataset.
5 papers · 0 benchmarks
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
5 papers · 0 benchmarks
SoMeSci (Software Mentions in Scientific Articles)
Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling.
5 papers · 0 benchmarks
The Song Describer Dataset (SDD) contains ~1.1k captions for 706 permissively licensed music recordings.
5 papers · 1 benchmark
SpaRTUN a dataset synthesized for transfer learning on spatial question answering (SQA) and spatial role labeling (SpRL).
5 papers · 0 benchmarks
SpaceNet 1: Building Detection v1 is a dataset for building footprint detection.
5 papers · 2 benchmarks
- Games dataset containing 100,000 Gameplay Images of 175 Video Games across 10 Sports Genres - AMERICAN FOOTBALL, BASKETBALL, BIKE RACING, CAR RACING, FIGHTING, HOCKEY, SOCCER, TABLE TENNIS, TENNIS.
5 papers · 2 benchmarks
The SUGARCREPE++ dataset evaluates the sensitivity of vision language models (VLMs) and unimodal language models (ULMs) to semantic and lexical alterations.
5 papers · 0 benchmarks
TRANCE (Transformation Driven Visual Reasoning)
TRANCE extends CLEVR by asking a uniform question, i.e.
5 papers · 0 benchmarks
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
This is a 16.2-million frame (50-hour) multimodal dataset of two-person face-to-face spontaneous conversations.
5 papers · 0 benchmarks
The Taskmaster-2 dataset consists of 17,289 dialogs in seven domains: restaurants (3276), food ordering (1050), movies (3047), hotels (2355), flights (2481), music (1602), and sports (3478).
5 papers · 0 benchmarks
The goal of this dataset is to probe video-language models for understanding of simple temporal relations like "before" and "after".
5 papers · 1 benchmark
A Dense-text Image Benchmark to evaluate large generation model's ability on text generation.
5 papers · 1 benchmark
Timers and Such is an open source dataset of spoken English commands for common voice control use cases involving numbers.
5 papers · 1 benchmark
Twitter-MEL is a multimodal entity linking (MEL) dataset built from Twitter.
5 papers · 0 benchmarks
Twitter100k is a large-scale dataset for weakly supervised cross-media retrieval.
5 papers · 0 benchmarks
UruDendro (UruDendro, a public dataset of cross-section images of pinus taeda)
UruDendro is a database of wood cross section images of commercially grown Pinus taeda trees from northern Uruguay.
5 papers · 1 benchmark
The dataset uses VGG-Sound which consists of 10s clips collected from YouTube for 309 sound classes.
5 papers · 0 benchmarks
VQA-VS (a new VQA benchmark considering Varying Shortcuts)
The current OOD benchmark VQA-CP v2 only considers one type of shortcut (from question type to answer) and thus still cannot guarantee that the modelrelies on the intended solution rather than a solution specific to this shortcut.
5 papers · 0 benchmarks
The Vent dataset is a large annotated dataset of text, emotions, and social connections.
5 papers · 0 benchmarks
A large-scale multi-modal dataset to facilitate research and studies that concentrate on vision-wireless systems.
5 papers · 1 benchmark
VidChapters-7M is a dataset of 817K user-chaptered videos including 7M chapters in total.
5 papers · 4 benchmarks
VidOR (Video Object Relation) dataset contains 10,000 videos (98.6 hours) from YFCC100M collection together with a large amount of fine-grained annotations for relation understanding.
5 papers · 1 benchmark
WEAR (WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition)
WEAR is an outdoor sports dataset for both vision- and inertial-based human activity recognition (HAR).
5 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.