Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 73 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3457–3504 of 3,998
ULI-RI (Unreal Labeled Images for Person Re-ID)
The ULI-RI dataset is generated using the Unreal Engine 4 to simulate various outdoor environments with 115 high-quality 3D human models.
1 paper · 0 benchmarks
ULS labeled data (UVA laser scanning labelled las data over tropical moist forest classified as leaf or wood points)
UAV Laser Scanning data collected over neotropical forest (Paracou French Guiana).
1 paper · 1 benchmark
UMC005 English-Urdu is a parallel corpus of texts in English and Urdu language with sentence alignments.
1 paper · 0 benchmarks
One-Shot Affordance Part Segmentation variant of the UMD dataset.
1 paper · 0 benchmarks
Repository for UML-English data This repository contains the data used for "Extraction of UML Class Diagrams from Natural Language Specification" (Yang et al.
1 paper · 1 benchmark
UNER v1 adds an NER annotation layer to 18 datasets (primarily treebanks from UD) and covers 12 geneologically and ty- pologically diverse languages: Cebuano, Danish, German, English, Croatian, Portuguese, Russian, Slovak, Serbian,…
1 paper · 31 benchmarks
A comprehensive dataset, merging all the aforementioned datasets.
1 paper · 0 benchmarks
In this dataset UR5 robot used 6 tools: metal-scissor, metal-whisk, plastic-knife, plastic-spoon, wooden-chopstick, and wooden-fork to perform 6 behaviors: look, stirring-slow, stirring-fast, stirring-twist, whisk, and poke.
1 paper · 0 benchmarks
USCOCO (Unexpected Situations of Common Objects in Context)
A test set of grammatically correct sentences and layouts (visual “imagined” situations), called Unexpected Situations of Common Objects in Context (USCOCO) describing compositions of entities and relations that are unlikely to be found in…
1 paper · 0 benchmarks
The USNA long-term scintillation study is a continuing effort to characterize and measure optical turbulence in the near-maritime boundary layer.
1 paper · 1 benchmark
The USNA long-term scintillation study is a continuing effort to characterize and measure optical turbulence in the near-maritime boundary layer.
1 paper · 2 benchmarks
The semantic segmentation of clothes is a challenging task due to the wide variety of clothing styles, layers and shapes.
1 paper · 1 benchmark
The UTRSet-Synth dataset is introduced as a complementary training resource to the UTRSet-Real Dataset, specifically designed to enhance the effectiveness of Urdu OCR models.
1 paper · 0 benchmarks
UV6K (Urban Vehicle Segmentation Dataset)
UV6K is a high-resolution remote sensing urban vehicle segmentation dataset.
1 paper · 1 benchmark
The Ubuntu Chat Corpus (UCC) is composed of archived chat logs from Ubuntu's Internet Relay Chat technical support channels.
1 paper · 0 benchmarks
This dataset extends the Semantic Segmentation of Underwater Imagery: Dataset and Benchmark, adding an uncertainty evaluation component.
1 paper · 0 benchmarks
This data contains the election polls for the 2004, 2008, 2012, and 2016 US presidential election by state including data on undecided voter proportions.
1 paper · 0 benchmarks
Uniswap (Replication Data for: Uniswap Daily Transaction Indices by Network)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The United-Syn-Med dataset is a specialized medical speech dataset designed to evaluate and improve Automatic Speech Recognition (ASR) systems within the healthcare domain.
1 paper · 0 benchmarks
Unpaired dataset: The dataset is built by ourselves, and there are all real haze images from websites.
1 paper · 0 benchmarks
The Unsplash Dataset is made up of over 350,000+ contributing global photographers and data sourced from hundreds of millions of searches across a nearly unlimited number of uses and contexts.
1 paper · 0 benchmarks
The prospective upper body thermal images SARS-CoV2 association study was designed to test the hypothesis that thermal videos may aid in the early diagnosis of COVID-19.
1 paper · 0 benchmarks
Urban Dict spelling variant is a variant spelling dataset for use of NLP research in the informal domain.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset includes User Story (or Issue) text descriptions, User Story titles, and Story Points from 33 software development projects, comprising a total of 20,479 User Stories (or issues) extracted from GitLab repositories, amounting…
1 paper · 0 benchmarks
V-MIND enhanced the MIND dataset with news pictures.
1 paper · 0 benchmarks
V-Rank (SIMULATED AIRCRAFT TRAJECTORY FOR THEORETICAL VELOCITY RANKING)
Abstract This data set is a data set used for aircraft theoretical velocity ranking.
1 paper · 0 benchmarks
V3C1 (the Vimeo Creative Commons Collection 1)
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, and will serve as evaluation basis for the Video Browser Showdown 2019-2021 and TREC Video Retrieval…
1 paper · 0 benchmarks
Uses same clean speech as VoiceBank+Demand but more noise types.
1 paper · 1 benchmark
VCG+112K (Video Instruction Dataset 112K)
Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs.
1 paper · 0 benchmarks
This task stems from the observation that text embedded in images is intrinsically different from common visual elements and natural language due to the need to align the modalities of vision, text, and text embedded in images.
1 paper · 0 benchmarks
VD-Ref is a dataset with ground-truth mappings from both noun phrases and pronouns to image regions.
1 paper · 0 benchmarks
VESUS (Varied Emotion in Syntactically Uniform Speech)
The Varied Emotion in Syntactically Uniform Speech (VESUS) repository is a lexically controlled database collected by the NSA lab.
1 paper · 0 benchmarks
VFD-2000 is a video fight detection dataset containing more than 2000 videos.
1 paper · 0 benchmarks
A synthetic dataset containing word images of 447 typefaces with font variations for each typeface, created for visual font recognition.
1 paper · 1 benchmark
A synthetic dataset containing 447 typefaces with only one font variation for each typeface, created for visual font recognition.
1 paper · 1 benchmark
A high-resolution version of VGGFace2 for academic face editing purposes.
1 paper · 0 benchmarks
VIRDO Dataset (VIRDO Simulated Kitchen Utensil Deformation Dataset)
From https://github.com/MMintLab/VIRDO/blob/master/data/datasetreadme.txt, 1.
1 paper · 0 benchmarks
VISEM-Tracking is a dataset consisting of 20 video recordings of 30s of spermatozoa with manually annotated bounding-box coordinates and a set of sperm characteristics analyzed by experts in the domain.
1 paper · 0 benchmarks
VMD (Virtual Moderation Dataset)
This dataset contains synthetically generated discussions and annotations using exclusively Large Language Model (LLM) agents.
1 paper · 0 benchmarks
VME & CDSI (Vehicles in the Middle East (VME) & Car Detection in Satellite Imagery (CDSI) datasets)
Vehicles in the Middle East (VME) dataset, designed explicitly for vehicle detection in high-resolution satellite images from Middle Eastern countries.
1 paper · 1 benchmark
The benchmark for VPData, the largest video inpainting dataset, which comprises over 390K clips (> 866.7 hours) and features precise masks and detailed video captions.
1 paper · 0 benchmarks
The largest video inpainting dataset comprises over 390K clips (> 866.7 hours), featuring precise masks and detailed video captions.
1 paper · 0 benchmarks
VSTaR-1M is a 1M instruction tuning dataset, created using Video-STaR, with the source datasets: Kinetics700 STAR-benchmark FineDiving The videos for VSTaR-1M can be found in the links above.
1 paper · 0 benchmarks
Combines CoVaxFrames and HpVaxFrames into a unified dataset of 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines and 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines.
1 paper · 0 benchmarks
A Natural Language Resource for Learning to Recognize Misinformation about the COVID-19 and HPV Vaccines.
1 paper · 0 benchmarks
Validity and Novelty are determined in a comparative setting between two conclusions at a time.
1 paper · 1 benchmark
ValiMath is a high-quality benchmark consisting of 2,147 carefully curated mathematical questions designed to evaluate an LLM's ability to verify the correctness of math questions based on multiple logic-based and structural criteria.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.