Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 83 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3937–3984 of 3,998
SynD (A Synthetic Energy Dataset for Non-Intrusive Load Monitoring in Households)
SynD is a synthetic energy dataset with a focus on residential buildings.
0 papers · 0 benchmarks
TCMP-300 (Traditional Chinese Medicinal Plant Dataset)
Traditional Chinese medicinal plants are often used to prevent and treat diseases for the human body.
0 papers · 1 benchmark
Dataset Introduction TFHAnnotatedDataset is an annotated patent dataset pertaining to thin film head technology in hard-disk.
0 papers · 0 benchmarks
Face detection and subsequent localization of facial landmarks are the primary steps in many face applications.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
0 papers · 0 benchmarks
TS-TR (Turkish Scene Text Recognition Dataset)
The Turkish Scene Text Recognition (TS-TR) dataset was primarily developed to fill the gap in non-English text recognition resources, specifically addressing the unique challenges presented by the Turkish language, such as special…
0 papers · 0 benchmarks
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection.
0 papers · 0 benchmarks
Collection of stream of consciousness.
0 papers · 0 benchmarks
The Reddit COVID Dataset is a dataset of 4.51M Reddit posts and 17.8M comments - all mentions of COVID until 2021-10-25 across the entire Reddit social network.
0 papers · 0 benchmarks
The Tornado Network (TorNet) dataset is a large, high-resolution benchmark dataset developed to support machine learning research in tornado detection and prediction.
0 papers · 0 benchmarks
Toronto NeuroFace Dataset: A New Dataset for Facial Motion Analysis in Individuals with Neurological Disorders Toronto NeuroFace Dataset is a public dataset with videos of oro-facial gestures performed by individuals with oro-facial…
0 papers · 0 benchmarks
Description: The Traffic Sign Recognition Dataset is designed to support the development of deep learning models, particularly for object detection and classification.
0 papers · 0 benchmarks
UAS-based Multispectral othomosaics of vineyards from central Portugal - 2 distinct vineyards - Multispectral and HD orthomosaics
0 papers · 0 benchmarks
UT Zappos50K (UT-Zap50K) is a large shoe dataset consisting of 50,025 catalog images collected from Zappos.com.
0 papers · 0 benchmarks
The Unsplash Dataset is created by over 200,000 contributing photographers and billions of searches across thousands of applications, uses, and contexts.
0 papers · 0 benchmarks
Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine.
0 papers · 0 benchmarks
The dataset contains more than 35000 images and 600 videos captured using 35 different portable devices of 11 major brands.
0 papers · 0 benchmarks
Vehicle-1M involves vehicle images captured across day and night, from head or rear, by multiple surveillance cameras installed in cities.
0 papers · 0 benchmarks
This dataset is an extremely challenging set of over 2000+ original Visiting card/ID card images captured and crowdsourced from over 300+ urban and rural areas, where each image is manually reviewed and verified by computer vision…
0 papers · 0 benchmarks
Visual Fields (UWHVF: A real-world, open source dataset of Humphrey Visual Fields (HVF) from the University of Washington)
28,943 Humphrey Visual Field (HVF) tests from 3,871 patients and 7,428 eyes.
0 papers · 0 benchmarks
Vript (🎬 Vript: A Video Is Worth Thousands of Words)
We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips).
0 papers · 0 benchmarks
The WASABI Song Corpus is a large corpus of songs enriched with metadata extracted from music databases on the Web, and resulting from the processing of song lyrics and from audio analysis.
0 papers · 0 benchmarks
reference paper YOUNIS, H., RATTROUT, A., & YOUNIS, M.
0 papers · 0 benchmarks
WTA/TLA (WTA/TLA: A UAV-captured Dataset for Semantic Segmentation of Energy Infrastructure)
WTA (Wind Turbine Aerial) and TLA (Transmission Line Aerial) are public datasets which contain a set of RGB images from wind turbine farms and transmission towers and power lines, along with semantic ground truth for relevant classes.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.