Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 45 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 2113–2160 of 3,998
VQDv1 (Visual Query Detection v1)
In Visual Query Detection (VQD), a system is given a query (prompt) natural language and an image, and then the system must produce 0 - N boxes that satisfy that query.
2 papers · 0 benchmarks
The dataset contains traffic traces collected from 3 different VR applications.
2 papers · 0 benchmarks
VTC (Videos, Titles and Comments)
VTC is a large-scale multimodal dataset containing video-caption pairs (~300k) alongside comments that can be used for multimodal representation learning.
2 papers · 0 benchmarks
Binary labels for Validity and Novelty respectively are given for each Conclusion.
2 papers · 1 benchmark
The code to create the dataset is available here.
2 papers · 2 benchmarks
VerilogEval Dataset The VerilogEval Dataset is a benchmark specifically designed to assess the ability of large language models (LLMs) to generate syntactically correct and functionally accurate Verilog code.
2 papers · 1 benchmark
WDC-PAVE (Web Data Commones - Product Attribute Value Extraction)
The datasets contains 1,420 human annotated product offers, systematically selected from the Web Data Commons Product Matching Corpus, featuring 24,582 annotated attribute-value pairs, making it a valuable resource for both product…
2 papers · 1 benchmark
WT-WT (Will-They-Won't-They)
Will-They-Won't-They (WT-WT) is a large dataset of English tweets targeted at stance detection for the rumor verification task.
2 papers · 0 benchmarks
The Wallhack1.8k dataset comprises 1,806 CSI amplitude spectrograms (and raw WiFi packet time series) corresponding to three activity classes: "no presence," "walking," and "walking + arm-waving." WiFi packets were transmitted at a…
2 papers · 0 benchmarks
Wastewater catchment area data are essential for wastewater treatment capacity planning and have recently become critical for operationalising wastewater-based epidemiology (WBE) for COVID-19.
2 papers · 0 benchmarks
WeatherKITTI is currently the most realistic all-weather simulated enhancement of the KITTI dataset.
2 papers · 0 benchmarks
We developed Web-Bench as a benchmark for evaluating the performance of LLMs on real-world web projects.
2 papers · 1 benchmark
This paper is a condensed report on the second year of the Touché shared task on argument retrieval held at CLEF 2021.
2 papers · 0 benchmarks
WetLinks (WetLinks: a Large-Scale Longitudinal Starlink Dataset with Contiguous Weather Data)
WetLinks: a Large-Scale Longitudinal Starlink Dataset with Contiguous Weather Data.
2 papers · 0 benchmarks
WikiCaps is a large-scale multilingual but non-parallel data set for multimodal machine translation and retrieval.
2 papers · 0 benchmarks
WikiTableSet is a large publicly available image-based table recognition dataset in three languages built from Wikipedia.
2 papers · 0 benchmarks
WildDESED (Wild Domestic Environment Sound Event Detection)
WildDESED is an extension of the original DESED dataset, created to reflect various domestic scenarios by incorporating complex and unpredictable background noises.
2 papers · 1 benchmark
Wyze Rule Recommendation Dataset.
2 papers · 0 benchmarks
X-WikiRE is a new, large-scale multilingual relation extraction dataset in which relation extraction is framed as a problem of reading comprehension to allow for generalization to unseen relations.
2 papers · 0 benchmarks
It consists of an extensive collection of a high quality cross-lingual fact-to-text dataset in 11 languages: Assamese (as), Bengali (bn), Gujarati (gu), Hindi (hi), Kannada (kn), Malayalam (ml), Marathi (mr), Oriya (or), Punjabi (pa),…
2 papers · 1 benchmark
XL-R2R (Cross-lingual Room-to-Room)
The XL-R2R dataset is built upon the R2R dataset and extends it with Chinese instructions.
2 papers · 0 benchmarks
YASO is a crowd-sourced TSA evaluation dataset, collected using a new annotation scheme for labeling targets and their sentiments.
2 papers · 0 benchmarks
We present YTSeg, a topically and structurally diverse benchmark for the text segmentation task based on YouTube transcriptions.
2 papers · 1 benchmark
The YouTube8M-MusicTextClips dataset consists of over 4k high-quality human text descriptions of music found in video clips from the YouTube8M dataset.
2 papers · 0 benchmarks
bcTCGA (The Cancer Genome Atlas Program)
This data set comes from breast cancer tissue samples deposited to The Cancer Genome Atlas (TCGA) project.
2 papers · 0 benchmarks
This dataset comprises 1344 expert annotated images of muscle-tendon junctions recorded with 3 ultrasound imaging systems (Aixplorer V6, Esaote MyLab60, Telemed ArtUs), on 2 muscles (Lateral Gastrocnemius, Medial Gastrocnemius), and 2…
2 papers · 0 benchmarks
This dataset contains pre and post destruction images and also segmentation labels for test images.
2 papers · 0 benchmarks
From the official description: > The corpus contains 10-K reports from many US companies during years > 1996-2006, as well as measured volatility of stock returns for the > twelve-month periods preceding and following each report.
2 papers · 0 benchmarks
This dataset was collected during a LoRaWAN measurement campaign in a multi-room indoor office environment in the University of Siegen, Germany.
2 papers · 0 benchmarks
kickstarter (Funding Successful Projects on Kickstarter)
Kickstarter is a community of more than 10 million people comprising of creative, tech enthusiasts who help in bringing creative project to life.
2 papers · 1 benchmark
Large language models such as ChatGPT and GPT-4 have recently achieved astonishing performance on a variety of natural language processing tasks.
2 papers · 0 benchmarks
neuronIO (Single cortical neuron (L5PC) input output simulation at 1ms temporal resolution)
Single cortical neurons as deep artificial neural networks This dataset contains training and testing subsets of the input/output relationship of a single cortical layer 5 pyramidal cell (L5PC) neuron at 1ms single spike temporal…
2 papers · 0 benchmarks
news20 (NewsWeeder: learning to filter netnews)
Two datasets featuring binary and multi-class classification.
2 papers · 0 benchmarks
robo-vln (Robotics Vision-and-Language Navigation)
The Robo-VLN dataset is a continuous control formulation of the VLN-CE dataset by Krantz et al ported over from Room-to-Room (R2R) dataset created by Anderson et al.
2 papers · 1 benchmark
satp-zsm-stage1 (Replication Data for: Crossing the Linguistic Causeway: A Binational Approach for Translating Soundscape Attributes to zsm)
This is the replication data for the paper: "Crossing the Linguistic Causeway: A Binational Approach for Translating Soundscape Attributes to Bahasa Melayu".
2 papers · 0 benchmarks
chinahate dataset contains a total of 2,172,333 tweets hashtagged #china posted during the time it was collected.
1 paper · 0 benchmarks
In one round of sequencing, 5 fecal pellets from 2 pro-inflammatory environments (Harvard BRI/Johns Hopkins) and 2 pro-survival environments (Broad Institute/Jackson Labs) were sequenced at the 16s rDNA locus.
1 paper · 0 benchmarks
Official dataset for Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models.
1 paper · 0 benchmarks
Dataset of low fidelity resolutions of the RANS equations over airfoils.
1 paper · 0 benchmarks
We provide here a new multi-view text dataset, collected from three well-known online news sources: BBC, Reuters, and The Guardian.
1 paper · 0 benchmarks
3D design file repository for the Stickbug Robot a 6 armed holonomic precision pollination robot
1 paper · 0 benchmarks
3D-BSLS-6D (3D scans of Bins by Structured-Light Scanner for 6D pose estimation)
Dataset consist of both real captures from Photoneo PhoXi structured light scanner devices annotated by hand and synthetic samples produced by custom generator.
1 paper · 1 benchmark
The datasets used and analysed from the glucose clamp study are available in this DIF file.
1 paper · 0 benchmarks
The datasets used and analysed from the glucose clamp study are available in this Excel file.
1 paper · 0 benchmarks
It contains 900 audio clips, annotated into 4 quadrants, according to Russell's model.
1 paper · 0 benchmarks
These are larger MATLAB .mat files required for reproducing plots from the sgbaird-5DOF/interp repository for grain boundary property interpolation.
1 paper · 0 benchmarks
The dataset contains 60,000 Stack Overflow questions from 2016-2020, classified into three categories: 1.
1 paper · 1 benchmark
6IMPOSE (Synthetic RGBD dataset for 6D pose estimation)
The dataset includes the synthetic data generated from rendering the 3D meshes of LM objects and several household objects in Blender for training 6D pose estimation algorithms.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.