Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 67 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3169–3216 of 3,998

A real-world stereo video dataset, containing 1200 frame pairs with real-world color and sharpness mismatches caused by beam splitter.
1 paper · 0 benchmarks
RealHDRTV dataset is the first real-world paired SDRTV-HDRTV dataset, which includes SDRTV-HDRTV pairs with 8K resolutions captured by a smartphone camera with the “SDR” and “HDR10” modes.
1 paper · 0 benchmarks
RealVul (RealVul-Vulnerability Dataset following realistic settings)
This is a C++ vulnerability detection dataset following realistic settings.
1 paper · 0 benchmarks
Reddit Engagement Dataset (RED), a distant-supervision set, with 80k single-turn conversations.
1 paper · 0 benchmarks
Dataset with articles posted in the r/Liberal and r/Conservative subreddits.
1 paper · 1 benchmark
This is a dataset of over 40K Reddit comments removed by moderators according to the specific type of macro norm being violated.
1 paper · 0 benchmarks
Reddit Posts Related To Eating Disorders and Dieting (Topic Annotations on Reddit Posts from Eating Disorders and Dieting Forums by Human and LLMs)
This dataset comprises 77,175 Reddit posts from 115 subreddit forums, annotated for the presence of 15 topics related to eating disorders and dieting.
1 paper · 0 benchmarks
https://arxiv.org/abs/2503.15222
1 paper · 0 benchmarks
Teaching assistants (TAs) are heavily used in computer science courses as a way to handle high enrollment and still being able to offer students individual tutoring and detailed assessments.
1 paper · 0 benchmarks
This dataset was acquired in a retrospective study from a cohort of pediatric patients admitted with abdominal pain to Children’s Hospital St.
1 paper · 0 benchmarks
The primary environmental health threat in the WHO European Region is air pollution, impacting the daily health and well-being of its citizens significantly.
1 paper · 0 benchmarks
This dataset is a collection of 5348 links from bug-introducing and bug-fixing commit sets extracted from Mozilla's Bugzilla with the use of bugbug.
1 paper · 0 benchmarks
Relicensing_Forks (Relicensed OSS projects and resulting forks)
The notebooks folder contains basic analysis of the organizational affiliation data for the contributors per open source project.
1 paper · 0 benchmarks
RepLab 2013 dataset uses Twitter data in English and Spanish (more than 142,000 tweets).
1 paper · 0 benchmarks
Replication Data for: "DAM" (Replication Data for: "DAM: A Universal Dual Attention Mechanism for Multimodal Timeseries Cryptocurrency Trend Forecasting")
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Bitcoin is a peer-to-peer electronic payment system that popularized rapidly in recent years.
1 paper · 0 benchmarks
Replication Data for: AI Ethics on Blockchain (Replication Data for: AI Ethics on Blockchain: Topic Analysis on Twitter Data for Blockchain Security Version 2.0)
Blockchain has empowered computer systems to be more secure using a distributed network.
1 paper · 0 benchmarks
Replication Data for: Blockchain Network Analysis (> Replication Data for: Blockchain Network Analysis: A Comparative Study of Decentralized Banks)
Decentralized finance (DeFi) is known for its unique mechanism design, which applies smart contracts to facilitate peer-to-peer transactions.
1 paper · 0 benchmarks
The dataset provides information about 450 HYIPs collected between November 2020 and September 2021.
1 paper · 0 benchmarks
Replication Data for: On the Mechanics of NFT Valuation (Replication Data for: On the Mechanics of NFT Valuation: AI Ethics and Social Media)
As CryptoPunks pioneers the innovation of non-fungible tokens (NFTs) in AI and art, the valuation mechanics of NFTs has become a trending topic.
1 paper · 0 benchmarks
The model forecasts for the sub-seasonal forecasting application considered in the Online Learning under Optimism and Delay paper experiments.
1 paper · 0 benchmarks
Replication Data for: Singapore Soundscape Site Selection Survey (S5) (Identification of Characteristic Soundscapes of Singapore via Weighted k-means Clustering)
This dataset contains the data used for all statistical analysis in our publication "Singapore Soundscape Site Selection Survey (S5): Identification of Characteristic Soundscapes of Singapore via Weighted k-means Clustering", summarised in…
1 paper · 0 benchmarks
This is our replication package for our study on Benchmarking scalability of stream processing frameworks deployed as microservices in the cloud.
1 paper · 0 benchmarks
This is the replication package for our systematic literature review and can be used for the reproducibility of the individual steps of our search and selection methodology.
1 paper · 0 benchmarks
This package contains the data and the reported results for the manuscript: Keo B, Li B, Younis W (2025) Measuring trade costs and analyzing the determinants of trade growth between Cambodia and major trading partners: 1993–2019.
1 paper · 0 benchmarks
This is a research artifact for the ICSE'22 paper "GitHub Sponsors: Exploring a New Way to Contribute to Open Source".
1 paper · 0 benchmarks
ReviewRobot Dataset Overview This repository contains data for paper ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis.
1 paper · 0 benchmarks
ata Set Name: Rice Dataset (Commeo and Osmancik) Abstract: A total of 3810 rice grain's images were taken for the two species (Cammeo and Osmancik), processed and feature inferences were made.
1 paper · 0 benchmarks
Riposte! (Riposte! A Large Corpus of Counter-Arguments)
From the Riposte!
1 paper · 0 benchmarks
Risholme-2021 contains >3.5K images of strawberries at various growth stages along with anomalous instances.
1 paper · 0 benchmarks
Risk-Aware Planning is a dataset that contains the overhead images and their semantic segmentation captured by a drone from the CityEnviron environment in AirSim simulator.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Ritter PoS (Ritter Twitter part-of-speech tagging)
PTB-tagged English Tweets
1 paper · 0 benchmarks
Dataset contains light curves of 6 rocket body types from Mini Mega Tortora database (MMT)[^1].
1 paper · 0 benchmarks
Robot@Home2 (Robot@Home2, a robotic dataset of home environments)
Robot@Home2, is an enhanced version aimed at improving usability and functionality for developing and testing mobile robotics and computer vision algorithms.
1 paper · 0 benchmarks
RoomEnv-v0 (The Room environment - v0)
The Room environment - v0 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
RoomEnv-v1 (The Room environment - v1)
The Room environment - v1 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
RoomEnv-v2 (The Room environment - v2)
The Room environment - v2 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
RoomSpace: a new benchmark designed to evaluate language models on spatial reasoning tasks demanding spatial relation knowledge and multi-hop reasoning.
1 paper · 0 benchmarks
Fact-based Text Editing dataset based on RotoWire dataset
1 paper · 1 benchmark
The RotoWire-Modified dataset is a cleaned extension of the RotoWire dataset, with writer information about each document.
1 paper · 0 benchmarks
S-BIAD843 (Individual 3D cell shapes of Drosophila Wing Disc)
Late third instar wing imaginal discs were cultured in Shields and Sang M3 media (Sigma) supplemented with 2% FBS (Sigma), 1% pen/strep (Gibco), 3ng/ml ecdysone (Sigma) and 2ng/ml insulin (Sigma).
1 paper · 0 benchmarks
SACID (Saliency Aware Compressed Images Dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SAD-Instruct (Situational Awareness Database for Instruct-Tuning)
The Situational Awareness Database for Instruct-Tuning (SAD-Instruct) is a dataset for dynamic task guidance.
1 paper · 0 benchmarks
SAGC-A68 (A space access graph dataset for the classification of spaces and space elements in apartment buildings)
The analysis of building models for usable area, building safety, and energy efficiency requires accurate classification data of spaces and space elements.
1 paper · 0 benchmarks
SAIL 2017 (Sentiment Analysis for Indian Languages)
India is a linguistic area with one of the longest histories of contact, influence, use, teaching and learning of English-in-diaspora in the world (Kachru and Nelson, 2006).
1 paper · 1 benchmark
SARA motion (Synthetic Actors and Real Actions)
Sara motion is a 3D motion dataset, named Synthetic Actors and Real Actions (SARA), for training a model to produce motion embeddings suitable for reasoning about motion similarity.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.