Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 69 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3265–3312 of 3,998

Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The satire dataset is a new multi-modal dataset of satirical and regular news articles.
1 paper · 0 benchmarks
Scan Entities in 3D (ScanEnts3D) is a large-scale dataset which provides explicit correspondences between 369k objects across 84k natural referentural sentences, covering 705 real-world scenes.
1 paper · 0 benchmarks
SciCo (Scientific Concept Induction Corpus)
SciCo is an expert-annotated dataset for hierarchical CDCR (cross-document coreference resolution) for concepts in scientific papers, with the goal of jointly inferring coreference clusters and hierarchy between them.
1 paper · 0 benchmarks
This resource contains 10.5 million paragraphs with associated statement labels, realized as one paragraph per file, one sentence per line.
1 paper · 0 benchmarks
Scifi TV Shows (Scifi TV Show Plot Summaries & Events)
A collection of long-running (80+ episodes) science fiction TV show synopses, scraped from Fandom.com wikis.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Scroll Readability Dataset contains scroll interactions of 598 participants reading advanced and elementary texts from the OneStopEnglish corpus.
1 paper · 0 benchmarks
1 paper · 0 benchmarks
This data set includes all raw data (e.g., collected certificates) of the WWW 2021 paper "Security of Alerting Authorities in the WWW: Measuring Namespaces, DNSSEC, and Web PKI".
1 paper · 0 benchmarks
This dataset is obtained during an ICON project (2017-2018) in collaboration with KU Leuven (ESAT-STADIUS), UZ Leuven, UCB, Byteflies and Pilipili.
1 paper · 0 benchmarks
Dataset Card for SemTabNet This dataset accompanies the following paper: Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar…
1 paper · 1 benchmark
Test dataset for Semantic Segmentation.
1 paper · 0 benchmarks
SemanticSugarBeets, a novel and high-quality dataset containing 953 monocular RGB images and 2920 annotations of sugar beets, enables a wide range of learning tasks including object detection, semantic segmentation, instance segmentation…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SensoDat is a dataset of self-driving car simulation data (30K executed simulations).
1 paper · 0 benchmarks
The dataset is based on a debate.org crawl.
1 paper · 0 benchmarks
Sentiment Merged (SST-3, DynaSent R1/R2)
This is a dataset for 3-way sentiment classification of reviews (negative, neutral, positive).
1 paper · 1 benchmark
SentimentArcs’ reference corpus for novels consists of 25 narratives selected to create a diverse set of well recognized novels that can serve as a benchmark for future studies.
1 paper · 0 benchmarks
This is the sequential instructions dataset from Understanding the Effects of RLHF on LLM Generalisation and Diversity.
1 paper · 0 benchmarks
ShadowLink dataset is designed to evaluate the impact of entity overshadowing on the task of entity disambiguation.
1 paper · 0 benchmarks
The ShapeIt dataset introduced by Alper et al.
1 paper · 0 benchmarks
Shaved Ice Snowflake VM Demand Dataset (Snowflake Dataset for "Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads" paper)
This repository contains documentation for the dataset that accompanies our ICPE 2025 paper, "Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads".
1 paper · 0 benchmarks
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts 🔥 Key Features - 3000+ hours of synthetic speech - Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including: - 📖 Reading Style - 🎙️…
1 paper · 0 benchmarks
ShopTC-100K Dataset The ShopTC-100K dataset is collected using TermMiner, an open-source data collection and topic modeling pipeline introduced in the paper: Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable…
1 paper · 0 benchmarks
In this Adjudicator ScoresShort Stories and Written Reflections folder: Four files from four student participants of the contest.
1 paper · 0 benchmarks
Siena Scalp EEG Database (Physionet Siena Scalp EEG Database)
The database consists of EEG recordings of 14 patients acquired at the Unit of Neurology and Neurophysiology of the University of Siena.
1 paper · 0 benchmarks
Procedural videos show step-by-step demonstrations of tasks like recipe preparation.
1 paper · 0 benchmarks
The SimBEV dataset is a collection of 320 scenes spread across all 11 CARLA maps and contains data from a variety of sensors, including five camera types (RGB, semantic segmentation, instance segmentation, depth, and optical flow), lidar,…
1 paper · 3 benchmarks
SimpEvalASSET is a dataset for learning learnable metrics using modern language models.
1 paper · 0 benchmarks
It consists of 32x32 pixel images of shapes with multiple attributes (size, location, rotation, color).
1 paper · 0 benchmarks
SimpleStories is a dataset of >2 million model-generated short stories.
1 paper · 0 benchmarks
Electromagnetic (EM) showers simulated dataset.
1 paper · 0 benchmarks
Simulated pulse Doppler radar signatures for four classes of helicopter-like targets.
1 paper · 0 benchmarks
SinGAN-Seg-polyps is a synthetic dataset for polyp segmentation consisting of 10,000 synthetic polyps and masks.
1 paper · 0 benchmarks
Single Point Corn Yield Data (Single Point Corn Yield Data - Weather, Soil, Cultivation Area, and Yield for Precision Agriculture)
This data comprises processed weather, soil, yield, and cultivation area for corn yield prediction in Sub-Sahara Africa, with emphasis on Nigeria.
1 paper · 0 benchmarks
Dataset of 374 photos of hand-drawn sketches of App Inventor apps used for development of the Sketch2aia model for automatic generation of App Inventor wireframes from hand-drawn sketches.
1 paper · 0 benchmarks
SketchyVR (Sparse 3D VR Sketches)
We present the first fine-grained dataset of 1,497 3D VR sketch and 3D shape pairs for 1,005 chair shapes with large shapes diversity from the ShapeNetCore dataset from 50 participants.
1 paper · 0 benchmarks
Skit-S2I (Skit-S2I: An Indian Accented Speech to Intent dataset)
This dataset for Intent classification from human speech covers 14 coarse-grained intents from the Banking domain.
1 paper · 0 benchmarks
Data related to 1040 patients with Covid-19 admitted to hospitals in Iran have been collected.
1 paper · 0 benchmarks
SoccerNet-Echoes (SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset.
1 paper · 0 benchmarks
The SoccerTrack dataset comprises top-view and wide-view video footage annotated with bounding boxes.
1 paper · 0 benchmarks
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity 🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain Dataset Statistics | Category | Samples | Description |…
1 paper · 0 benchmarks
The SNS data (Valente et al., 2013) is a four-wave survey conducted in Los Angeles county, the United States, which features a sample of 1,795 high-school students.
1 paper · 0 benchmarks
The scene derives from photo-realistic HM3D datasets.
1 paper · 0 benchmarks
The scene derives from photo-realistic MP3D datasets.
1 paper · 0 benchmarks
SoliDiffy Differencing Contract Pairs and Edit Scripts Dataset The project creates and maintains two main datasets to assist with research and evaluation of Solidity smart contract differencing: Mutated Contracts Dataset: The mutated…
1 paper · 0 benchmarks
Songdo Traffic (Songdo Traffic: High Accuracy Georeferenced Vehicle Trajectories from a Large-Scale Study in a Smart City)
The Songdo Traffic dataset delivers precisely georeferenced vehicle trajectories captured through high-altitude bird's-eye view (BeV) drone footage over Songdo International Business District, South Korea.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.