Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 69 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3265–3312 of 3,998
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The satire dataset is a new multi-modal dataset of satirical and regular news articles.
1 paper · 0 benchmarks
Scan Entities in 3D (ScanEnts3D) is a large-scale dataset which provides explicit correspondences between 369k objects across 84k natural referentural sentences, covering 705 real-world scenes.
1 paper · 0 benchmarks
SciCo (Scientific Concept Induction Corpus)
SciCo is an expert-annotated dataset for hierarchical CDCR (cross-document coreference resolution) for concepts in scientific papers, with the goal of jointly inferring coreference clusters and hierarchy between them.
1 paper · 0 benchmarks
This resource contains 10.5 million paragraphs with associated statement labels, realized as one paragraph per file, one sentence per line.
1 paper · 0 benchmarks
A collection of long-running (80+ episodes) science fiction TV show synopses, scraped from Fandom.com wikis.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Scroll Readability Dataset contains scroll interactions of 598 participants reading advanced and elementary texts from the OneStopEnglish corpus.
1 paper · 0 benchmarks
This data set includes all raw data (e.g., collected certificates) of the WWW 2021 paper "Security of Alerting Authorities in the WWW: Measuring Namespaces, DNSSEC, and Web PKI".
1 paper · 0 benchmarks
This dataset is obtained during an ICON project (2017-2018) in collaboration with KU Leuven (ESAT-STADIUS), UZ Leuven, UCB, Byteflies and Pilipili.
1 paper · 0 benchmarks
Dataset Card for SemTabNet This dataset accompanies the following paper: Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar…
1 paper · 1 benchmark
Test dataset for Semantic Segmentation.
1 paper · 0 benchmarks
SemanticSugarBeets, a novel and high-quality dataset containing 953 monocular RGB images and 2920 annotations of sugar beets, enables a wide range of learning tasks including object detection, semantic segmentation, instance segmentation…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SensoDat is a dataset of self-driving car simulation data (30K executed simulations).
1 paper · 0 benchmarks
The dataset is based on a debate.org crawl.
1 paper · 0 benchmarks
This is a dataset for 3-way sentiment classification of reviews (negative, neutral, positive).
1 paper · 1 benchmark
SentimentArcs’ reference corpus for novels consists of 25 narratives selected to create a diverse set of well recognized novels that can serve as a benchmark for future studies.
1 paper · 0 benchmarks
This is the sequential instructions dataset from Understanding the Effects of RLHF on LLM Generalisation and Diversity.
1 paper · 0 benchmarks
ShadowLink dataset is designed to evaluate the impact of entity overshadowing on the task of entity disambiguation.
1 paper · 0 benchmarks
The ShapeIt dataset introduced by Alper et al.
1 paper · 0 benchmarks
This repository contains documentation for the dataset that accompanies our ICPE 2025 paper, "Shaved Ice: Optimal Compute Resource Commitments for Dynamic Multi-Cloud Workloads".
1 paper · 0 benchmarks
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts 🔥 Key Features - 3000+ hours of synthetic speech - Diverse Distribution Shifts: The dataset spans 7 key distribution shifts, including: - 📖 Reading Style - 🎙️…
1 paper · 0 benchmarks
ShopTC-100K Dataset The ShopTC-100K dataset is collected using TermMiner, an open-source data collection and topic modeling pipeline introduced in the paper: Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable…
1 paper · 0 benchmarks
In this Adjudicator ScoresShort Stories and Written Reflections folder: Four files from four student participants of the contest.
1 paper · 0 benchmarks
The database consists of EEG recordings of 14 patients acquired at the Unit of Neurology and Neurophysiology of the University of Siena.
1 paper · 0 benchmarks
Procedural videos show step-by-step demonstrations of tasks like recipe preparation.
1 paper · 0 benchmarks
The SimBEV dataset is a collection of 320 scenes spread across all 11 CARLA maps and contains data from a variety of sensors, including five camera types (RGB, semantic segmentation, instance segmentation, depth, and optical flow), lidar,…
1 paper · 3 benchmarks
SimpEvalASSET is a dataset for learning learnable metrics using modern language models.
1 paper · 0 benchmarks
It consists of 32x32 pixel images of shapes with multiple attributes (size, location, rotation, color).
1 paper · 0 benchmarks
SimpleStories is a dataset of >2 million model-generated short stories.
1 paper · 0 benchmarks
Electromagnetic (EM) showers simulated dataset.
1 paper · 0 benchmarks
Simulated pulse Doppler radar signatures for four classes of helicopter-like targets.
1 paper · 0 benchmarks
SinGAN-Seg-polyps is a synthetic dataset for polyp segmentation consisting of 10,000 synthetic polyps and masks.
1 paper · 0 benchmarks
This data comprises processed weather, soil, yield, and cultivation area for corn yield prediction in Sub-Sahara Africa, with emphasis on Nigeria.
1 paper · 0 benchmarks
Dataset of 374 photos of hand-drawn sketches of App Inventor apps used for development of the Sketch2aia model for automatic generation of App Inventor wireframes from hand-drawn sketches.
1 paper · 0 benchmarks
We present the first fine-grained dataset of 1,497 3D VR sketch and 3D shape pairs for 1,005 chair shapes with large shapes diversity from the ShapeNetCore dataset from 50 participants.
1 paper · 0 benchmarks
Skit-S2I (Skit-S2I: An Indian Accented Speech to Intent dataset)
This dataset for Intent classification from human speech covers 14 coarse-grained intents from the Banking domain.
1 paper · 0 benchmarks
Data related to 1040 patients with Covid-19 admitted to hospitals in Iran have been collected.
1 paper · 0 benchmarks
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset.
1 paper · 0 benchmarks
The SoccerTrack dataset comprises top-view and wide-view video footage annotated with bounding boxes.
1 paper · 0 benchmarks
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity 🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain Dataset Statistics | Category | Samples | Description |…
1 paper · 0 benchmarks
The SNS data (Valente et al., 2013) is a four-wave survey conducted in Los Angeles county, the United States, which features a sample of 1,795 high-school students.
1 paper · 0 benchmarks
The scene derives from photo-realistic HM3D datasets.
1 paper · 0 benchmarks
The scene derives from photo-realistic MP3D datasets.
1 paper · 0 benchmarks
SoliDiffy Differencing Contract Pairs and Edit Scripts Dataset The project creates and maintains two main datasets to assist with research and evaluation of Solidity smart contract differencing: Mutated Contracts Dataset: The mutated…
1 paper · 0 benchmarks
Songdo Traffic (Songdo Traffic: High Accuracy Georeferenced Vehicle Trajectories from a Large-Scale Study in a Smart City)
The Songdo Traffic dataset delivers precisely georeferenced vehicle trajectories captured through high-altitude bird's-eye view (BeV) drone footage over Songdo International Business District, South Korea.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.