Home › Datasets › language › English

English datasets

archive 2025-07-28

3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 68 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

English datasets 3217–3264 of 3,998

SAT-MTB-VSR is a large-scale dataset for satellite video super-resolution made from original videos of Jilin-1, which is a subset of the satellite video multitasking dataset SAT-MTB.
1 paper · 1 benchmark
SBCoseg (SBCoseg Dataset)
The SBCoseg dataset includes 889 groups of images and each group consists of 18 images with a common object, leading to 16002 images in total.
1 paper · 1 benchmark
SC2EGSet: StarCraft II Esport Game State Dataset Pre-processed data that was generated from the SC2ReSet: StarCraft II Esports Replaypack Set Data Modeling Our aplication programing interface (API) implementation supports downloading,…
1 paper · 0 benchmarks
Raw StarCraft II data is subject to processing under the Blizzard end user license agreement (EULA), and in special cases Blizzard AI and Machine Learning License may be applied.
1 paper · 0 benchmarks
SCG (SCG Dataset from Graph Neural Networks in Supply Chain Analytics and Optimization: Concepts, Perspectives, Dataset & Benchmarks)
Abstract: Graph Neural Networks (GNNs) have recently gained traction in transportation, bioinformatics, language and image processing, but research on their application to supply chain management remains limited.
1 paper · 1 benchmark
SCI (Self-Contradictory Instructions)
Large multimodal models (LMMs) excel in adhering to human instructions.
1 paper · 0 benchmarks
SCIAN (SCIAN Gold-standard for Morphological Sperm Analysis)
Dataset of sperm head images with expert-classification labels.
1 paper · 0 benchmarks
SDSS_Cosmic_Web (SDSS-IV Cosmic Web Catalog)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SE-PEF (Stack Exchange - Personalized Expert Finding)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SE-PQA (SE-PQA: a Resource for Personalized Community Question Answering)
Personalization in Information Retrieval is a topic studied for a long time.
1 paper · 0 benchmarks
SECO (Seasonal Contrast)
Read more about the dataset here: https://github.com/ServiceNow/seasonal-contrast
1 paper · 0 benchmarks
Click to add a brisef description of the datdaset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A Benchmark Dataset for Deep Learning-based Methods for 3D Topology Optimization.
1 paper · 0 benchmarks
SF20K (Short-Films 20K)
Short-Films 20K (SF20K) is the largest publicly available movie dataset.
1 paper · 0 benchmarks
SG-NLG (Schema-Guided Natural Language Generation)
The SG-NLG dataset is a pre-processed version of the DSTC8 Schema-Guided Dialogue SGD dataset, designed specifically for data-to-text Natural Language Generation (NLG).
1 paper · 0 benchmarks
SGXSTest (Singapore XSTest)
For testing refusal behavior in a cultural setting, we introduce SGXSTest — a set of manually curated prompts designed to measure exaggerated safety within the context of Singaporean culture.
1 paper · 0 benchmarks
SHADR (sythetic SDoH Human Annotated Demographic Robustness dataset (SHADR))
SDoH Human Annotated Demoographic Robustness (SHADR) Dataset Overview The Social determinants of health (SDoH) play a pivotal role in determining patient outcomes.
1 paper · 0 benchmarks
SHD - Adding (Spiking Heidelberg Digits - Adding)
This dataset is based on the Spiking Heidelberg Digits (SHD) dataset.
1 paper · 1 benchmark
A large collection of human-written natural language questions and their corresponding SPARQL queries over federated bioinformatics knowledge graphs (KGs) collected for several years across different research groups at the SIB Swiss…
1 paper · 0 benchmarks
SICKLE (Satellite Imagery for Cropping annotated with Keyparameter LabEls)
The availability of well-curated datasets has driven the success of Machine Learning (ML) models.
1 paper · 1 benchmark
Smartphone cameras are ubiquitous in daily life, yet their performance can be severely impacted by dirty lenses, leading to degraded image quality.
1 paper · 0 benchmarks
The Sequence labellIng evaLuatIon benChmark fOr spoken laNguagE (SILICONE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems specifically designed for spoken language.
1 paper · 1 benchmark
SKF-BLS Dataset (SKF Heterogeneous Test-rig Bearing Load Sensing Dataset)
This dataset is developed to estimate bearing loads under various operating conditions (rotational speed, axial and radial loads) using data from temperature and vibration sensors.
1 paper · 0 benchmarks
SMR IU X-Ray (Simplified Medical Reports)
This paper introduces CPIR-MR (Chained Prompting for Improved Readability of Medical Reports), a method designed to simplify complex chest X-ray reports for better patient understanding.
1 paper · 0 benchmarks
The SOLO Corpus comprises over 4 million English tweets, each of which contains at least one of the following tokens: solitude, lonely, and loneliness.
1 paper · 0 benchmarks
SOMD (SOftware Mention Detection)
The dataset contains the training and test data for the SOftware Mention Detection challenge.
1 paper · 0 benchmarks
SOMPT22 (Surveillance Oriented Multi-Pedestrian Tracking Dataset (SOMPT22))
SOMPT22 is a multi-object tracking (MOT) benchmark focused on surveillance-style pedestrian tracking.
1 paper · 0 benchmarks
SOTIF-PCOD (SOTIF-related Use Case Dataset)
SOTIF-PCOD is a dataset generated using the CARLA simulator, specifically designed for Safety of the Intended Functionality (SOTIF) research.
1 paper · 0 benchmarks
Curated QA Benchmark on State of the Union Address 2023.
1 paper · 0 benchmarks
SOTVerse is a user-defined task space of single object tracking.
1 paper · 0 benchmarks
SPAVE-28G (Signal Propagation Analyses in V2X Ecosystems (S.P.A.V.E) at 28 GHz on the NSF POWDER testbed)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SPAVE-28G on NSF POWDER (Propagation Measurements and Analyses at 28 GHz via an Autonomous Beam-Steering Platform)
This paper details the design of an autonomous alignment and tracking platform to mechanically steer directional horn antennas in a sliding correlator channel sounder setup for 28 GHz V2X propagation modeling.
1 paper · 0 benchmarks
SPIQA Dataset Card Dataset Details Dataset Name: SPIQA (Scientific Paper Image Question Answering) Paper: SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers Github: SPIQA eval and metrics code repo Dataset Summary:…
1 paper · 0 benchmarks
SPOT-10 (Animal Pattern Benchmark Dataset for Machine Learning Algorithms)
The SPOTS-10 dataset is an extensive collection of grayscale images showcasing diverse patterns found in ten animal species.
1 paper · 1 benchmark
SPRIGHT is the first, large-scale vision-language dataset that focuses on spatial relationships.
1 paper · 0 benchmarks
Synthetic Question Answering dataset in Serbian, acquired by automatic translation of SQuAD.
1 paper · 0 benchmarks
Confocal fluorescence microscopy is one of the most accessible and widely used imaging techniques for the study of biological processes at the cellular and subcellular levels.
1 paper · 0 benchmarks
SSBI Dataset (Synthetic Signature Bankcheck Images)
The Synthetic Signature Bankcheck Images (SSBI) Dataset is the first publicly available dataset of bank check images with annotations for detecting handwritten components, including names, amounts, dates, and signatures.
1 paper · 0 benchmarks
SSv2-Spatio-Temporal (Something Someting v2-Spatio-Temporal)
We use Something-Something v2 dataset to obtain the generation prompts and ground truth masks from real action videos.
1 paper · 0 benchmarks
StEduCov, a dataset annotated for stances toward online education during the COVID-19 pandemic.
1 paper · 1 benchmark
STN PLAD (STN Power Line Assets Dataset)
STN PLAD is a high-resolution and real-world image dataset of multiple high-voltage power line components.
1 paper · 1 benchmark
SUDO is a benchmark of 50 real-world malicious tasks designed to evaluate LLM-based computer agents in live desktop and web environments.
1 paper · 1 benchmark
SUDOER (System/User Dataset for Obedience Evaluation in Responses)
The dataset aims to provide system prompts and user prompts for assistant.
1 paper · 0 benchmarks
A RGB-D dataset converted from SUN-RGBD into COCO-style instance segmentation format.
1 paper · 2 benchmarks
A labeled dataset that presents fake news surrounding the conflict in Syria.
1 paper · 0 benchmarks
SVLD (Social Vision and Language Dataset)
The social vision and language dataset is a large-scale multimodal dataset designed for research into social contextual learning.
1 paper · 0 benchmarks
SVRT (Synthetic Visual Reasoning Task)
The Synthetic Visual Reasoning Test (SVRT) is a series of 23 classification problems involving images of randomly generated shapes.
1 paper · 0 benchmarks
SafeEdit encompasses 4,050 training, 2,700 validation, and 1,350 test instances.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.