Home › Datasets › language › English
English datasets
archive 2025-07-28
3,998 datasets carry the language tag "English", ordered by the archive's paper count. Page 68 of 84: 48 shown of 3,998. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
English datasets 3217–3264 of 3,998
SAT-MTB-VSR is a large-scale dataset for satellite video super-resolution made from original videos of Jilin-1, which is a subset of the satellite video multitasking dataset SAT-MTB.
1 paper · 1 benchmark
The SBCoseg dataset includes 889 groups of images and each group consists of 18 images with a common object, leading to 16002 images in total.
1 paper · 1 benchmark
SC2EGSet: StarCraft II Esport Game State Dataset Pre-processed data that was generated from the SC2ReSet: StarCraft II Esports Replaypack Set Data Modeling Our aplication programing interface (API) implementation supports downloading,…
1 paper · 0 benchmarks
Raw StarCraft II data is subject to processing under the Blizzard end user license agreement (EULA), and in special cases Blizzard AI and Machine Learning License may be applied.
1 paper · 0 benchmarks
SCG (SCG Dataset from Graph Neural Networks in Supply Chain Analytics and Optimization: Concepts, Perspectives, Dataset & Benchmarks)
Abstract: Graph Neural Networks (GNNs) have recently gained traction in transportation, bioinformatics, language and image processing, but research on their application to supply chain management remains limited.
1 paper · 1 benchmark
SCI (Self-Contradictory Instructions)
Large multimodal models (LMMs) excel in adhering to human instructions.
1 paper · 0 benchmarks
SCIAN (SCIAN Gold-standard for Morphological Sperm Analysis)
Dataset of sperm head images with expert-classification labels.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SE-PEF (Stack Exchange - Personalized Expert Finding)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SE-PQA (SE-PQA: a Resource for Personalized Community Question Answering)
Personalization in Information Retrieval is a topic studied for a long time.
1 paper · 0 benchmarks
Read more about the dataset here: https://github.com/ServiceNow/seasonal-contrast
1 paper · 0 benchmarks
Click to add a brisef description of the datdaset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A Benchmark Dataset for Deep Learning-based Methods for 3D Topology Optimization.
1 paper · 0 benchmarks
Short-Films 20K (SF20K) is the largest publicly available movie dataset.
1 paper · 0 benchmarks
SG-NLG (Schema-Guided Natural Language Generation)
The SG-NLG dataset is a pre-processed version of the DSTC8 Schema-Guided Dialogue SGD dataset, designed specifically for data-to-text Natural Language Generation (NLG).
1 paper · 0 benchmarks
For testing refusal behavior in a cultural setting, we introduce SGXSTest — a set of manually curated prompts designed to measure exaggerated safety within the context of Singaporean culture.
1 paper · 0 benchmarks
SHADR (sythetic SDoH Human Annotated Demographic Robustness dataset (SHADR))
SDoH Human Annotated Demoographic Robustness (SHADR) Dataset Overview The Social determinants of health (SDoH) play a pivotal role in determining patient outcomes.
1 paper · 0 benchmarks
This dataset is based on the Spiking Heidelberg Digits (SHD) dataset.
1 paper · 1 benchmark
A large collection of human-written natural language questions and their corresponding SPARQL queries over federated bioinformatics knowledge graphs (KGs) collected for several years across different research groups at the SIB Swiss…
1 paper · 0 benchmarks
SICKLE (Satellite Imagery for Cropping annotated with Keyparameter LabEls)
The availability of well-curated datasets has driven the success of Machine Learning (ML) models.
1 paper · 1 benchmark
Smartphone cameras are ubiquitous in daily life, yet their performance can be severely impacted by dirty lenses, leading to degraded image quality.
1 paper · 0 benchmarks
The Sequence labellIng evaLuatIon benChmark fOr spoken laNguagE (SILICONE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems specifically designed for spoken language.
1 paper · 1 benchmark
This dataset is developed to estimate bearing loads under various operating conditions (rotational speed, axial and radial loads) using data from temperature and vibration sensors.
1 paper · 0 benchmarks
This paper introduces CPIR-MR (Chained Prompting for Improved Readability of Medical Reports), a method designed to simplify complex chest X-ray reports for better patient understanding.
1 paper · 0 benchmarks
The SOLO Corpus comprises over 4 million English tweets, each of which contains at least one of the following tokens: solitude, lonely, and loneliness.
1 paper · 0 benchmarks
SOMD (SOftware Mention Detection)
The dataset contains the training and test data for the SOftware Mention Detection challenge.
1 paper · 0 benchmarks
SOMPT22 (Surveillance Oriented Multi-Pedestrian Tracking Dataset (SOMPT22))
SOMPT22 is a multi-object tracking (MOT) benchmark focused on surveillance-style pedestrian tracking.
1 paper · 0 benchmarks
SOTIF-PCOD is a dataset generated using the CARLA simulator, specifically designed for Safety of the Intended Functionality (SOTIF) research.
1 paper · 0 benchmarks
Curated QA Benchmark on State of the Union Address 2023.
1 paper · 0 benchmarks
SOTVerse is a user-defined task space of single object tracking.
1 paper · 0 benchmarks
SPAVE-28G (Signal Propagation Analyses in V2X Ecosystems (S.P.A.V.E) at 28 GHz on the NSF POWDER testbed)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This paper details the design of an autonomous alignment and tracking platform to mechanically steer directional horn antennas in a sliding correlator channel sounder setup for 28 GHz V2X propagation modeling.
1 paper · 0 benchmarks
SPIQA Dataset Card Dataset Details Dataset Name: SPIQA (Scientific Paper Image Question Answering) Paper: SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers Github: SPIQA eval and metrics code repo Dataset Summary:…
1 paper · 0 benchmarks
SPOT-10 (Animal Pattern Benchmark Dataset for Machine Learning Algorithms)
The SPOTS-10 dataset is an extensive collection of grayscale images showcasing diverse patterns found in ten animal species.
1 paper · 1 benchmark
SPRIGHT is the first, large-scale vision-language dataset that focuses on spatial relationships.
1 paper · 0 benchmarks
Synthetic Question Answering dataset in Serbian, acquired by automatic translation of SQuAD.
1 paper · 0 benchmarks
Confocal fluorescence microscopy is one of the most accessible and widely used imaging techniques for the study of biological processes at the cellular and subcellular levels.
1 paper · 0 benchmarks
The Synthetic Signature Bankcheck Images (SSBI) Dataset is the first publicly available dataset of bank check images with annotations for detecting handwritten components, including names, amounts, dates, and signatures.
1 paper · 0 benchmarks
We use Something-Something v2 dataset to obtain the generation prompts and ground truth masks from real action videos.
1 paper · 0 benchmarks
StEduCov, a dataset annotated for stances toward online education during the COVID-19 pandemic.
1 paper · 1 benchmark
STN PLAD (STN Power Line Assets Dataset)
STN PLAD is a high-resolution and real-world image dataset of multiple high-voltage power line components.
1 paper · 1 benchmark
SUDO is a benchmark of 50 real-world malicious tasks designed to evaluate LLM-based computer agents in live desktop and web environments.
1 paper · 1 benchmark
SUDOER (System/User Dataset for Obedience Evaluation in Responses)
The dataset aims to provide system prompts and user prompts for assistant.
1 paper · 0 benchmarks
A RGB-D dataset converted from SUN-RGBD into COCO-style instance segmentation format.
1 paper · 2 benchmarks
A labeled dataset that presents fake news surrounding the conflict in Syria.
1 paper · 0 benchmarks
SVLD (Social Vision and Language Dataset)
The social vision and language dataset is a large-scale multimodal dataset designed for research into social contextual learning.
1 paper · 0 benchmarks
SVRT (Synthetic Visual Reasoning Task)
The Synthetic Visual Reasoning Test (SVRT) is a series of 23 classification problems involving images of randomly generated shapes.
1 paper · 0 benchmarks
SafeEdit encompasses 4,050 training, 2,700 validation, and 1,350 test instances.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.