Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 58 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2737–2784 of 3,130

SGXSTest (Singapore XSTest)
For testing refusal behavior in a cultural setting, we introduce SGXSTest — a set of manually curated prompts designed to measure exaggerated safety within the context of Singaporean culture.
1 paper · 0 benchmarks
SHADR (sythetic SDoH Human Annotated Demographic Robustness dataset (SHADR))
SDoH Human Annotated Demoographic Robustness (SHADR) Dataset Overview The Social determinants of health (SDoH) play a pivotal role in determining patient outcomes.
1 paper · 0 benchmarks
SHAJ (Spoken Hate in the Albanian Jargon)
This is an abusive/offensive language detection dataset for Albanian.
1 paper · 1 benchmark
A large collection of human-written natural language questions and their corresponding SPARQL queries over federated bioinformatics knowledge graphs (KGs) collected for several years across different research groups at the SIB Swiss…
1 paper · 0 benchmarks
The Sequence labellIng evaLuatIon benChmark fOr spoken laNguagE (SILICONE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems specifically designed for spoken language.
1 paper · 1 benchmark
SMR IU X-Ray (Simplified Medical Reports)
This paper introduces CPIR-MR (Chained Prompting for Improved Readability of Medical Reports), a method designed to simplify complex chest X-ray reports for better patient understanding.
1 paper · 0 benchmarks
The SOLO Corpus comprises over 4 million English tweets, each of which contains at least one of the following tokens: solitude, lonely, and loneliness.
1 paper · 0 benchmarks
SOMD (SOftware Mention Detection)
The dataset contains the training and test data for the SOftware Mention Detection challenge.
1 paper · 0 benchmarks
Curated QA Benchmark on State of the Union Address 2023.
1 paper · 0 benchmarks
SPIQA Dataset Card Dataset Details Dataset Name: SPIQA (Scientific Paper Image Question Answering) Paper: SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers Github: SPIQA eval and metrics code repo Dataset Summary:…
1 paper · 0 benchmarks
SQuAD-it is derived from the SQuAD dataset and it is obtained through semi-automatic translation of the SQuAD dataset into Italian.
1 paper · 0 benchmarks
SSD_ID (Sub-Slot Dialogue dataset id number domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 0 benchmarks
SSD_NAME (Sub-Slot Dialogue dataset name domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 1 benchmark
SSD_PLATE (Sub-Slot Dialogue dataset license plate number domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
1 paper · 0 benchmarks
SSv2-Spatio-Temporal (Something Someting v2-Spatio-Temporal)
We use Something-Something v2 dataset to obtain the generation prompts and ground truth masks from real action videos.
1 paper · 0 benchmarks
A large-scale Japanese video caption dataset consisting of 79,822 videos and 399,233 captions.
1 paper · 0 benchmarks
With the rise of social media, user-generated content has surged, and hate speech has proliferated.
1 paper · 0 benchmarks
The STR-2021 dataset has 5,500 English sentence pairs manually annotated for semantic relatedness using a comparative annotation framework.
1 paper · 0 benchmarks
SUDO is a benchmark of 50 real-world malicious tasks designed to evaluate LLM-based computer agents in live desktop and web environments.
1 paper · 1 benchmark
SUDOER (System/User Dataset for Obedience Evaluation in Responses)
The dataset aims to provide system prompts and user prompts for assistant.
1 paper · 0 benchmarks
A labeled dataset that presents fake news surrounding the conflict in Syria.
1 paper · 0 benchmarks
SVLD (Social Vision and Language Dataset)
The social vision and language dataset is a large-scale multimodal dataset designed for research into social contextual learning.
1 paper · 0 benchmarks
SWS (Smart Word Suggestions Benchmark)
Smart Word Suggestions (SWS) is a task and benchmark.
1 paper · 0 benchmarks
SaGA (The Bielefeld Speech and Gesture Alignment Corpus (SaGA))
The primary data of the SaGA corpus are made up of 25 dialogs of interlocutors (50), who engage in a spatial communication task combining direction-giving and sight description.
1 paper · 0 benchmarks
The satire dataset is a new multi-modal dataset of satirical and regular news articles.
1 paper · 0 benchmarks
Scan Entities in 3D (ScanEnts3D) is a large-scale dataset which provides explicit correspondences between 369k objects across 84k natural referentural sentences, covering 705 real-world scenes.
1 paper · 0 benchmarks
SciCo (Scientific Concept Induction Corpus)
SciCo is an expert-annotated dataset for hierarchical CDCR (cross-document coreference resolution) for concepts in scientific papers, with the goal of jointly inferring coreference clusters and hierarchy between them.
1 paper · 0 benchmarks
ScienceExamCER is a collection of resources for studying explanation-centered inference, including explanation graphs for 1,680 questions, with 4,950 tablestore rows, and other analyses of the knowledge required to answer elementary and…
1 paper · 0 benchmarks
This resource contains 10.5 million paragraphs with associated statement labels, realized as one paragraph per file, one sentence per line.
1 paper · 0 benchmarks
Scifi TV Shows (Scifi TV Show Plot Summaries & Events)
A collection of long-running (80+ episodes) science fiction TV show synopses, scraped from Fandom.com wikis.
1 paper · 0 benchmarks
Scroll Readability Dataset contains scroll interactions of 598 participants reading advanced and elementary texts from the OneStopEnglish corpus.
1 paper · 0 benchmarks
Search4Code is a large-scale web query based dataset of code search queries for C# and Java.
1 paper · 0 benchmarks
Dataset Card for SemTabNet This dataset accompanies the following paper: Title: Statements: Universal Information Extraction from Tables with Large Language Models for ESG KPIs Authors: Lokesh Mishra, Sohayl Dhibi, Yusik Kim, Cesar…
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Sentiment Merged (SST-3, DynaSent R1/R2)
This is a dataset for 3-way sentiment classification of reviews (negative, neutral, positive).
1 paper · 1 benchmark
SentimentArcs’ reference corpus for novels consists of 25 narratives selected to create a diverse set of well recognized novels that can serve as a benchmark for future studies.
1 paper · 0 benchmarks
This is the sequential instructions dataset from Understanding the Effects of RLHF on LLM Generalisation and Diversity.
1 paper · 0 benchmarks
ShadowLink dataset is designed to evaluate the impact of entity overshadowing on the task of entity disambiguation.
1 paper · 0 benchmarks
The ShapeIt dataset introduced by Alper et al.
1 paper · 0 benchmarks
ShopTC-100K Dataset The ShopTC-100K dataset is collected using TermMiner, an open-source data collection and topic modeling pipeline introduced in the paper: Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable…
1 paper · 0 benchmarks
In this Adjudicator ScoresShort Stories and Written Reflections folder: Four files from four student participants of the contest.
1 paper · 0 benchmarks
The proposed dataset includes 1,309 short text instances from Adobe Spark.
1 paper · 0 benchmarks
ShortPersianEmo is a new data set for emotion recognition in Persian short texts.
1 paper · 1 benchmark
Procedural videos show step-by-step demonstrations of tasks like recipe preparation.
1 paper · 0 benchmarks
SimpEvalASSET is a dataset for learning learnable metrics using modern language models.
1 paper · 0 benchmarks
It consists of 32x32 pixel images of shapes with multiple attributes (size, location, rotation, color).
1 paper · 0 benchmarks
SimpleStories is a dataset of >2 million model-generated short stories.
1 paper · 0 benchmarks
Skit-S2I (Skit-S2I: An Indian Accented Speech to Intent dataset)
This dataset for Intent classification from human speech covers 14 coarse-grained intents from the Banking domain.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.