Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 41 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1921–1968 of 3,130

RuWorldTree is a QA dataset with multiple-choice elementary-level science questions, which evaluate the understanding of core science facts.
2 papers · 1 benchmark
S-TEST is a benchmark for measuring the specificity of the language of pre-trained language models.
2 papers · 0 benchmarks
SART is a collection of three datasets for Similarity, Analogies and Relatedness for the Tatar language.
2 papers · 0 benchmarks
SIMARA (SIMARA: a database for key-value information extraction from full-page handwritten documents)
Description We propose a new database for information extraction from historical handwritten documents.
2 papers · 2 benchmarks
The dataset consists of source code and LLVM IR pairs generated from accepted and de-duped programming contest solutions.
2 papers · 0 benchmarks
Contents (As on March 4, 2019) -------- The text corpus contains running text from various free licensed sources.
2 papers · 0 benchmarks
SQL-Eval is an open-source PostgreSQL evaluation dataset released by Defog, constructed based on Spider.
2 papers · 1 benchmark
SSD (Sub-Slot Dialogue dataset)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
2 papers · 0 benchmarks
SSD_PHONE (Sub-Slot Dialogue dataset phone domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
2 papers · 0 benchmarks
Saint Gall dataset contains handwritten historical manuscripts written in Latin that date back to the 9th century.
2 papers · 1 benchmark
SciGen is a challenge dataset for the task of reasoning-aware data-to-text generation consisting of tables from scientific articles and their corresponding descriptions.
2 papers · 0 benchmarks
SciHTC is a dataset for hierarchical multi-label text classification (HMLTC) of scientific papers which contains 186,160 papers and 1,233 categories from the ACM CCS tree.
2 papers · 0 benchmarks
The SegmentedTables dataset is a collection of almost 2,000 tables extracted from 352 machine learning papers.
2 papers · 0 benchmarks
SemClinBr (A multi‑institutional and multi‑specialty semantically annotated corpus for Portuguese clinical NLP tasks)
Background: The high volume of research focusing on extracting patient information from electronic health records (EHRs) has led to an increase in the demand for annotated corpora, which are a precious resource for both the development and…
2 papers · 1 benchmark
The Sentimental LIAR dataset is a modified and further extended version of the LIAR extension introduced by Kirilin et al.
2 papers · 0 benchmarks
The Signal Media One-Million News Articles Dataset dataset by Signal Media was released to facilitate researching news articles.
2 papers · 0 benchmarks
LLMs' lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarcity of relevant data.
2 papers · 0 benchmarks
Social media attributions of YouTube comments (Social media attributions dataset of YouTube comments in the context of water crisis)
Data set constructed from YouTube comments (72,098 comments posted by 43,859 users on 623 relevant videos to the crisis)
2 papers · 1 benchmark
SpCQL (Text-to-CQL)
The first dataset contains annotated natural language queries (i.e.
2 papers · 0 benchmarks
Spades (Semantic PArsing of DEclarative Sentences)
Datasets Spades contains 93,319 questions derived from clueweb09 sentences.
2 papers · 0 benchmarks
This resource is designed to allow for research into Natural Language Generation.
2 papers · 0 benchmarks
Schema2QA is the first large question answering dataset over real-world Schema.org data.
2 papers · 0 benchmarks
A multimodal empathetic dialogue dataset.
2 papers · 0 benchmarks
StoryBench (StoryBench: A Multifaceted Benchmark for Continuous Story Visualization)
StoryBench is a multi-task benchmark to reliably evaluate the ability of text-to-video models to generate stories from a sequence of captions and their duration.
2 papers · 1 benchmark
Text-Vison Cross-Modal Place Recognition Dataset
2 papers · 0 benchmarks
Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study File Descriptions File | Description --- | --- commitcategorizations.csv | Categorizations for the commits in our dataset.
2 papers · 0 benchmarks
StyleKQC is a style-variant paraphrase corpus for korean questions and commands.
2 papers · 0 benchmarks
SubEdits is a human-annnoated post-editing dataset of neural machine translation outputs, compiled from in-house NMT outputs and human post-edits of subtitles form Rakuten Viki.
2 papers · 0 benchmarks
This is a discourse dataset with multiple and subjective interpretations of English conversation in the form of perceived conversation acts and intents.
2 papers · 0 benchmarks
TLDR9+ is a large-scale summarization dataset containing over 9 million training instances extracted from Reddit discussion forum.
2 papers · 1 benchmark
TOMG-Bench (Text-based Open Molecule Generation Benchmark)
In this paper, we propose Text-based Open Molecule Generation Benchmark (TOMG-Bench), the first benchmark to evaluate the open-domain molecule generation capability of LLMs.
2 papers · 1 benchmark
TUSC (Tweets from US and Canada)
Tweets from US and Canada (TUSC) is a large dataset of more than 45 million geo-located tweets posted between 2015 and 2021 from US and Canada (TUSC), especially curated for natural language analysis
2 papers · 0 benchmarks
TVL Dataset (Touch-Vision-Language Dataset)
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model.
2 papers · 0 benchmarks
Talk2Nav is a large-scale dataset with verbal navigation instructions.
2 papers · 0 benchmarks
The TaoDescribe dataset contains 2,129,187 product titles and descriptions in Chinese.
2 papers · 0 benchmarks
TempQA-WD is a benchmark dataset for temporal reasoning designed to encourage research in extending the present approaches to target a more challenging set of complex reasoning tasks.
2 papers · 1 benchmark
Text2KGBench is a benchmark to evaluate the capabilities of language models to generate KGs from natural language text guided by an ontology.
2 papers · 0 benchmarks
Tweets and items from psychological scales for sexism detection with counterfactual examples.
2 papers · 0 benchmarks
ThreatGram 101 - Extreme Telegram Data (ThreatGram 101 - Extreme Telegram Replies Data with Threat Levels)
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
A movie ticketing dialog dataset with 23,789 annotated conversations.
2 papers · 0 benchmarks
Tilde MODEL Corpus (Tilde Multilingual Open Data for European Languages)
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
ToM-in-AMC is a novel NLP benchmark, Short for Theory-of-Mind meta-learning Assessment with Movie Characters.
2 papers · 0 benchmarks
A benchmark for suppositional reasoning based on the principles of knights and knaves puzzles.
2 papers · 0 benchmarks
TwinViews-13k is a dataset of 13,855 pairs of left-leaning and right-leaning political statements, each pair matched by topic.
2 papers · 0 benchmarks
Twitch-FIFA is video-context, many-speaker dialogue dataset based on live-broadcast soccer game videos and chats from Twitch.tv.
2 papers · 0 benchmarks
The data set contains 2500 manually-stance-labeled tweets, 1250 for each candidate (Joe Biden and Donald Trump).
2 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.