Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 41 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1921–1968 of 3,130
RuWorldTree is a QA dataset with multiple-choice elementary-level science questions, which evaluate the understanding of core science facts.
2 papers · 1 benchmark
S-TEST is a benchmark for measuring the specificity of the language of pre-trained language models.
2 papers · 0 benchmarks
SART is a collection of three datasets for Similarity, Analogies and Relatedness for the Tatar language.
2 papers · 0 benchmarks
SIMARA (SIMARA: a database for key-value information extraction from full-page handwritten documents)
Description We propose a new database for information extraction from historical handwritten documents.
2 papers · 2 benchmarks
The dataset consists of source code and LLVM IR pairs generated from accepted and de-duped programming contest solutions.
2 papers · 0 benchmarks
Contents (As on March 4, 2019) -------- The text corpus contains running text from various free licensed sources.
2 papers · 0 benchmarks
SQL-Eval is an open-source PostgreSQL evaluation dataset released by Defog, constructed based on Spider.
2 papers · 1 benchmark
SSD (Sub-Slot Dialogue dataset)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
2 papers · 0 benchmarks
SSD_PHONE (Sub-Slot Dialogue dataset phone domain)
SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".
2 papers · 0 benchmarks
Saint Gall dataset contains handwritten historical manuscripts written in Latin that date back to the 9th century.
2 papers · 1 benchmark
SciGen is a challenge dataset for the task of reasoning-aware data-to-text generation consisting of tables from scientific articles and their corresponding descriptions.
2 papers · 0 benchmarks
SciHTC is a dataset for hierarchical multi-label text classification (HMLTC) of scientific papers which contains 186,160 papers and 1,233 categories from the ACM CCS tree.
2 papers · 0 benchmarks
The SegmentedTables dataset is a collection of almost 2,000 tables extracted from 352 machine learning papers.
2 papers · 0 benchmarks
SemClinBr (A multi‑institutional and multi‑specialty semantically annotated corpus for Portuguese clinical NLP tasks)
Background: The high volume of research focusing on extracting patient information from electronic health records (EHRs) has led to an increase in the demand for annotated corpora, which are a precious resource for both the development and…
2 papers · 1 benchmark
The Sentimental LIAR dataset is a modified and further extended version of the LIAR extension introduced by Kirilin et al.
2 papers · 0 benchmarks
The Signal Media One-Million News Articles Dataset dataset by Signal Media was released to facilitate researching news articles.
2 papers · 0 benchmarks
LLMs' lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarcity of relevant data.
2 papers · 0 benchmarks
Data set constructed from YouTube comments (72,098 comments posted by 43,859 users on 623 relevant videos to the crisis)
2 papers · 1 benchmark
The first dataset contains annotated natural language queries (i.e.
2 papers · 0 benchmarks
Spades (Semantic PArsing of DEclarative Sentences)
Datasets Spades contains 93,319 questions derived from clueweb09 sentences.
2 papers · 0 benchmarks
This resource is designed to allow for research into Natural Language Generation.
2 papers · 0 benchmarks
Schema2QA is the first large question answering dataset over real-world Schema.org data.
2 papers · 0 benchmarks
A multimodal empathetic dialogue dataset.
2 papers · 0 benchmarks
StoryBench (StoryBench: A Multifaceted Benchmark for Continuous Story Visualization)
StoryBench is a multi-task benchmark to reliably evaluate the ability of text-to-video models to generate stories from a sequence of captions and their duration.
2 papers · 1 benchmark
Text-Vison Cross-Modal Place Recognition Dataset
2 papers · 0 benchmarks
Challenges in Migrating Imperative Deep Learning Programs to Graph Execution: An Empirical Study File Descriptions File | Description --- | --- commitcategorizations.csv | Categorizations for the commits in our dataset.
2 papers · 0 benchmarks
StyleKQC is a style-variant paraphrase corpus for korean questions and commands.
2 papers · 0 benchmarks
SubEdits is a human-annnoated post-editing dataset of neural machine translation outputs, compiled from in-house NMT outputs and human post-edits of subtitles form Rakuten Viki.
2 papers · 0 benchmarks
This is a discourse dataset with multiple and subjective interpretations of English conversation in the form of perceived conversation acts and intents.
2 papers · 0 benchmarks
TLDR9+ is a large-scale summarization dataset containing over 9 million training instances extracted from Reddit discussion forum.
2 papers · 1 benchmark
TOMG-Bench (Text-based Open Molecule Generation Benchmark)
In this paper, we propose Text-based Open Molecule Generation Benchmark (TOMG-Bench), the first benchmark to evaluate the open-domain molecule generation capability of LLMs.
2 papers · 1 benchmark
TUSC (Tweets from US and Canada)
Tweets from US and Canada (TUSC) is a large dataset of more than 45 million geo-located tweets posted between 2015 and 2021 from US and Canada (TUSC), especially curated for natural language analysis
2 papers · 0 benchmarks
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model.
2 papers · 0 benchmarks
Talk2Nav is a large-scale dataset with verbal navigation instructions.
2 papers · 0 benchmarks
The TaoDescribe dataset contains 2,129,187 product titles and descriptions in Chinese.
2 papers · 0 benchmarks
TempQA-WD is a benchmark dataset for temporal reasoning designed to encourage research in extending the present approaches to target a more challenging set of complex reasoning tasks.
2 papers · 1 benchmark
Text2KGBench is a benchmark to evaluate the capabilities of language models to generate KGs from natural language text guided by an ontology.
2 papers · 0 benchmarks
Tweets and items from psychological scales for sexism detection with counterfactual examples.
2 papers · 0 benchmarks
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
A movie ticketing dialog dataset with 23,789 annotated conversations.
2 papers · 0 benchmarks
Tilde MODEL Corpus is a multilingual corpora for European languages – particularly focused on the smaller languages.
2 papers · 0 benchmarks
ToM-in-AMC is a novel NLP benchmark, Short for Theory-of-Mind meta-learning Assessment with Movie Characters.
2 papers · 0 benchmarks
A benchmark for suppositional reasoning based on the principles of knights and knaves puzzles.
2 papers · 0 benchmarks
TwinViews-13k is a dataset of 13,855 pairs of left-leaning and right-leaning political statements, each pair matched by topic.
2 papers · 0 benchmarks
Twitch-FIFA is video-context, many-speaker dialogue dataset based on live-broadcast soccer game videos and chats from Twitch.tv.
2 papers · 0 benchmarks
The data set contains 2500 manually-stance-labeled tweets, 1250 for each candidate (Joe Biden and Donald Trump).
2 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.