Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 13 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 577–624 of 3,130
HeadQA is a multi-choice question answering testbed to encourage research on complex reasoning.
23 papers · 1 benchmark
LitBank is an annotated dataset of 100 works of English-language fiction to support tasks in natural language processing and the computational humanities, described in more detail in the following publications: - David Bamman, Sejal Popat…
23 papers · 1 benchmark
Simplified Chinese dataset for NER in The Third International Chinese Language Processing Bakeoff (2006), provided by Microsoft Research Asia (MSRA).
23 papers · 3 benchmarks
We study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video-language understanding.
23 papers · 1 benchmark
The NetHack Learning Environment (NLE) is a Reinforcement Learning environment based on NetHack 3.6.6.
23 papers · 1 benchmark
PIE-Bench (Prompt-based Image Editing Benchmark)
PIE-Bench comprises 700 images featuring 10 distinct editing types.
23 papers · 1 benchmark
RECCON is a dataset for the task of recognizing emotion cause in conversations.
23 papers · 2 benchmarks
ShapeWorld is a new evaluation methodology and framework for multimodal deep learning models, with a focus on formal-semantic style generalization capabilities.
23 papers · 0 benchmarks
SituatedQA is an open-retrieval QA dataset where systems must produce the correct answer to a question given the temporal or geographical context.
23 papers · 0 benchmarks
mMARCO is a multilingual version of the MS MARCO passage ranking dataset comprising 8 languages that was created using machine translation.
23 papers · 0 benchmarks
AMR Bank (Abstract Meaning Representation)
The AMR Bank is a set of English sentences paired with simple, readable semantic representations.
22 papers · 1 benchmark
BookTest is a new dataset similar to the popular Children’s Book Test (CBT), however more than 60 times larger.
22 papers · 0 benchmarks
CxC (Crisscrossed Captions)
Crisscrossed Captions (CxC) contains 247,315 human-labeled annotations including positive and negative associations between image pairs, caption pairs and image-caption pairs.
22 papers · 1 benchmark
The Django dataset is a dataset for code generation comprising of 16000 training, 1000 development and 1805 test annotations.
22 papers · 1 benchmark
Do-Not-Answer is a dataset to evaluate safeguards in large language models, and deploy safer open-source LLMs at a low cost.
22 papers · 0 benchmarks
We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
22 papers · 0 benchmarks
KdConv (Knowledge-driven Conversation)
KdConv is a Chinese multi-domain Knowledge-driven Conversation dataset, grounding the topics in multi-turn conversations to knowledge graphs.
22 papers · 0 benchmarks
100 tasks from LIBERO-100 suite.
22 papers · 1 benchmark
MOROCO (MOldavian and ROmanian Dialectal COrpus)
The MOldavian and ROmanian Dialectal COrpus (MOROCO) is a corpus that contains 33,564 samples of text (with over 10 million tokens) collected from the news domain.
22 papers · 0 benchmarks
Motion-X is a large-scale 3D expressive whole-body motion dataset, which comprises 15.6M precise 3D whole-body pose annotations (i.e., SMPL-X) covering 81.1K motion sequences from massive scenes, meanwhile providing corresponding semantic…
22 papers · 1 benchmark
PIT (Paraphrase and Semantic Similarity in Twitter)
Paraphrase and Semantic Similarity in Twitter (PIT) presents a constructed Twitter Paraphrase Corpus that contains 18,762 sentence pairs.
22 papers · 1 benchmark
Fact-checking (FC) articles which contains pairs (multimodal tweet and a FC-article) from snopes.com.
22 papers · 1 benchmark
TVBench is a new benchmark specifically created to evaluate temporal understanding in video QA.
22 papers · 1 benchmark
This dataset is aimed to study the existing reading comprehension models' capability to perform temporal reasoning, and see whether they are sensitive to the temporal description in the given question.
22 papers · 0 benchmarks
TopiOCQA (pronounced Tapioca) is an open-domain conversational dataset with topic switches on Wikipedia.
22 papers · 0 benchmarks
WebSRC (WebSRC: A Dataset for Web-Based Structural Reading Comprehension)
WebSRC is a novel Web-based Structural Reading Comprehension dataset.
22 papers · 2 benchmarks
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
XLCoST (Cross-Lingual Code Snippet)
XLCoST is a benchmark dataset for cross-lingual code intelligence.
22 papers · 0 benchmarks
iSarcasmEval is the first shared task to target intended sarcasm detection: the data for this task was provided and labelled by the authors of the texts themselves.
22 papers · 0 benchmarks
iVQA (Instructional Video Question Answering)
An open-ended VideoQA benchmark that aims to: i) provide a well-defined evaluation by including five correct answer annotations per question and ii) avoid questions which can be answered without the video.
22 papers · 2 benchmarks
COVID-Fact is a FEVER-like dataset of claims concerning the COVID-19 pandemic.
21 papers · 0 benchmarks
CliCR is a new dataset for domain specific reading comprehension used to construct around 100,000 cloze queries from clinical case reports.
21 papers · 1 benchmark
Event2Mind is a corpus of 25,000 event phrases covering a diverse range of everyday events and situations.
21 papers · 2 benchmarks
FQuAD (French Question Answering Dataset)
A French Native Reading Comprehension dataset of questions and answers on a set of Wikipedia articles that consists of 25,000+ samples for the 1.0 version and 60,000+ samples for the 1.1 version.
21 papers · 1 benchmark
KLUE (Korean Language Understanding Evaluation)
Korean Language Understanding Evaluation (KLUE) benchmark is a series of datasets to evaluate natural language understanding capability of Korean language models.
21 papers · 1 benchmark
MIR-1K (Multimedia Information Retrieval lab, 1000 song clips) is a dataset designed for singing voice separation.
21 papers · 0 benchmarks
Publicly available dataset of naturally occurring factual claims for the purpose of automatic claim verification.
21 papers · 0 benchmarks
STREUSLE stands for Supersense-Tagged Repository of English with a Unified Semantics for Lexical Expressions.
21 papers · 1 benchmark
The Terms of Service dataset is a law dataset corresponding to the task of identifying whether contractual terms are potentially unfair.
21 papers · 1 benchmark
WIQA (What-If Question Answering)
The WIQA dataset V1 has 39705 questions containing a perturbation and a possible effect in the context of a paragraph.
21 papers · 0 benchmarks
Contains one million naturally occurring sentence rewrites, providing sixty times more distinct split examples and a ninety times larger vocabulary than the WebSplit corpus introduced by Narayan et al.
21 papers · 0 benchmarks
CLOTH (CLOze test by TeacHers)
The Cloze Test by Teachers (CLOTH) benchmark is a collection of nearly 100,000 4-way multiple-choice cloze-style questions from middle- and high school-level English language exams, where the answer fills a blank in a given text.
20 papers · 0 benchmarks
ETHOS (multi-labEl haTe speecH detectiOn dataSet)
ETHOS is a hate speech detection dataset.
20 papers · 2 benchmarks
The George Washington dataset contains 20 pages of letters written by George Washington and his associates in 1755 and thereby categorized into historical collection.
20 papers · 0 benchmarks
JNLPBA is a biomedical dataset that comes from the GENIA version 3.02 corpus (Kim et al., 2003).
20 papers · 2 benchmarks
MINTAKA is a complex, natural, and multilingual dataset designed for experimenting with end-to-end question-answering models.
20 papers · 0 benchmarks
MLQE-PE (Multilingual Quality Estimation and Automatic Post-editing Dataset)
The Multilingual Quality Estimation and Automatic Post-editing (MLQE-PE) Dataset is a dataset for Machine Translation (MT) Quality Estimation (QE) and Automatic Post-Editing (APE).
20 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.