Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 8 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 337–384 of 3,130
Quora Question Pairs (QQP) dataset consists of over 400,000 question pairs, and each question pair is annotated with a binary value indicating whether the two questions are paraphrase of each other.
55 papers · 8 benchmarks
The Shifts Dataset is a dataset for evaluation of uncertainty estimates and robustness to distributional shift.
55 papers · 1 benchmark
XTREME (Cross-Lingual Transfer Evaluation of Multilingual Encoders)
The Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark was introduced to encourage more research on multilingual transfer learning,.
55 papers · 2 benchmarks
Emotion-cause pair extraction (ECPE) aims to extract the potential pairs of emotions and corresponding causes in a document.
55 papers · 0 benchmarks
ASSET is a new dataset for assessing sentence simplification in English.
54 papers · 1 benchmark
C3 is a free-form multiple-Choice Chinese machine reading Comprehension dataset.
54 papers · 0 benchmarks
CrossTask dataset contains instructional videos, collected for 83 different tasks.
54 papers · 1 benchmark
DAQUAR (DAtaset for QUestion Answering on Real-world images) is a dataset of human question answer pairs about images.
54 papers · 0 benchmarks
The DDIExtraction 2013 task relies on the DDI corpus which contains MedLine abstracts on drug-drug interactions as well as documents describing drug-drug interactions from the DrugBank database.
54 papers · 3 benchmarks
The Multi-Domain Sentiment Dataset contains product reviews taken from Amazon.com from many product types (domains).
54 papers · 1 benchmark
Room-Across-Room (RxR) is a multilingual dataset for Vision-and-Language Navigation (VLN) for Matterport3D environments.
54 papers · 1 benchmark
WikiSum is a dataset based on English Wikipedia and suitable for a task of multi-document abstractive summarization.
54 papers · 0 benchmarks
This corpus includes annotations of cancer-related PubMed articles, covering 3 full papers (PMID:24651010, PMID:11777939, PMID:15630473) as well as the result sections of 46 additional PubMed papers.
53 papers · 1 benchmark
DRCD (Delta Reading Comprehension Dataset)
Delta Reading Comprehension Dataset (DRCD) is an open domain traditional Chinese machine reading comprehension (MRC) dataset.
53 papers · 0 benchmarks
Large language models (LLMs), after being aligned with vision models and integrated into vision-language models (VLMs), can bring impressive improvement in image reasoning tasks.
53 papers · 1 benchmark
MLDoc (Multilingual Document Classification Corpus)
Multilingual Document Classification Corpus (MLDoc) is a cross-lingual document classification dataset covering English, German, French, Spanish, Italian, Russian, Japanese and Chinese.
53 papers · 8 benchmarks
The MMVP (Multimodal Visual Patterns) Benchmark focuses on identifying "CLIP-blind pairs" – images that appear similar to the CLIP model despite having clear visual differences.
53 papers · 1 benchmark
PAQ (Probably Asked Questions)
Probably Asked Questions (PAQ) is a very large resource of 65M automatically-generated QA-pairs.
53 papers · 0 benchmarks
VoiceBank + DEMAND (Noisy speech database for training speech enhancement algorithms and TTS models)
VoiceBank+DEMAND is a noisy speech database for training speech enhancement algorithms and TTS models.
53 papers · 1 benchmark
EntailmentBank is a dataset that contains multistep entailment trees.
52 papers · 0 benchmarks
InfographicVQA is a dataset that comprises a diverse collection of infographics along with natural language questions and answers annotations.
52 papers · 1 benchmark
The Machine Translation of Noisy Text (MTNT) dataset is a Machine Translation dataset that consists of noisy comments on Reddit and professionally sourced translation.
52 papers · 0 benchmarks
Pick-a-Pic dataset was created by logging user interactions with the Pick-a-Pic web application for text-to image generation.
52 papers · 0 benchmarks
There exist previous works [6, 10] that constructed referring segmentation datasets for videos.
52 papers · 3 benchmarks
The Weibo NER dataset is a Chinese Named Entity Recognition dataset drawn from the social media website Sina Weibo.
52 papers · 2 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
ACE 2004 (ACE 2004 Multilingual Training Corpus)
ACE 2004 Multilingual Training Corpus contains the complete set of English, Arabic and Chinese training data for the 2004 Automatic Content Extraction (ACE) technology evaluation.
51 papers · 6 benchmarks
QReCC contains 14K conversations with 81K question-answer pairs.
51 papers · 0 benchmarks
The Tumblr GIF (TGIF) dataset contains 100K animated GIFs and 120K sentences describing visual content of the animated GIFs.
51 papers · 1 benchmark
TOFU (Task of Fictitious Unlearning)
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks.
51 papers · 0 benchmarks
NomBank is an annotation project at New York University that is related to the PropBank project at the University of Colorado.
50 papers · 0 benchmarks
PlotQA is a VQA dataset with 28.9 million question-answer pairs grounded over 224,377 plots on data from real-world sources and questions based on crowd-sourced question templates.
50 papers · 5 benchmarks
Quoref is a QA dataset which tests the coreferential reasoning capability of reading comprehension systems.
50 papers · 0 benchmarks
emrQA has 1 million question-logical form and 400,000+ questionanswer evidence pairs.
50 papers · 0 benchmarks
CMRC 2018 (Chinese Machine Reading Comprehension 2018)
CMRC 2018 is a dataset for Chinese Machine Reading Comprehension.
49 papers · 0 benchmarks
DVQA (Data Visualizations via Question Answering)
DVQA is a synthetic question-answering dataset on images of bar-charts.
49 papers · 1 benchmark
MKQA (Multilingual Knowledge Questions and Answers)
Multilingual Knowledge Questions and Answers (MKQA) is an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically diverse languages (260k question-answer pairs in total).
49 papers · 0 benchmarks
The KIT Motion-Language is a dataset linking human motion and natural language.
48 papers · 2 benchmarks
MedMentions is a new manually annotated resource for the recognition of biomedical concepts.
48 papers · 1 benchmark
TQA (Textbook Question Answering)
The TextbookQuestionAnswering (TQA) dataset is drawn from middle school science curricula.
48 papers · 1 benchmark
BillSum is the first dataset for summarization of US Congressional and California state bills.
47 papers · 2 benchmarks
ECHR is an English legal judgment prediction dataset of cases from the European Court of Human Rights (ECHR).
47 papers · 1 benchmark
MultiCoNER is a large multilingual dataset (11 languages) for Named Entity Recognition.
47 papers · 0 benchmarks
QUASAR (QUestion Answering by Search And Reading)
The Question Answering by Search And Reading (QUASAR) is a large-scale dataset consisting of QUASAR-S and QUASAR-T.
47 papers · 1 benchmark
Reddit TIFU dataset is a newly collected Reddit dataset, where TIFU denotes the name of /r/tifu subbreddit.
47 papers · 1 benchmark
ScreenSpot Evaluation Benchmark ScreenSpot is an evaluation benchmark for GUI grounding, comprising over 1,200 instructions from various environments, including iOS, Android, macOS, Windows, and Web.
47 papers · 1 benchmark
CoS-E (Commonsense Explanations Dataset)
CoS-E consists of human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations Source: Explain Yourself!
46 papers · 0 benchmarks
FEVEROUS (Fact Extraction and VERification Over Unstructured and Structured information)
FEVEROUS (Fact Extraction and VERification Over Unstructured and Structured information) is a fact verification dataset which consists of 87,026 verified claims.
46 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.