Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 10 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 433–480 of 3,130

ConvFinQA (Conversational Finance Question Answering)
ConvFinQA is a dataset designed to study the chain of numerical reasoning in conversational question answering.
38 papers · 2 benchmarks
DeepFix consists of a program repair dataset (fix compiler errors in C programs).
38 papers · 1 benchmark
InsuranceQA is a question answering dataset for the insurance domain, the data stemming from the website Insurance Library.
38 papers · 0 benchmarks
Tatoeba is a free collection of example sentences with translations geared towards foreign language learners.
38 papers · 2 benchmarks
CV-Bench (Cambrian Vision-Centric Benchmark)
The Cambrian Vision-Centric Benchmark (CV-Bench) is designed to address the limitations of existing vision-centric benchmarks by providing a comprehensive evaluation framework for multimodal large language models (MLLMs).
37 papers · 0 benchmarks
HOC (Hallmarks of Cancer)
The Hallmarks of Cancer (HOC) corpus consists of 1852 PubMed publication abstracts manually annotated by experts according to the Hallmarks of Cancer taxonomy.
37 papers · 1 benchmark
SQA (SequentialQA)
The SQA dataset was created to explore the task of answering sequences of inter-related questions on HTML tables.
37 papers · 1 benchmark
The SemEval-2018 hypernym discovery evaluation benchmark (Camacho-Collados et al.
37 papers · 3 benchmarks
TextOCR is a dataset to benchmark text recognition on arbitrary shaped scene-text.
37 papers · 0 benchmarks
WMT 2018 is a collection of datasets used in shared tasks of the Third Conference on Machine Translation.
37 papers · 4 benchmarks
The Yelp Reviews Polarity dataset is obtained from the Yelp Dataset Challenge in 2015 (1,569,264 samples that have review text).
37 papers · 0 benchmarks
AdvGLUE (Adversarial GLUE)
Adversarial GLUE (AdvGLUE) is a new multi-task benchmark to quantitatively and thoroughly explore and evaluate the vulnerabilities of modern large-scale language models under various types of adversarial attacks.
36 papers · 1 benchmark
DialogRE is the first human-annotated dialogue-based relation extraction dataset, containing 1,788 dialogues originating from the complete transcripts of a famous American television situation comedy Friends.
36 papers · 1 benchmark
Doc2Dial (Doc2Dial: Document-grounded Dialogue)
For goal-oriented document-grounded dialogs, it often involves complex contexts for identifying the most relevant information, which requires better understanding of the inter-relations between conversations and documents.
36 papers · 0 benchmarks
EBM-NLP annotates PICO (Participants, Interventions, Comparisons and Outcomes) spans in clinical trial abstracts.
36 papers · 1 benchmark
Fashion-Gen consists of 293,008 high definition (1360 x 1360 pixels) fashion images paired with item descriptions provided by professional stylists.
36 papers · 0 benchmarks
InfoSeek (Visual Information Seeking)
In this project, we introduce InfoSeek, a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge.
36 papers · 2 benchmarks
MAD (Movie Audio Descriptions) is an automatically curated large-scale dataset for the task of natural language grounding in videos or natural language moment retrieval.
36 papers · 2 benchmarks
The Open Table-and-Text Question Answering (OTT-QA) dataset contains open questions which require retrieving tables and text from the web to answer.
36 papers · 1 benchmark
PanLex translates words in thousands of languages.
36 papers · 0 benchmarks
VisualMRC (VisualMRC: Machine Reading Comprehension on Document Images)
VisualMRC is a visual machine reading comprehension dataset that proposes a task: given a question and a document image, a model produces an abstractive answer.
36 papers · 1 benchmark
stocknet (stocknet-dataset)
stocknet-dataset This repository releases a comprehensive dataset for stock movement prediction from tweets and historical stock prices.
36 papers · 1 benchmark
Attribution, Relation, and Order (ARO) benchmark to systematically evaluate the ability of VLMs to understand different types of relationships, attributes, and order information.
35 papers · 0 benchmarks
CIRCO (Composed Image Retrieval on Common Objects in context)
CIRCO (Composed Image Retrieval on Common Objects in context) is an open-domain benchmarking dataset for Composed Image Retrieval (CIR) based on real-world images from COCO 2017 unlabeled set.
35 papers · 1 benchmark
The Open Entity dataset is a collection of about 6,000 sentences with fine-grained entity types annotations.
35 papers · 2 benchmarks
Spot-the-diff is a dataset consisting of 13,192 image pairs along with corresponding human provided text annotations stating the differences between the two images.
35 papers · 0 benchmarks
A new dataset of goal-oriented dialogues that are grounded in the associated documents.
35 papers · 0 benchmarks
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
GrailQA (Strongly Generalizable Question Answering)
GrailQA is a new large-scale, high-quality dataset for question answering on knowledge bases (KBQA) on Freebase with 64,331 questions annotated with both answers and corresponding logical forms in different syntax (i.e., SPARQL,…
34 papers · 4 benchmarks
The ProPara dataset is designed to train and test comprehension of simple paragraphs describing processes (e.g., photosynthesis), designed for the task of predicting, tracking, and answering questions about how entities change during the…
34 papers · 0 benchmarks
xP3 is a multilingual dataset for multitask prompted finetuning.
34 papers · 0 benchmarks
CCAligned consists of parallel or comparable web-document pairs in 137 languages aligned with English.
33 papers · 0 benchmarks
The Dialog State Tracking Challenges 2 & 3 (DSTC2&3) were research challenge focused on improving the state of the art in tracking the state of spoken dialog systems.
33 papers · 5 benchmarks
GLUCOSE is a large-scale dataset of implicit commonsense causal knowledge, encoded as causal mini-theories about the world, each grounded in a narrative context.
33 papers · 0 benchmarks
A new multitask action quality assessment (AQA) dataset, the largest to date, comprising of more than 1600 diving samples; contains detailed annotations for fine-grained action recognition, commentary generation, and estimating the AQA…
33 papers · 2 benchmarks
MeQSum is a dataset for medical question summarization.
33 papers · 1 benchmark
SciBench is a large-scale scientific problem-solving benchmark suite that aims to systematically examine the reasoning capabilities required for complex scientific problem solving.
33 papers · 0 benchmarks
SCIREX is a document level IE dataset that encompasses multiple IE tasks, including salient entity identification and document level N-ary relation identification from scientific articles.
33 papers · 2 benchmarks
WMT 2015 is a collection of datasets used in shared tasks of the Tenth Workshop on Statistical Machine Translation.
33 papers · 2 benchmarks
WMT 2020 is a collection of datasets used in shared tasks of the Fifth Conference on Machine Translation.
33 papers · 0 benchmarks
WikiEvents is a document-level event extraction benchmark dataset which includes complete event and coreference annotation.
33 papers · 1 benchmark
Dress Code is a new dataset for image-based virtual try-on composed of image pairs coming from different catalogs of YOOX NET-A-PORTER.
32 papers · 1 benchmark
IPM NEL (Derczynski IPM Named Entity Linking)
This data is for the task of named entity recognition and linking/disambiguation over tweets.
32 papers · 1 benchmark
ImageNet-P consists of noise, blur, weather, and digital distortions.
32 papers · 1 benchmark
MS-CXR (Making the Most of Text Semantics to Improve Biomedical Vision-Language Processing)
The MS-CXR dataset provides 1162 image–sentence pairs of bounding boxes and corresponding phrases, collected across eight different cardiopulmonary radiological findings, with an approximately equal number of pairs for each finding.
32 papers · 0 benchmarks
MedQuAD (Medical Question Answering Dataset)
MedQuAD includes 47,457 medical question-answer pairs created from 12 NIH websites (e.g.
32 papers · 0 benchmarks
P-Stance: A Large Dataset for Stance Detection in Political Domain 2021
32 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.