Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 3 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 97–144 of 3,130
CodeXGLUE is a benchmark dataset and open challenge for code intelligence.
205 papers · 10 benchmarks
LAION 5B is a large-scale dataset for research purposes consisting of 5,85B CLIP-filtered image-text pairs.
205 papers · 0 benchmarks
TACRED (The TAC Relation Extraction Dataset)
TACRED is a large-scale relation extraction dataset with 106,264 examples built over newswire and web text from the corpus used in the yearly TAC Knowledge Base Population (TAC KBP) challenges.
204 papers · 2 benchmarks
COCO Captions contains over one and a half million captions describing over 330,000 images.
203 papers · 4 benchmarks
Clotho is an audio captioning dataset, consisting of 4981 audio samples, and each audio sample has five captions (a total of 24 905 captions).
202 papers · 3 benchmarks
HumanML3D is a 3D human motion-language dataset that originates from a combination of HumanAct12 and Amass dataset.
201 papers · 2 benchmarks
YouCook2 is the largest task-oriented, instructional video dataset in the vision community.
198 papers · 7 benchmarks
WikiQA (Wikipedia open-domain Question Answering)
The WikiQA corpus is a publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.
196 papers · 2 benchmarks
BioASQ (Biomedical Semantic Indexing and Question Answering)
BioASQ is a question answering dataset.
192 papers · 1 benchmark
BC5CDR (BioCreative V CDR corpus)
BC5CDR corpus consists of 1500 PubMed articles with 4409 annotated chemicals, 5818 diseases and 3116 chemical-disease interactions.
191 papers · 4 benchmarks
CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) is the largest dataset of sentence-level sentiment analysis and emotion recognition in online videos.
190 papers · 3 benchmarks
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance.
190 papers · 1 benchmark
FewRel (Few-Shot Relation Classification Dataset)
The FewRel (Few-Shot Relation Classification Dataset) contains 100 relations and 70,000 instances from Wikipedia.
189 papers · 3 benchmarks
LRW (Lip Reading in the Wild)
The Lip Reading in the Wild (LRW) dataset a large-scale audio-visual database that contains 500 different words from over 1,000 speakers.
188 papers · 8 benchmarks
SGD (Schema-Guided Dialogue)
The Schema-Guided Dialogue (SGD) dataset consists of over 20k annotated multi-domain, task-oriented conversations between a human and a virtual assistant.
186 papers · 2 benchmarks
A new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality.
183 papers · 1 benchmark
The SciQ dataset contains 13,679 crowdsourced science exam questions about Physics, Chemistry and Biology, among others.
183 papers · 1 benchmark
APPS (Automated Programming Progress Standard)
The APPS dataset consists of problems collected from different open-access coding websites such as Codeforces, Kattis, and more.
180 papers · 1 benchmark
Dataset of hate speech annotated on Internet forum posts in English at sentence-level.
180 papers · 1 benchmark
IFEval (Instruction Following Evaluation Datset)
This dataset evaluates instruction following ability of large language models.
180 papers · 1 benchmark
Long-range arena (LRA) is an effort toward systematic evaluation of efficient transformer models.
180 papers · 1 benchmark
FUNSD (Form Understanding in Noisy Scanned Documents)
Form Understanding in Noisy Scanned Documents (FUNSD) comprises 199 real, fully annotated, scanned forms.
179 papers · 3 benchmarks
VCR (Visual Commonsense Reasoning)
Visual Commonsense Reasoning (VCR) is a large-scale dataset for cognition-level visual understanding.
179 papers · 13 benchmarks
The AI2’s Reasoning Challenge (ARC) dataset is a multiple-choice question-answering dataset, containing questions from science exams from grade 3 to grade 9.
178 papers · 3 benchmarks
QuAC (Question Answering in Context)
Question Answering in Context is a large-scale dataset that consists of around 14K crowdsourced Question Answering dialogs with 98K question-answer pairs in total.
178 papers · 1 benchmark
WMT 2016 is a collection of datasets used in shared tasks of the First Conference on Machine Translation.
178 papers · 16 benchmarks
The Hateful Memes data set is a multimodal dataset for hateful meme detection (image + text) that contains 10,000+ new multimodal examples created by Facebook AI.
177 papers · 3 benchmarks
NELL (Never Ending Language Learning)
NELL is a dataset built from the Web via an intelligent agent called Never-Ending Language Learner.
177 papers · 2 benchmarks
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
174 papers · 3 benchmarks
R2R is a dataset for visually-grounded natural language navigation in real buildings.
174 papers · 2 benchmarks
ALFRED (Action Learning From Realistic Environments and Directives)
ALFRED (Action Learning From Realistic Environments and Directives), is a new benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks.
173 papers · 0 benchmarks
PAWS-X contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean.
172 papers · 0 benchmarks
ATOMIC is an atlas of everyday commonsense reasoning, organized through 877k textual descriptions of inferential knowledge.
170 papers · 0 benchmarks
LAION-400M is a dataset with CLIP-filtered 400 million image-text pairs, their CLIP embeddings and kNN indices that allow efficient similarity search.
169 papers · 1 benchmark
SentEval is a toolkit for evaluating the quality of universal sentence representations.
168 papers · 1 benchmark
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
CORD-19 is a free resource of tens of thousands of scholarly articles about COVID-19, SARS-CoV-2, and related coronaviruses for use by the global research community.
163 papers · 1 benchmark
SWAG (Situations With Adversarial Generations)
Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine").
163 papers · 2 benchmarks
This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS.
162 papers · 6 benchmarks
MultiRC (Multi-Sentence Reading Comprehension)
MultiRC (Multi-Sentence Reading Comprehension) is a dataset of short paragraphs and multi-sentence questions, i.e., questions that can be answered by combining information from multiple sentences of the paragraph.
162 papers · 1 benchmark
MathQA significantly enhances the AQuA dataset with fully-specified operational programs.
159 papers · 1 benchmark
PAWS (Paraphrase Adversaries from Word Scrambling)
Paraphrase Adversaries from Word Scrambling (PAWS) is a dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of…
159 papers · 0 benchmarks
Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset.
159 papers · 6 benchmarks
ELI5 is a dataset for long-form question answering.
158 papers · 1 benchmark
At the end of 2017 the Civil Comments platform shut down and chose make their ~2m public comments from their platform available in a lasting open archive so that researchers could understand and improve civility in online conversations for…
156 papers · 1 benchmark
DocRED (Document-Level Relation Extraction Dataset) is a relation extraction dataset constructed from Wikipedia and Wikidata.
155 papers · 4 benchmarks
MIND (MIcrosoft News Dataset)
MIcrosoft News Dataset (MIND) is a large-scale dataset for news recommendation research.
155 papers · 0 benchmarks
MTEB (Massive Text Embedding Benchmark)
MTEB is a benchmark that spans 8 embedding tasks covering a total of 56 datasets and 112 languages.
155 papers · 6 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.