Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 11 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 481–528 of 3,130
PUBHEALTH is a comprehensive dataset for explainable automated fact-checking of public health claims.
32 papers · 0 benchmarks
Worldtree is a corpus of explanation graphs, explanatory role ratings, and associated tablestore.
32 papers · 0 benchmarks
The Extreme Summarization (XSum) dataset is a dataset for evaluation of abstractive single-document summarization systems.
32 papers · 5 benchmarks
The Actor-Action Dataset (A2D) by Xu et al.
31 papers · 1 benchmark
ArtEmis is a large-scale dataset aimed at providing a detailed understanding of the interplay between visual content, its emotional effect, and explanations for the latter in language.
31 papers · 0 benchmarks
ContractNLI is a dataset for document-level natural language inference (NLI) on contracts whose goal is to automate/support a time-consuming procedure of contract review.
31 papers · 0 benchmarks
IGLUE (Image-Grounded Language Understanding Evaluation)
The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded…
31 papers · 0 benchmarks
The Image Paragraph Captioning dataset allows researchers to benchmark their progress in generating paragraphs that tell a story about an image.
31 papers · 1 benchmark
The IMAGE-CHAT dataset is a large collection of (image, style trait for speaker A, style trait for speaker B, dialogue between A & B) tuples that we collected using crowd-workers, Each dialogue consists of consecutive turns by speaker A…
31 papers · 2 benchmarks
KVQA (Knowledge-aware VQA)
It contains manually verified 183K question-answer pairs about more than 18K persons and 24K images.
31 papers · 0 benchmarks
Multimodal C4 (MMC4) is an augmentation of the popular text-only c4 corpus with images interleaved.
31 papers · 0 benchmarks
The MQ2008 dataset is a dataset for Learning to Rank.
31 papers · 0 benchmarks
OpinionQA is a dataset for evaluating the alignment of LM opinions with those of 60 US demographic groups over topics ranging from abortion to automation.
31 papers · 0 benchmarks
PhraseCut is a dataset consisting of 77,262 images and 345,486 phrase-region pairs.
31 papers · 1 benchmark
The SemEval-2013 Task 2 dataset contains data for two subtasks: A, an expression-level subtask, and B, a message-level subtask.
31 papers · 0 benchmarks
TIMIT (TIMIT Acoustic-Phonetic Continuous Speech Corpus)
The TIMIT Acoustic-Phonetic Continuous Speech Corpus is a standard dataset used for evaluation of automatic speech recognition systems.
31 papers · 6 benchmarks
WCEP (Wikipedia Current Events Portal)
The WCEP dataset for multi-document summarization (MDS) consists of short, human-written summaries about news events, obtained from the Wikipedia Current Events Portal (WCEP), each paired with a cluster of news articles associated with an…
31 papers · 1 benchmark
AGQA (Action Genome Question Answering)
Action Genome Question Answering (AGQA) is a benchmark for compositional spatio-temporal reasoning.
30 papers · 0 benchmarks
A testbed for commonsense reasoning about entity knowledge, bridging fact-checking about entities with commonsense inferences.
30 papers · 0 benchmarks
ECtHR (European Court of Human Rights Cases)
ECtHR is a dataset comprising European Court of Human Rights cases, including annotations for paragraph-level rationales.
30 papers · 0 benchmarks
EntityQuestions is a dataset of simple, entity-rich questions based on facts from Wikidata (e.g., "Where was Arve Furset born?
30 papers · 1 benchmark
FakeNewsNet is collected from two fact-checking websites: GossipCop and PolitiFact containing news contents with labels annotated by professional journalists and experts, along with social context information.
30 papers · 0 benchmarks
The HELP dataset is an automatically created natural language inference (NLI) dataset that embodies the combination of lexical and logical inferences focusing on monotonicity (i.e., phrase replacement-based reasoning).
30 papers · 1 benchmark
Holl-E is a dataset containing movie chats wherein each response is explicitly generated by copying and/or modifying sentences from unstructured background knowledge such as plots, comments and reviews about the movie.
30 papers · 0 benchmarks
QuaRTz is a crowdsourced dataset of 3864 multiple-choice questions about open domain qualitative relationships.
30 papers · 0 benchmarks
VALUE (Video-And-Language Understanding Evaluation)
VALUE is a Video-And-Language Understanding Evaluation benchmark to test models that are generalizable to diverse tasks, domains, and datasets.
30 papers · 0 benchmarks
Video Instruction Dataset is used to train Video-ChatGPT.
30 papers · 7 benchmarks
CODAH (COmmonsense Dataset Adversarially-authored by Humans)
The COmmonsense Dataset Adversarially-authored by Humans (CODAH) is an evaluation set for commonsense question-answering in the sentence completion style of SWAG.
29 papers · 2 benchmarks
MaRVL (Multicultural Reasoning over Vision and Language)
Multicultural Reasoning over Vision and Language (MaRVL) is a dataset based on an ImageNet-style hierarchy representative of many languages and cultures (Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish).
29 papers · 1 benchmark
MassiveText is a collection of large English-language text datasets from multiple sources: web pages, books, news articles, and code.
29 papers · 0 benchmarks
Pushshift makes available all the submissions and comments posted on Reddit between June 2005 and April 2019.
29 papers · 0 benchmarks
WI-LOCNESS (Cambridge English Write & Improve & LOCNESS)
WI-LOCNESS is part of the Building Educational Applications 2019 Shared Task for Grammatical Error Correction.
29 papers · 2 benchmarks
A human-to-human Chinese dialog dataset (about 10k dialogs, 156k utterances), which contains multiple sequential dialogs for every pair of a recommendation seeker (user) and a recommender (bot).
28 papers · 0 benchmarks
EVALution dataset is evenly distributed among the three classes (hypernyms, co-hyponyms and random) and involves three types of parts of speech (noun, verb, adjective).
28 papers · 0 benchmarks
To collect How2QA for video QA task, the same set of selected video clips are presented to another group of AMT workers for multichoice QA annotation.
28 papers · 2 benchmarks
IndicCorp is a large monolingual corpora with around 9 billion tokens covering 12 of the major Indian languages.
28 papers · 0 benchmarks
KaggleDBQA (KaggleDBQA: Realistic Text-to-SQL dataset)
KaggleDBQA is a challenging cross-domain and complex evaluation dataset of real Web databases, with domain-specific data types, original formatting, and unrestricted questions.
28 papers · 1 benchmark
MR Movie Reviews is a dataset for use in sentiment-analysis experiments.
28 papers · 3 benchmarks
Consists of parallel sentences which pair 13 major languages of India with English.
28 papers · 0 benchmarks
Resume contains eight fine-grained entity categories -score from 74.5% to 86.88%.
28 papers · 1 benchmark
The WoZ 2.0 dataset is a newer dialogue state tracking dataset whose evaluation is detached from the noisy output of speech recognition systems.
28 papers · 1 benchmark
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
decaNLP (Natural Language Decathlon Benchmark)
Natural Language Decathlon Benchmark (decaNLP) is a challenge that spans ten tasks: question answering, machine translation, summarization, natural language inference, sentiment analysis, semantic role labeling, zero-shot relation…
28 papers · 0 benchmarks
The 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across 20 different newsgroups.
27 papers · 5 benchmarks
CONAN (COunter NArratives through Nichesourcing)
COunter NArratives through Nichesourcing (CONAN) is a dataset that consists of 4,078 pairs over the 3 languages.
27 papers · 0 benchmarks
CaseHOLD (Case Holdings On Legal Decisions)
CaseHOLD (Case Holdings On Legal Decisions) is a law dataset comprised of over 53,000+ multiple choice questions to identify the relevant holding of a cited case.
27 papers · 2 benchmarks
Chaos NLI is a Natural Language Inference (NLI) dataset with 100 annotations per example (for a total of 464,500 annotations) for some existing data points in the development sets of SNLI, MNLI, and Abductive NLI.
27 papers · 0 benchmarks
Evidence Inference is a corpus for this task comprising 10,000+ prompts coupled with full-text articles describing RCTs.
27 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.