Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 12 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 529–576 of 3,130
We contribute an IntentQA dataset with diverse intents in daily social activities.
27 papers · 2 benchmarks
KPTimes is a large-scale dataset of news texts paired with editor-curated keyphrases.
27 papers · 3 benchmarks
LDC2017T10 (Abstract Meaning Representation (AMR) Annotation Release 2.0)
Abstract Meaning Representation (AMR) Annotation Release 2.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
27 papers · 1 benchmark
Multi-Modal-CelebA-HQ is a large-scale face image dataset that has 30,000 high-resolution face images selected from the CelebA dataset by following CelebA-HQ.
27 papers · 3 benchmarks
ProofNet is a benchmark for autoformalization and formal proving of undergraduate-level mathematics.
27 papers · 0 benchmarks
Sentiment analysis of codemixed tweets.
27 papers · 0 benchmarks
TG-ReDial is a a topic-guided conversational recommendation dataset for research on conversational/interactive recommender systems.
27 papers · 0 benchmarks
We have created three new Reading Comprehension datasets constructed using an adversarial model-in-the-loop.
26 papers · 2 benchmarks
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
DUDE (Document UnderstanDing of Everything)
DUDE is formulated as an instance of Document Question Answering (DocQA) to evaluate how well current solutions deal with multi-page documents, if they can navigate and reason over the layout, and if they can generalize these skills to…
26 papers · 0 benchmarks
FairytaleQA is a dataset focusing on narrative comprehension of kindergarten to eighth-grade students.
26 papers · 2 benchmarks
The GenericsKB contains 3.4M+ generic sentences about the world, i.e., sentences expressing general truths such as "Dogs bark," and "Trees remove carbon dioxide from the atmosphere." Generics are potentially useful as a knowledge source…
26 papers · 0 benchmarks
IIRC (Incomplete Information Reading Comprehension)
Contains more than 13K questions over paragraphs from English Wikipedia that provide only partial information to answer them, with the missing information occurring in one or more linked documents.
26 papers · 0 benchmarks
MultiDoc2Dial (MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents)
MultiDoc2Dial is a new task and dataset on modeling goal-oriented dialogues grounded in multiple documents.
26 papers · 0 benchmarks
PixelHelp includes 187 multi-step instructions of 4 task categories deined in https://support.google.com/pixelphone and annotated by human.
26 papers · 0 benchmarks
Screen2Words is a large-scale screen summarization dataset annotated by human workers.
26 papers · 0 benchmarks
The SentiCap dataset contains several thousand images with captions with positive and negative sentiments.
26 papers · 0 benchmarks
VALSE (VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena)
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic…
26 papers · 12 benchmarks
WikiReading is a large-scale natural language understanding task and publicly-available dataset with 18 million instances.
26 papers · 0 benchmarks
XStoryCloze consists of the professionally translated version of the English StoryCloze dataset (Spring 2016 version) to 10 non-English languages.
26 papers · 0 benchmarks
BioRED is a first-of-its-kind biomedical relation extraction dataset with multiple entity types (e.g.
25 papers · 3 benchmarks
CH-SIMS is a Chinese single- and multimodal sentiment analysis dataset which contains 2,281 refined video segments in the wild with both multimodal and independent unimodal annotations.
25 papers · 1 benchmark
CrisisMMD is a large multi-modal dataset collected from Twitter during different natural disasters.
25 papers · 0 benchmarks
CrossWOZ is the first large-scale Chinese Cross-Domain Wizard-of-Oz task-oriented dataset.
25 papers · 0 benchmarks
ELEVATER (Evaluation of Language-augmented Visual Task-level Transfer)
The ELEVATER benchmark is a collection of resources for training, evaluating, and analyzing language-image models on image classification and object detection.
25 papers · 2 benchmarks
HONEST (Hurtful Sentence Completion in English Language Models)
The HONEST dataset is a template-based corpus for testing the hurtfulness of sentence completions in language models (e.g., BERT) in six different languages (English, Italian, French, Portuguese, Romanian, and Spanish).
25 papers · 1 benchmark
The MULTEXT-East resources are a multilingual dataset for language engineering research and development.
25 papers · 0 benchmarks
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
TOPv2 (Task Oriented Parsing v2)
Task Oriented Parsing v2 (TOPv2) representations for intent-slot based dialog systems.
25 papers · 0 benchmarks
ZESHEL is a zero-shot entity linking dataset, which places more emphasis on understanding the unstructured descriptions of entities to resolve the ambiguity of mentions on four unseen domains.
25 papers · 1 benchmark
A benchmark dataset for the Aspect Sentiment Triplet Extraction, an updated version of ASTE-Data-V1.
24 papers · 1 benchmark
FlickrStyle10K is collected and built on Flickr30K image caption dataset.
24 papers · 2 benchmarks
GeoS is a dataset for automatic math problem solving.
24 papers · 1 benchmark
InterHuman is a multimodal dataset, named InterHuman.
24 papers · 1 benchmark
KELM is a large-scale synthetic corpus of Wikidata KG as natural text.
24 papers · 0 benchmarks
MCScript is used as the official dataset of SemEval2018 Task11.
24 papers · 0 benchmarks
Moral Stories is a crowd-sourced dataset of structured narratives that describe normative and norm-divergent actions taken by individuals to accomplish certain intentions in concrete situations, and their respective consequences.
24 papers · 0 benchmarks
The Natural Stories dataset consists of English texts edited to contain many low-frequency syntactic constructions while still sounding fluent to native speakers.
24 papers · 0 benchmarks
ROPES (Reasoning Over Paragraph Effects in Situations)
ROPES is a QA dataset which tests a system's ability to apply knowledge from a passage of text to a new situation.
24 papers · 0 benchmarks
RecipeQA is a dataset for multimodal comprehension of cooking recipes.
24 papers · 1 benchmark
SIMMC (Situated and Interactive Multimodal Conversations)
Situated Interactive MultiModal Conversations (SIMMC) is the task of taking multimodal actions grounded in a co-evolving multimodal input content in addition to the dialog history.
24 papers · 0 benchmarks
News translation is a recurring WMT task.
24 papers · 0 benchmarks
The Wiki-ZSL (Wiki Zero-Shot Learning) dataset contains 113 relations and 94,383 instances from Wikipedia.
24 papers · 1 benchmark
We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements.
23 papers · 2 benchmarks
CoSQA (Code Search and Question Answering)
CoSQA (Code Search and Question Answering) It includes 20,604 labels for pairs of natural language queries and codes, each annotated by at least 3 human annotators.
23 papers · 0 benchmarks
EURLEX57K is a new publicly available legal LMTC dataset, dubbed EURLEX57K, containing 57k English EU legislative documents from the EUR-LEX portal, tagged with ∼4.3k labels (concepts) from the European Vocabulary (EUROVOC).
23 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.