Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 20 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 913–960 of 3,130

UGIF is a multi-lingual, multi-modal UI grounded dataset for step-by-step task completion on the smartphone.
10 papers · 0 benchmarks
ViP-Bench (Making Large Multimodal Models Understand Arbitrary Visual Prompts)
ViP-Bench is a comprehensive benchmark designed to assess the capability of multimodal models in understanding visual prompts across multiple dimensions.
10 papers · 1 benchmark
VideoCC3M (Video-Conceptual-Captions)
We propose a new, scalable video-mining pipeline which transfers captioning supervision from image datasets to video and audio.
10 papers · 0 benchmarks
WikiAtomicEdits is a corpus of 43 million atomic edits across 8 languages.
10 papers · 0 benchmarks
A corpus that encompasses the complete history of conversations between contributors to Wikipedia, one of the largest online collaborative communities.
10 papers · 0 benchmarks
XFORMAL is a multilingual formal style transfer benchmark of multiple formal reformulations of informal text in Brazilian Portuguese, French, and Italian.
10 papers · 0 benchmarks
e-ViL is a benchmark for explainable vision-language tasks.
10 papers · 0 benchmarks
2012 i2b2 Temporal Relations (2012 i2b2 Temporal Relations Corpus)
The Sixth Informatics for Integrating Biology and the Bedside (i2b2) Natural Language Processing Challenge for Clinical Records focused on the temporal relations in clinical narratives.
9 papers · 2 benchmarks
Our dataset which consists of multiple indoor and outdoor experiments for up to 30 m gNB-UE link.
9 papers · 0 benchmarks
ART Dataset (Abductive Reasoning in narrative Text)
ART consists of over 20k commonsense narrative contexts and 200k explanations.
9 papers · 0 benchmarks
Is an acronym disambiguation (AD) dataset for scientific domain with 62,441 samples which is significantly larger than the previous scientific AD dataset.
9 papers · 0 benchmarks
This dataset contains 8.9M commonsense assertions extracted by the Ascent pipeline developed at the Max Planck Institute for Informatics.
9 papers · 0 benchmarks
BIMCV-COVID19+ dataset is a large dataset with chest X-ray images CXR (CR, DX) and computed tomography (CT) imaging of COVID-19 patients along with their radiographic findings, pathologies, polymerase chain reaction (PCR), immunoglobulin G…
9 papers · 0 benchmarks
Bactrian-X is a comprehensive multilingual parallel dataset of 3.4 million instruction-response pairs across 52 languages.
9 papers · 0 benchmarks
BiToD is a bilingual multi-domain dataset for end-to-end task-oriented dialogue modeling.
9 papers · 0 benchmarks
The Japanese-English business conversation corpus, namely Business Scene Dialogue corpus, was constructed in 3 steps: 1.
9 papers · 2 benchmarks
CQASUMM is a dataset for CQA (Community Question Answering) summarization, constructed from the 4.4 million Yahoo!
9 papers · 0 benchmarks
CaSiNo is a dataset of 1030 negotiation dialogues in English.
9 papers · 0 benchmarks
ClueWeb22 is the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information.
9 papers · 0 benchmarks
ComQA is a large dataset of real user questions that exhibit different challenging aspects such as compositionality, temporal reasoning, and comparisons.
9 papers · 0 benchmarks
DL-HARD (Deep Learning Hard)
Deep Learning Hard (DL-HARD) is an annotated dataset designed to more effectively evaluate neural ranking models on complex topics.
9 papers · 0 benchmarks
E-KAR (Benchmark for Explainable Knowledge-intensive Analogical Reasoning)
The ability to recognize analogies is fundamental to human cognition.
9 papers · 0 benchmarks
The Earning Calls dataset consists of processed earning conference calls data (text and audio).
9 papers · 0 benchmarks
Natural Language Inference (NLI), also called Textual Entailment, is an important task in NLP with the goal of determining the inference relationship between a premise p and a hypothesis h.
9 papers · 1 benchmark
FinRED is a relation extraction dataset curated from financial news and earning call transcripts containing relations from the finance domain.
9 papers · 0 benchmarks
GeoWebNews provides test/train examples and enable fine-grained Geotagging and Toponym Resolution (Geocoding).
9 papers · 0 benchmarks
Global Voices is a multilingual dataset for evaluating cross-lingual summarization methods.
9 papers · 0 benchmarks
HappyDB is a corpus of 100,000 crowdsourced happy moments.
9 papers · 0 benchmarks
Paper | Github | Dataset| Model As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e.
9 papers · 1 benchmark
Amazon Mechanical Turk (AMT) is used to collect annotations on HowTo100M videos.
9 papers · 0 benchmarks
The IndoSum dataset is a benchmark dataset for Indonesian text summarization.
9 papers · 0 benchmarks
KaMed is a knowledge-aware medical dialogue dataset, which contains over 60,000 medical dialogue sessions with 5,682 entities (such as Asthma and Atropine).
9 papers · 0 benchmarks
KnowIT VQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory.
9 papers · 0 benchmarks
LDC2020T02 (Abstract Meaning Representation (AMR) Annotation Release 3.0)
Abstract Meaning Representation (AMR) Annotation Release 3.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
9 papers · 1 benchmark
LEAF-QA, a comprehensive dataset of 250,000 densely annotated figures/charts, constructed from real-world open data sources, along with ~2 million question-answer (QA) pairs querying the structure and semantics of these charts.
9 papers · 0 benchmarks
MFRC (Moral Foundations Reddit Corpus)
Moral Foundations Reddit Corpus (MFRC) is a collection of 16,123 Reddit comments that have been curated from 12 distinct subreddits, hand-annotated by at least three trained annotators for 8 categories of moral sentiment (i.e., Care,…
9 papers · 0 benchmarks
MMCU (Measuring Massive Multitask Chinese Understanding)
We propose a test to measure the multitask accuracy of large Chinese language models.
9 papers · 0 benchmarks
MOD (Meme incorporated Open-domain Dialogue)
MOD is a large-scale open-domain multimodal dialogue dataset incorporating abundant Internet memes into utterances.
9 papers · 0 benchmarks
MVK (Marine Video Kit)
The dataset contains single-shot videos taken from moving cameras in underwater environments.
9 papers · 1 benchmark
MedNLI (Medical Natural Language Inference)
The MedNLI dataset consists of the sentence pairs developed by Physicians from the Past Medical History section of MIMIC-III clinical notes annotated for Definitely True, Maybe True and Definitely False.
9 papers · 2 benchmarks
MuSe-CaR (Multimodal Sentiment Analysis in Car Reviews)
The MuSe-CAR database is a large, multimodal (video, audio, and text) dataset which has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that…
9 papers · 0 benchmarks
NELA-GT-2018 is a dataset for the study of misinformation that consists of 713k articles collected between 02/2018-11/2018.
9 papers · 0 benchmarks
OntoGUM is an OntoNotes-like coreference dataset converted from GUM, an English corpus covering 12 genres using deterministic rules.
9 papers · 1 benchmark
RAVEN-FAIR is a modified version of the RAVEN dataset.
9 papers · 0 benchmarks
RadQA (A Question Answering Dataset to Improve Comprehension of Radiology Reports)
RadQA is a radiology question answering dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans (resulting in a total of 6148 question-answer evidence pairs) by physicians.
9 papers · 1 benchmark
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.
9 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.