Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 20 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 913–960 of 3,130
UGIF is a multi-lingual, multi-modal UI grounded dataset for step-by-step task completion on the smartphone.
10 papers · 0 benchmarks
ViP-Bench (Making Large Multimodal Models Understand Arbitrary Visual Prompts)
ViP-Bench is a comprehensive benchmark designed to assess the capability of multimodal models in understanding visual prompts across multiple dimensions.
10 papers · 1 benchmark
We propose a new, scalable video-mining pipeline which transfers captioning supervision from image datasets to video and audio.
10 papers · 0 benchmarks
WikiAtomicEdits is a corpus of 43 million atomic edits across 8 languages.
10 papers · 0 benchmarks
A corpus that encompasses the complete history of conversations between contributors to Wikipedia, one of the largest online collaborative communities.
10 papers · 0 benchmarks
XFORMAL is a multilingual formal style transfer benchmark of multiple formal reformulations of informal text in Brazilian Portuguese, French, and Italian.
10 papers · 0 benchmarks
e-ViL is a benchmark for explainable vision-language tasks.
10 papers · 0 benchmarks
The Sixth Informatics for Integrating Biology and the Bedside (i2b2) Natural Language Processing Challenge for Clinical Records focused on the temporal relations in clinical narratives.
9 papers · 2 benchmarks
Our dataset which consists of multiple indoor and outdoor experiments for up to 30 m gNB-UE link.
9 papers · 0 benchmarks
ART consists of over 20k commonsense narrative contexts and 200k explanations.
9 papers · 0 benchmarks
Is an acronym disambiguation (AD) dataset for scientific domain with 62,441 samples which is significantly larger than the previous scientific AD dataset.
9 papers · 0 benchmarks
This dataset contains 8.9M commonsense assertions extracted by the Ascent pipeline developed at the Max Planck Institute for Informatics.
9 papers · 0 benchmarks
BIMCV-COVID19+ dataset is a large dataset with chest X-ray images CXR (CR, DX) and computed tomography (CT) imaging of COVID-19 patients along with their radiographic findings, pathologies, polymerase chain reaction (PCR), immunoglobulin G…
9 papers · 0 benchmarks
Bactrian-X is a comprehensive multilingual parallel dataset of 3.4 million instruction-response pairs across 52 languages.
9 papers · 0 benchmarks
BiToD is a bilingual multi-domain dataset for end-to-end task-oriented dialogue modeling.
9 papers · 0 benchmarks
The Japanese-English business conversation corpus, namely Business Scene Dialogue corpus, was constructed in 3 steps: 1.
9 papers · 2 benchmarks
LRW-1000 has been renamed as CAS-VSR-W1k.
9 papers · 1 benchmark
CQASUMM is a dataset for CQA (Community Question Answering) summarization, constructed from the 4.4 million Yahoo!
9 papers · 0 benchmarks
CaSiNo is a dataset of 1030 negotiation dialogues in English.
9 papers · 0 benchmarks
ClueWeb22 is the newest iteration of the ClueWeb line of datasets, provides 10 billion web pages affiliated with rich information.
9 papers · 0 benchmarks
ComQA is a large dataset of real user questions that exhibit different challenging aspects such as compositionality, temporal reasoning, and comparisons.
9 papers · 0 benchmarks
Deep Learning Hard (DL-HARD) is an annotated dataset designed to more effectively evaluate neural ranking models on complex topics.
9 papers · 0 benchmarks
E-KAR (Benchmark for Explainable Knowledge-intensive Analogical Reasoning)
The ability to recognize analogies is fundamental to human cognition.
9 papers · 0 benchmarks
The Earning Calls dataset consists of processed earning conference calls data (text and audio).
9 papers · 0 benchmarks
EgoProceL is a large-scale dataset for procedure learning.
9 papers · 0 benchmarks
Natural Language Inference (NLI), also called Textual Entailment, is an important task in NLP with the goal of determining the inference relationship between a premise p and a hypothesis h.
9 papers · 1 benchmark
FinRED is a relation extraction dataset curated from financial news and earning call transcripts containing relations from the finance domain.
9 papers · 0 benchmarks
GeoWebNews provides test/train examples and enable fine-grained Geotagging and Toponym Resolution (Geocoding).
9 papers · 0 benchmarks
Global Voices is a multilingual dataset for evaluating cross-lingual summarization methods.
9 papers · 0 benchmarks
HappyDB is a corpus of 100,000 crowdsourced happy moments.
9 papers · 0 benchmarks
Paper | Github | Dataset| Model As a part of our research efforts toward making LLMs more safe for public use, we create HarmfulQA i.e.
9 papers · 1 benchmark
Amazon Mechanical Turk (AMT) is used to collect annotations on HowTo100M videos.
9 papers · 0 benchmarks
The IndoSum dataset is a benchmark dataset for Indonesian text summarization.
9 papers · 0 benchmarks
KaMed is a knowledge-aware medical dialogue dataset, which contains over 60,000 medical dialogue sessions with 5,682 entities (such as Asthma and Atropine).
9 papers · 0 benchmarks
KnowIT VQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory.
9 papers · 0 benchmarks
LDC2020T02 (Abstract Meaning Representation (AMR) Annotation Release 3.0)
Abstract Meaning Representation (AMR) Annotation Release 3.0 was developed by the Linguistic Data Consortium (LDC), SDL/Language Weaver, Inc., the University of Colorado's Computational Language and Educational Research group and the…
9 papers · 1 benchmark
LEAF-QA, a comprehensive dataset of 250,000 densely annotated figures/charts, constructed from real-world open data sources, along with ~2 million question-answer (QA) pairs querying the structure and semantics of these charts.
9 papers · 0 benchmarks
MFRC (Moral Foundations Reddit Corpus)
Moral Foundations Reddit Corpus (MFRC) is a collection of 16,123 Reddit comments that have been curated from 12 distinct subreddits, hand-annotated by at least three trained annotators for 8 categories of moral sentiment (i.e., Care,…
9 papers · 0 benchmarks
MMCU (Measuring Massive Multitask Chinese Understanding)
We propose a test to measure the multitask accuracy of large Chinese language models.
9 papers · 0 benchmarks
MOD (Meme incorporated Open-domain Dialogue)
MOD is a large-scale open-domain multimodal dialogue dataset incorporating abundant Internet memes into utterances.
9 papers · 0 benchmarks
The dataset contains single-shot videos taken from moving cameras in underwater environments.
9 papers · 1 benchmark
MedNLI (Medical Natural Language Inference)
The MedNLI dataset consists of the sentence pairs developed by Physicians from the Past Medical History section of MIMIC-III clinical notes annotated for Definitely True, Maybe True and Definitely False.
9 papers · 2 benchmarks
MuSe-CaR (Multimodal Sentiment Analysis in Car Reviews)
The MuSe-CAR database is a large, multimodal (video, audio, and text) dataset which has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that…
9 papers · 0 benchmarks
NELA-GT-2018 is a dataset for the study of misinformation that consists of 713k articles collected between 02/2018-11/2018.
9 papers · 0 benchmarks
OntoGUM is an OntoNotes-like coreference dataset converted from GUM, an English corpus covering 12 genres using deterministic rules.
9 papers · 1 benchmark
RAVEN-FAIR is a modified version of the RAVEN dataset.
9 papers · 0 benchmarks
RadQA (A Question Answering Dataset to Improve Comprehension of Radiology Reports)
RadQA is a radiology question answering dataset with 3074 questions posed against radiology reports and annotated with their corresponding answer spans (resulting in a total of 6148 question-answer evidence pairs) by physicians.
9 papers · 1 benchmark
Speech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.
9 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.