Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 23 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1057–1104 of 3,130

The Medical Dataset for Abbreviation Disambiguation for Natural Language Understanding (MeDAL) is a large medical text dataset curated for abbreviation disambiguation, designed for natural language understanding pre-training in the medical…
7 papers · 0 benchmarks
The MedDialog dataset (Chinese) contains conversations (in Chinese) between doctors and patients.
7 papers · 0 benchmarks
In MutualFriends, two agents, A and B, each have a private knowledge base, which contains a list of friends with multiple attributes (e.g., name, school, major, etc.).
7 papers · 0 benchmarks
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
7 papers · 1 benchmark
New3, a set of 527 instances from AMR 3.0, whose original source was the LORELEI DARPA project – not included in the AMR 2.0 training set – consisting of excerpts from newswires and online forum.
7 papers · 1 benchmark
OPOSUM is a dataset for the training and evaluation of Opinion Summarization models which contains Amazon reviews from six product domains: Laptop Bags, Bluetooth Headsets, Boots, Keyboards, Televisions, and Vacuums.
7 papers · 0 benchmarks
PARus (Choice of Plausible Alternatives for Russian language)
Choice of Plausible Alternatives for Russian language (PARus) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning.
7 papers · 1 benchmark
PDNC (Project Dialogism Novel Corpus)
A annotated dataset of quotations and within-quotation-mentions in 22 full-length English novels.
7 papers · 0 benchmarks
PET (PET: A new Dataset for Process Extraction from Natural Language Text)
The dataset contains 45 documents containing narrative description of business process and their annotations.
7 papers · 0 benchmarks
PFN-PIC (PFN Picking Instructions for Commodities Dataset)
This dataset is a collection of spoken language instructions for a robotic system to pick and place common objects.
7 papers · 0 benchmarks
PHM2017 is a new dataset consisting of 7,192 English tweets across six diseases and conditions: Alzheimer’s Disease, heart attack (any severity), Parkinson’s disease, cancer (any type), Depression (any severity), and Stroke.
7 papers · 0 benchmarks
PhoMT is a high-quality and large-scale Vietnamese-English parallel dataset of 3.02M sentence pairs for machine translation.
7 papers · 1 benchmark
RWSD (The Winograd Schema Challenge (Russian))
A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its resolution.
7 papers · 1 benchmark
Reddit Corpus is part of a repository of conversational datasets consisting of hundreds of millions of examples, and a standardised evaluation procedure for conversational response selection models using '1-of-100 accuracy'.
7 papers · 0 benchmarks
A data set containing citations, citation contexts, and papers.
7 papers · 0 benchmarks
English subset of the SLAKE dataset, comprising 642 images and more than 7,000 question–answer pairs.
7 papers · 0 benchmarks
SME (Standard Multimodal Explanation)
SME is a new dataset for Multi-modal Explanation for Visual Question Answering comprising 1,028,230 samples, with 1,656 visual objects requiring detection in explanations.
7 papers · 1 benchmark
SNARE, short for ShapeNet Annotated with Referring Expressions, is a benchmark requires a model to choose which of two objects is being referenced by a natural language description.
7 papers · 0 benchmarks
A multimodal agent benchmark on professional data science and engineering.
7 papers · 0 benchmarks
The TACRED-Revisited dataset improves the crowd-sourced TACRED dataset for relation extraction by relabeling the dev and test sets using expert linguistic annotators.
7 papers · 1 benchmark
TERRa (Textual Entailment Recognition for Russian)
Textual Entailment Recognition has been proposed recently as a generic task that captures major semantic inference needs across many NLP applications, such as Question Answering, Information Retrieval, Information Extraction, and Text…
7 papers · 1 benchmark
TIAGE is a topic-shift aware dialog benchmark constructed utilizing human annotations on topic shifts.
7 papers · 0 benchmarks
UIT-ViCTSD (UIT Vietnamese Constructive and Toxic Speech Detection)
UIT-ViCTSD (Vietnamese Constructive and Toxic Speech Detection) is a dataset for constructive and toxic speech detection in Vietnamese.
7 papers · 0 benchmarks
UPAR (Unified Pedestrian Attribute Recognition)
The Task: The challenge will use an extension of the UPAR Dataset [1], which consists of images of pedestrians annotated for 40 binary attributes.
7 papers · 1 benchmark
This dataset was collected with the goal of assessing dialog evaluation metrics.
7 papers · 1 benchmark
This dataset was collected with the goal of assessing dialog evaluation metrics.
7 papers · 1 benchmark
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
VQA-CE (VQA Counterexamples)
This dataset provides a new split of VQA v2 (similarly to VQA-CP v2), which is built of questions that are hard to answer for biased models.
7 papers · 1 benchmark
ViHOS (Hate Speech Spans Detection for Vietnamese)
The first human-annotated corpus containing 26k spans on 11k comments
7 papers · 0 benchmarks
The WeChat dataset for fake news detection contains more than 20k news labelled as fake news or not.
7 papers · 1 benchmark
WikiDetox (Wikipedia Detox)
An annotated dataset of 1m crowd-sourced annotations that cover 100k talk page diffs (with 10 judgements per diff) for personal attacks, aggression, and toxicity.
7 papers · 0 benchmarks
This dataset is a new knowledge-base (KB) of hasPart relationships, extracted from a large corpus of generic statements.
7 papers · 0 benchmarks
k-qa (K-QA: A Real-World Medical Q&A Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
7 papers · 0 benchmarks
The AMI Meeting Corpus is a multi-modal data set comprising 100 hours of meeting recordings.
6 papers · 1 benchmark
The AQUAINT Corpus consists of newswire text data in English, drawn from three sources: the Xinhua News Service (People's Republic of China), the New York Times News Service, and the Associated Press Worldstream News Service.
6 papers · 1 benchmark
ArCOV-19 is an Arabic COVID-19 Twitter dataset that covers the period from 27th of January till 30th of April 2020.
6 papers · 0 benchmarks
ArmanEmo is a human-labeled emotion dataset of more than 7000 Persian sentences labeled for seven categories.
6 papers · 1 benchmark
BC4CHEMD (BioCreative IV Chemical compound and drug name recognition)
Introduced by Krallinger et al.
6 papers · 1 benchmark
BMELD is a bilingual (English-Chinese) dialogue corpus for Neural chat translation.
6 papers · 0 benchmarks
BSARD (Belgian Statutory Article Retrieval Dataset)
The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native corpus for studying statutory article retrieval.
6 papers · 1 benchmark
BiPaR is a manually annotated bilingual parallel novel-style machine reading comprehension (MRC) dataset, developed to support monolingual, multilingual and cross-lingual reading comprehension on novels.
6 papers · 0 benchmarks
CAIS (Chinese Artificial Intelligence Speakers)
We collect utterances from the Chinese Artificial Intelligence Speakers (CAIS), and annotate them with slot tags and intent labels.
6 papers · 2 benchmarks
CELLS is a large (63k pairs) and broadest-ranging (12 journals) parallel corpus for lay language generation.
6 papers · 0 benchmarks
CITE is a crowd-sourced resource for multimodal discourse: this resource characterises inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations.
6 papers · 1 benchmark
CLAMS (Cross-linguistic Analysis of Models on Syntax)
Targeted syntactic evaluation datasets in 5 languages: English, French, German, Russian, and Hebrew.
6 papers · 0 benchmarks
CLUECorpus2020 is a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation.
6 papers · 0 benchmarks
CoDa (The Color Dataset)
The Color Dataset (CoDa) is a probing dataset to evaluate the representation of visual properties in language models.
6 papers · 0 benchmarks
The CoarseWSD-20 dataset is a coarse-grained sense disambiguation dataset built from Wikipedia (nouns only) targeting 2 to 5 senses of 20 ambiguous words.
6 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.