Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 14 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 625–672 of 3,130
Paralex learns from a collection of 18 million question-paraphrase pairs scraped from WikiAnswers.
20 papers · 1 benchmark
PhotoChat, the first dataset that casts light on the photo sharing behavior in online messaging.
20 papers · 2 benchmarks
A large-scale English dataset for coreference resolution.
20 papers · 1 benchmark
PsyQA is a Chinese Dataset for generating long counseling text for mental health support.
20 papers · 0 benchmarks
PubMed 200k RCT is new dataset based on PubMed for sequential sentence classification.
20 papers · 0 benchmarks
QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms.
20 papers · 0 benchmarks
RCTW-17 (Reading Chinese Text in the Wild)
Features a large-scale dataset with 12,263 annotated images.
20 papers · 0 benchmarks
RoboCup is an initiative in which research groups compete by enabling their robots to play football matches.
20 papers · 0 benchmarks
SUBJ (Subjectivity dataset)
Available are collections of movie-review documents labeled with respect to their overall sentiment polarity (positive or negative) or subjective rating (e.g., "two and a half stars") and sentences labeled with respect to their…
20 papers · 1 benchmark
Large-scale manually-annotated corpus for 1,000 scientific papers (on computational linguistics) for automatic summarization.
20 papers · 0 benchmarks
The purpose of this dataset was to study gender bias in occupations.
19 papers · 1 benchmark
Ten years (2008-2018) ChFinAnn documents and human-summarized event knowledge bases to conduct the DS-based event labeling.
19 papers · 1 benchmark
CoAuthor is a dataset designed for revealing GPT-3's capabilities in assisting creative and argumentative writing.
19 papers · 1 benchmark
DDXPlus (DDXPlus: A New Dataset For Automatic Medical Diagnosis)
There has been a rapidly growing interest in Automatic Symptom Detection (ASD) and Automatic Diagnosis (AD) systems in the machine learning research literature, aiming to assist doctors in telemedicine services.
19 papers · 0 benchmarks
DIRHA (Distant-speech Interaction for Robust Home Applications)
DIRHA-English is a multi-microphone database composed of real and simulated sequences of 1-minute.
19 papers · 1 benchmark
FNC-1 (Fake News Challenge Stage 1)
FNC-1 was designed as a stance detection dataset and it contains 75,385 labeled headline and article pairs.
19 papers · 2 benchmarks
FoCus (Call for Customized Conversation: Customized Conversation Grounding Persona and Knowledge)
We introduce a new dataset, called FoCus, that supports knowledge-grounded answers that reflect user’s persona.
19 papers · 0 benchmarks
Funcom is a collection of ~2.1 million Java methods and their associated Javadoc comments.
19 papers · 0 benchmarks
This dataset code generates mathematical question and answer pairs, from a range of question types at roughly school-level difficulty.
19 papers · 1 benchmark
OVEN (Open-domain Visual Entity Recognition)
In this project, we formally present the task of Open-domain Visual Entity recognitioN (OVEN), where a model need to link an image onto a Wikipedia entity with respect to a text query.
19 papers · 1 benchmark
QAMPARI is an ODQA benchmark, where question answers are lists of entities, spread across many paragraphs.
19 papers · 0 benchmarks
A dataset on asking Questions for Lack of Clarity in open-domain information-seeking conversations.
19 papers · 0 benchmarks
SUTD-TrafficQA (Singapore University of Technology and Design - Traffic Question Answering) is a dataset which takes the form of video QA based on 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive…
19 papers · 1 benchmark
There are now many computer programs for automatically determining the sense of a word in context (Word Sense Disambiguation or WSD).
19 papers · 0 benchmarks
Taskmaster-1 is a dialog dataset consisting of 13,215 task-based dialogs in English, including 5,507 spoken and 7,708 written dialogs created with two distinct procedures.
19 papers · 0 benchmarks
Torque is an English reading comprehension benchmark built on 3.2k news snippets with 21k human-generated questions querying temporal relationships.
19 papers · 1 benchmark
Touchdown is a corpus for executing navigation instructions and resolving spatial descriptions in visual real-world environments.
19 papers · 1 benchmark
With social media becoming increasingly popular on which lots of news and real-time events are reported, developing automated question answering systems is critical to the effectiveness of many applications that rely on real-time knowledge.
19 papers · 1 benchmark
UMLS (Unified Medical Language System)
The Unified Medical Language System (UMLS) is a comprehensive resource that integrates and disseminates essential terminology, classification standards, and coding systems.
19 papers · 1 benchmark
WanJuan is a large-scale training corpus that includes multiple modalities.
19 papers · 0 benchmarks
2010 i2b2/VA is a biomedical dataset for relation classification and entity typing.
18 papers · 4 benchmarks
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
18 papers · 3 benchmarks
We collect a new dataset of human-posed free-form natural language questions about CLEVR images.
18 papers · 1 benchmark
CMB (Comprehensive Medical Benchmark in Chinese)
CMB is a comprehensive, multi-level Medical Benchmark in Chinese.
18 papers · 0 benchmarks
DWIE (Deutsche Welle corpus for Information Extraction)
The 'Deutsche Welle corpus for Information Extraction' (DWIE) is a multi-task dataset that combines four main Information Extraction (IE) annotation sub-tasks: (i) Named Entity Recognition (NER), (ii) Coreference Resolution, (iii) Relation…
18 papers · 5 benchmarks
The objective in extreme multi-label classification is to learn feature architectures and classifiers that can automatically tag a data point with the most relevant subset of labels from an extremely large label set.
18 papers · 0 benchmarks
This dataset was collected and prepared by the CALO Project (A Cognitive Assistant that Learns and Organizes).
18 papers · 1 benchmark
ICDAR2017 is a dataset for scene text detection.
18 papers · 1 benchmark
InfoTabS comprises of human-written textual hypotheses based on premises that are tables extracted from Wikipedia info-boxes.
18 papers · 0 benchmarks
KorNLI is a Korean Natural Language Inference (NLI) dataset.
18 papers · 0 benchmarks
MED (Monotonicity Entailment Dataset)
MED is a new evaluation dataset that covers a wide range of monotonicity reasoning that was created by crowdsourcing and collected from linguistics publications.
18 papers · 1 benchmark
The MMD (MultiModal Dialogs) dataset is a dataset for multimodal domain-aware conversations.
18 papers · 0 benchmarks
SALAD-Bench (A Hierarchical and Comprehensive Safety Benchmark for Large Language Models)
In the rapidly evolving landscape of Large Language Models (LLMs), ensuring robust safety measures is paramount.
18 papers · 0 benchmarks
SCICAP is a large-scale image captioning dataset that contains real-world scientific figures and captions.
18 papers · 1 benchmark
SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically.
18 papers · 0 benchmarks
TuringBench is a benchmark environment that contains : - Benchmark tasks- Turing Test (i.e., human vs.
18 papers · 2 benchmarks
TempEval-3 (TempEval-3: events, times, and temporal relations)
Within the SemEval-2013 evaluation exercise, the TempEval-3 shared task aims to advance research on temporal information processing.
18 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.