Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 32 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1489–1536 of 3,130
The high-quality multi-turn dialogue dataset, which has a total of 3,134 multi-turn consultation dialogues.
3 papers · 0 benchmarks
This dataset is constructed and based on the online free-access fictions that are tagged with sci-fi, urban novel, love story, youth, etc.
3 papers · 0 benchmarks
ClarQ, consists of ∼2M examples distributed across 173 domains of stackexchange.
3 papers · 0 benchmarks
CoScript is a constrained language planning dataset, which consists of 55,000 scripts.
3 papers · 0 benchmarks
CoVaxLies v1 includes 17 known Misinformation Targets (MisTs) found on Twitter about the covid-19 vaccines.
3 papers · 0 benchmarks
CoWeSe (Corpus Web Salud Espanol)
CoWeSe is a Spanish biomedical corpus consisting of 4.5GB (about 750M tokens) of clean plain text.
3 papers · 0 benchmarks
CommitBART is a benchmark for researching commit-related task such as denoising, cross-modal generation and contrastive learning.
3 papers · 0 benchmarks
Commonsense-Dialogues is a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense.
3 papers · 0 benchmarks
ConcurrentQA is a textual multi-hop QA benchmark to require concurrent retrieval over multiple data-distributions (i.e.
3 papers · 1 benchmark
Conversational Stance Detection (CSD) is a dataset with annotations of stances and the structures of conversation threads.
3 papers · 0 benchmarks
The 2021 SIGIR workshop on eCommerce is hosting the Coveo Data Challenge for "In-session prediction for purchase intent and recommendations".
3 papers · 1 benchmark
DAWT (Densely Annotated Wikipedia Texts)
The DAWT dataset consists of Densely Annotated Wikipedia Texts across multiple languages.
3 papers · 0 benchmarks
The dataset contains the main components of the news articles published online by the newspaper named Gazzetta di Modena: url of the web page, title, sub-title, text, date of publication, crime category assigned to each news article by the…
3 papers · 0 benchmarks
DISRPT2021 (DISRPT2021 shared task on Discourse Unit Segmentation, Connective Detection and Discourse Relation Classification)
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism…
3 papers · 0 benchmarks
The DeepMind Q&A Dataset consists of two datasets for Question Answering, CNN and DailyMail.
3 papers · 0 benchmarks
DR.BENCH (Diagnostic Reasoning Benchmark for clinical natural language processing)
DR.BENCH is a dataset for developing and evaluating cNLP models with clinical diagnostic reasoning ability.
3 papers · 0 benchmarks
We present a dataset, DANFEVER, intended for claim verification in Danish.
3 papers · 1 benchmark
The dermatology differential diagnoses (ddx) dataset for skin condition classification includes expert annotations and model predictions for 1947 cases.
3 papers · 0 benchmarks
DiFair serves as a meticulous endeavor to address the oversight in evaluating the impact of bias mitigation on useful gender knowledge while assessing gender neutrality in pretrained language models.
3 papers · 0 benchmarks
Disfl-QA is a targeted dataset for contextual disfluencies in an information seeking setting, namely question answering over Wikipedia passages.
3 papers · 0 benchmarks
DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
This repository contains gzipped files containing more than 2 million tokens (words) from answers submitted by more than 6,000 students over the course of their first 30 days of using Duolingo.
3 papers · 0 benchmarks
This is a gzipped CSV file containing the 13 million Duolingo student learning traces used in experiments by Settles & Meeder (2016).
3 papers · 0 benchmarks
EC-FUNSD is introduced in [[arXiv:2402.02379]](https://arxiv.org/abs/2402.02379) as a benchmark of semantic entity recognition (SER) and entity linking (EL), designed for the entity-centric robustness evaluation of pre-trained…
3 papers · 2 benchmarks
EMOTyDA (Emotion aware Dialogue Act)
EMOTyDA is a multimodal Emotion aware Dialogue Act dataset collected from open-sourced dialogue datasets.
3 papers · 1 benchmark
The EXEQ-300k dataset contains 290,479 detailed questions with corresponding math headlines from Mathematics Stack Exchange.
3 papers · 0 benchmarks
Ego4D-HCap is a hierarchical video captioning dataset comprised of a three-tier hierarchy of captions: short clip-level captions, medium-length video segment descriptions, and long-range video-level summaries.
3 papers · 0 benchmarks
Email Thread Summarization (EmailSum) is a dataset which contains human-annotated short (<30 words) and long (<100 words) summaries of 2,549 email threads (each containing 3 to 10 emails) over a wide variety of topics.
3 papers · 2 benchmarks
Emotional Dialogue Acts data contains dialogue act labels for existing emotion multi-modal conversational datasets.
3 papers · 0 benchmarks
ErAConD (Error Annotated Conversational Dialog Dataset for Grammatical Error Correction)
ErAConD is a novel GEC dataset consisting of parallel original and corrected utterances drawn from open-domain chatbot conversations.
3 papers · 0 benchmarks
FFHQ-Text is a small-scale face image dataset with large-scale facial attributes, designed for text-to-face generation & manipulation, text-guided facial image manipulation, and other vision-related tasks.
3 papers · 0 benchmarks
FOBIE (Focused Open Biological Information Extraction)
The Focused Open Biology Information Extraction (FOBIE) dataset aims to support IE from Computer-Aided Biomimetics.
3 papers · 0 benchmarks
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
FakeNewsAMT & Celebrity include two novel datasets for the task of fake news detection, covering seven different news domains.
3 papers · 0 benchmarks
This is a dataset for segmentation and classification of epistemic activities in diagnostic reasoning texts.
3 papers · 0 benchmarks
FixMyPose is a dataset for automated pose correction.
3 papers · 0 benchmarks
FloDial (Flowchart Grounded Dialogs Dataset)
Flowchart Grounded Dialog Dataset (FloDial) is a corpus of troubleshooting dialogs between a user and an agent collected using Amazon Mechanical Turk.
3 papers · 0 benchmarks
FreSaDa is a French satire dataset for cross-domain satire detection, which is composed of 11,570 articles from the news domain.
3 papers · 0 benchmarks
FunQA is a challenging video question answering (QA) dataset specifically designed to evaluate and enhance the depth of video reasoning based on counter-intuitive and fun videos.
3 papers · 0 benchmarks
A high-quality dataset for machine translation evaluation that aims at being one of the first non-synthetic gender-balanced test datasets.
3 papers · 0 benchmarks
GeoGLUE (GeoGraphic Language Understanding Evaluation Benchmark)
GeoGLUE is a GeoGraphic Language Understanding Evaluation benchmark, which consists of six geographic text-related tasks, including geographic textual similarity on recall, geotagged geographic elements tagging, geographic composition…
3 papers · 0 benchmarks
Geoclidean-Constraints dataset consists of 20 concepts and 40 tasks, created from permutations of line and circle construction rules with various constraints describing the relationship between objects.
3 papers · 0 benchmarks
Are you the kind of person who makes a lot of typos when writing code?
3 papers · 0 benchmarks
GlotScript-R is a resource that provides the attested writing systems for more than 7,000 languages.
3 papers · 0 benchmarks
GroOT (Grounded Multiple Object Tracking)
One of the recent trends in vision problems is to use natural language captions to describe the objects of interest.
3 papers · 0 benchmarks
Multilingual text collection extracted from the Internet Archive and Common Crawl archives.
3 papers · 0 benchmarks
This dataset is built from Twitter and contains 1290 hate tweet and counterspeech reply pairs.
3 papers · 0 benchmarks
HateScore (HateScore : Human-in-the-Loop and Neutral Korean Multi-label Online Hate Speech Dataset)
2.2K neutral sentences from Wikipedia 1.7K additionally labeled sentences generated by the Human-in-the-Loop procedure (based on Korean Unsmile Dataset Base Model) 7.1K rule-generated neutral sentences
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.