Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 26 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1201–1248 of 3,130

Coached Conversational Preference Elicitation is a dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
5 papers · 0 benchmarks
ComFact is a benchmark for commonsense fact linking, where models are given contexts and trained to identify situationally-relevant commonsense knowledge from KGs.
5 papers · 0 benchmarks
CrossRE is a cross-domain benchmark for Relation Extraction (RE), which comprises six distinct text domains and includes multi-label annotations.
5 papers · 0 benchmarks
A corpus of Offensive Language and Hate Speech Detection for Danish This DKhate dataset contains 3600 comments from the web annotated for offensive language, following the Zampieri et al.
5 papers · 1 benchmark
DaN+ is a new multi-domain corpus and annotation guidelines for Danish nested named entities (NEs) and lexical normalization to support research on cross-lingual cross-domain learning for a less-resourced language.
5 papers · 0 benchmarks
The Distress Analysis Interview Corpus/Wizard-of-Oz set (DAIC-WOZ) dataset [24, 25] comprises voice and text samples from 189 interviewed healthy and control persons and their PHQ-8 depression detection questionnaire.
5 papers · 0 benchmarks
The EDT dataset is designed for corporate event detection and text-based stock prediction (trading strategy) benchmark.
5 papers · 0 benchmarks
FLD (Formal Logic Deduction)
A deductive reasoning benchmark based on formal logic theory.
5 papers · 0 benchmarks
GAD (Gene Associations Database)
GAD, or Gene Associations Database, is a corpus of gene-disease associations curated from genetic association studies.
5 papers · 1 benchmark
Repair AST parse (syntax) errors in Python code
5 papers · 1 benchmark
Grep-BiasIR (Gender Representation-Bias for Information Retrieval)
Grep-BiasIR is a novel thoroughly-audited dataset which aim to facilitate the studies of gender bias in the retrieved results of IR systems.
5 papers · 0 benchmarks
HJDataset is a large dataset of Historical Japanese Documents with Complex Layouts.
5 papers · 0 benchmarks
HaDes is a token-level, reference-free hallucination detection dataset named HAllucination DEtection dataSet.
5 papers · 0 benchmarks
The Headlines dataset for sarcasm detection is collected from two news website.
5 papers · 0 benchmarks
We introduce HourVideo, a benchmark dataset for hour-long video-language understanding.
5 papers · 0 benchmarks
HowMany-Qa is a object counting dataset.
5 papers · 1 benchmark
A dataset of 69,270,581 video clip, question and answer triplets (v, q, a).
5 papers · 0 benchmarks
Hummingbird is a dataset to examine stylistic lexical cues from human perception and BERT used to characterize their discrepancy.
5 papers · 0 benchmarks
IAM(line-level) (Line-level Handwritten Text Recognition on IAM)
The IAM database contains 13,353 images of handwritten lines of text created by 657 writers.
5 papers · 1 benchmark
IG-1B-Targeted is an internal Facebook AI Research dataset that consists of 940 million public images with 1.5K hashtags matching with 1000 ImageNet1K synsets.
5 papers · 0 benchmarks
IGLU is a dataset designed for interactive grounded language understanding.
5 papers · 0 benchmarks
Illness-dataset (Illness multi-domain textual dataset)
A dataset for evaluating text classification, domain adaptation, and active learning models.
5 papers · 0 benchmarks
The ImplicitQA dataset was introduced in the paper ImplicitQA: Going beyond frames towards Implicit Video Reasoning.
5 papers · 1 benchmark
We collect, organize and open-source the large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples.
5 papers · 0 benchmarks
Instructional-DT (Instr-DT) (Instructional Discourse Treebank)
This discourse treebank includes annotated instructional texts originally assembled at the Information Technology Research Institute, University of Brighton.
5 papers · 1 benchmark
ItaCoLA is a corpus for monolingual and cross-lingual acceptability judgments which contains almost 10,000 sentences with acceptability judgments.
5 papers · 1 benchmark
Dataset for lyrics alignment and transcription evaluation.
5 papers · 0 benchmarks
6.5 million anonymous ratings of jokes by users of the Jester Joke Recommender System.
5 papers · 0 benchmarks
KETOD (Knowledge-Enriched Task-Oriented Dialogue)
KETOD (Knowledge-Enriched Task-Oriented Dialogue) is a dataset containing system responses designed for enriching task-oriented dialogues with chit-chat based on relevant entity knowledge.
5 papers · 0 benchmarks
Klexikon (Klexikon: A German Dataset for Joint Summarization and Simplification)
The dataset introduces document alignments between German Wikipedia and the children's lexicon Klexikon.
5 papers · 1 benchmark
LINNAEUS is a general-purpose dictionary matching software, capable of processing multiple types of document formats in the biomedical domain (MEDLINE, PMC, BMC, OTMI, text, etc.).
5 papers · 1 benchmark
LiDiRus (Linguistic Diagnostic for Russian)
LiDiRus is a diagnostic dataset that covers a large volume of linguistic phenomena, while allowing you to evaluate information systems on a simple test of textual entailment recognition.
5 papers · 1 benchmark
We propose a novel long-context benchmark, 🐉 Loong, aligning with realistic scenarios through extended multi-document question answering (QA).
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
MATINF (Maternal and Infant Dataset)
Maternal and Infant (MATINF) Dataset is a large-scale dataset jointly labeled for classification, question answering and summarization in the domain of maternity and baby caring in Chinese.
5 papers · 0 benchmarks
MLQE (MultiLingual Quality Estimation)
The MLQE dataset is a dataset for sentence-level Machine Translation Quality Estimation.
5 papers · 0 benchmarks
- A large scale Chinese multi-modal dialogue corpus (120.84K dialogues and 198.82 K images).
5 papers · 0 benchmarks
MatSynth MatSynth is a Physically Based Rendering (PBR) materials dataset designed for modern AI applications.
5 papers · 0 benchmarks
A large, realistic multimodal dataset consisting of real personal photos and crowd-sourced questions/answers.
5 papers · 1 benchmark
MindCraft is a fine-grained dataset of collaborative tasks performed by pairs of human subjects in the 3D virtual blocks world of Minecraft.
5 papers · 0 benchmarks
MoA (MoA_Long_ModelQA)
This is the dataset used by the automatic sparse attention compression method MoA.
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
NELA-GT-2019 is an updated version of the NELA-GT-2018 dataset.
5 papers · 0 benchmarks
NFCorpus is a full-text English retrieval data set for Medical Information Retrieval.
5 papers · 1 benchmark
NLPeer is a multidomain corpus of more than 5k papers and 11k review reports from five different venues.
5 papers · 0 benchmarks
Given two entities, generating a coherent sentence describing the relation between them.
5 papers · 0 benchmarks
Open6DOR V2 (Benchmarking Open-instruction 6-DoF Object Rearrangement and A VLM-based Approach)
We introduce a challenging and comprehensive benchmark for open-instruction 6-DoF object rearrangement tasks, termed Open6DOR.
5 papers · 1 benchmark
PQuAD (Persian Question Answering Dataset)
Persian Question Answering Dataset (PQuAD) is a crowdsourced reading comprehension dataset on Persian Wikipedia articles.
5 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.