Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 24 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1105–1152 of 3,130
ConvoSumm is a suite of four datasets to evaluate a model’s performance on a broad spectrum of conversation data.
6 papers · 0 benchmarks
DaLAJ 1.0, a dataset for Linguistic Acceptability Judgments for Swedish, comprising 9,596 sentences in its first version; and the initial experiment using it for the binary classification task.
6 papers · 1 benchmark
DaNE (Danish Dependency Treebank)
Danish Dependency Treebank (DaNE) is a named entity annotation for the Danish Universal Dependencies treebank using the CoNLL-2003 annotation scheme.
6 papers · 5 benchmarks
DebateSum consists of 187328 debate documents, arguments (also can be thought of as abstractive summaries, or queries), word-level extractive summaries, citations, and associated metadata organized by topic-year.
6 papers · 1 benchmark
Cant (also known as doublespeak, cryptolect, argot, anti-language or secret language) is important for understanding advertising, comedies and dog-whistle politics.
6 papers · 0 benchmarks
EHR-RelB is a benchmark dataset for biomedical concept relatedness, consisting of 3630 concept pairs sampled from electronic health records (EHRs).
6 papers · 0 benchmarks
EarthVQA (A multi-modal multi-task VQA dataset for remote sensing)
Earth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning.
6 papers · 1 benchmark
This dataset contains around 10000 videos generated by various methods using the Prompt list.
6 papers · 1 benchmark
French TimeBank, a corpus for French annotated in ISO-TimeML.
6 papers · 1 benchmark
French Wikipedia is a dataset used for pretraining the CamemBERT French language model.
6 papers · 0 benchmarks
FrenchMedMCQA (FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain)
This paper introduces FrenchMedMCQA, the first publicly available Multiple-Choice Question Answering (MCQA) dataset in French for medical domain.
6 papers · 1 benchmark
FusedChat is an inter-mode dialogue dataset.
6 papers · 1 benchmark
Geo-Diverse Visual Commonsense Reasoning (GD-VCR) is a new dataset to test vision-and-language models' ability to understand cultural and geo-location-specific commonsense.
6 papers · 1 benchmark
This dataset is a new benchmark, grounded in real-world usages is developed to support more authentic and comprehensive evaluation of image editing models.
6 papers · 2 benchmarks
GICoref (Gender Inclusive Coreference)
GICoref is a fully annotated coreference resolution dataset written by and about trans people.
6 papers · 0 benchmarks
The German Lipreading dataset consists of 250,000 publicly available videos of the faces of speakers of the Hessian Parliament, which was processed for word-level lip reading using an automatic pipeline.
6 papers · 0 benchmarks
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
HRS-Bench (Holistic, Reliable, and Scalable Benchmark)
HRS-Bench is a concrete evaluation benchmark for T2I models that is Holistic, Reliable, and Scalable.
6 papers · 0 benchmarks
We introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency.
6 papers · 0 benchmarks
HurricaneEmo is an emotion dataset that contains 15,000 English tweets spanning three hurricanes: Harvey, Irma, and Maria.
6 papers · 0 benchmarks
The IBM-Rank-30k is a dataset for the task of argument quality ranking.
6 papers · 0 benchmarks
IIIT-ILST is a dataset and benchmark for scene text recognition for three Indic scripts - Devanagari, Telugu and Malayalam.
6 papers · 0 benchmarks
JParaCrawl is a parallel corpus for English-Japanese, for which the amount of publicly available parallel corpora is still limited.
6 papers · 0 benchmarks
JerichoWorld is a dataset that enables the creation of learning agents that can build knowledge graph-based world models of interactive narratives.
6 papers · 2 benchmarks
KAMEL (Knowledge Analysis with Multitoken Entities in Language Models)
KAMEL comprises knowledge about 234 relations from Wikidata with a large training, validation, and test dataset.
6 papers · 1 benchmark
KnowledgeNet is a benchmark dataset for the task of automatically populating a knowledge base (Wikidata) with facts expressed in natural language text on the web.
6 papers · 0 benchmarks
Kompetencer (Danish Job Postings Classification Dataset)
Kompetencer (en: competences) is a Danish job posting dataset annotated for nested spans of competences.
6 papers · 0 benchmarks
LCQMC (Large-scale Chinese Question Matching Corpus)
LCQMC is a large-scale Chinese question matching corpus.
6 papers · 0 benchmarks
LSSED, a challenging large-scale english dataset for speech emotion recognition.
6 papers · 1 benchmark
Lyra is a dataset for code generation that consists on Python code with embedded SQL.
6 papers · 0 benchmarks
The MEDIA French corpus is dedicated to semantic extraction from speech in a context of human/machine dialogues.
6 papers · 0 benchmarks
MOLD (Marathi Offensive Language Dataset)
MOLD is a Marathi dataset for offensive language identification
6 papers · 0 benchmarks
The MidiCaps dataset [1] is a large-scale dataset of 168,385 midi music files with descriptive text captions, and a set of extracted musical features.
6 papers · 0 benchmarks
A quantitative benchmark for developing and understanding video of fill-in-the-blank question-answering dataset with over 300,000 examples, based on descriptive video annotations for the visually impaired.
6 papers · 0 benchmarks
Moviescope is a large-scale dataset of 5,000 movies with corresponding video trailers, posters, plots and metadata.
6 papers · 0 benchmarks
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio.
6 papers · 0 benchmarks
MuSeRC (Russian Multi-Sentence Reading Comprehension)
We present a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences.
6 papers · 1 benchmark
The MusicBench dataset is a music audio-text pair dataset that was designed for text-to-music generation purpose and released along with Mustango text-to-music model.
6 papers · 1 benchmark
NLU++ (NLLU++ : A Multi-Label, Slot-Rich, Generalisable Dataset for Natural Language Understanding in Task-Oriented Dialogue)
nlu++ is a dataset for natural language understanding (NLU) in task-oriented dialogue (ToD) systems, with the aim to provide a much more challenging evaluation environment for dialogue NLU models, up to date with the current application…
6 papers · 0 benchmarks
There are two versions of the NLmaps corpus.
6 papers · 0 benchmarks
Naamapadam is a Named Entity Recognition (NER) dataset for the 11 major Indian languages from two language families.
6 papers · 0 benchmarks
OMICS (Open Mind Indoor Common Sense)
OMICS is an extensive collection of knowledge for indoor service robots gathered from internet users.
6 papers · 0 benchmarks
OTTers is a dataset of human one-turn topic transitions.
6 papers · 0 benchmarks
Open Images is a computer vision dataset covering ~9 million images with labels spanning thousands of object categories.
6 papers · 0 benchmarks
PLABA (Plain Language Adaptation of Biomedical Abstracts)
Plain Language Adaptation of Biomedical Abstracts (PLABA) is a dataset designed for automatic adaptation that is both document- and sentence-aligned.
6 papers · 0 benchmarks
QA-SRL Bank 2.0 is a large-scale corpus of Question-Answer driven Semantic Role Labeling (QA-SRL) annotations.
6 papers · 0 benchmarks
RCB (Russian Commitment Bank)
The Russian Commitment Bank is a corpus of naturally occurring discourses whose final sentence contains a clause-embedding predicate under an entailment cancelling operator (question, modal, negation, antecedent of conditional).
6 papers · 1 benchmark
This dataset arises from the READ project (Horizon 2020).
6 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.