Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 37 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1729–1776 of 3,130

CoVaxLies v2 includes 47 Misinformation Targets (MisTs) found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
CodeQueries Benchmark dataset consists of instances of semantic queries, code context and code spans in the context corresponding to the semantic queries.
2 papers · 0 benchmarks
CodeSyntax is a large-scale dataset of programs annotated with the syntactic relationships in their corresponding abstract syntax trees.
2 papers · 0 benchmarks
The Common Crawl corpus contains petabytes of data collected over 12 years of web crawling.
2 papers · 0 benchmarks
Comparative Question Completion is a dataset to evaluate what do large Language Models learn.
2 papers · 0 benchmarks
Concise has two datasets of 2000 sentences each, that were annotated by two and five human annotators, respectively.
2 papers · 0 benchmarks
Covid-HeRA is a dataset for health risk assessment and severity-informed decision making in the presence of COVID19 misinformation.
2 papers · 0 benchmarks
Given an English article, generate a short summary in the target language.
2 papers · 0 benchmarks
The Curiosity dataset consists of 14K dialogs (with 181K utterances) with fine-grained knowledge groundings, dialog act annotations, and other auxiliary annotation.
2 papers · 0 benchmarks
Czech restaurant information is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language.
2 papers · 1 benchmark
DACCORD is a new dataset dedicated to the task of automatically detecting contradictions between sentences in French.
2 papers · 0 benchmarks
DAST (Danish Stance)
This is an SDQC stance-annotated Reddit dataset for the Danish language generated within a thesis project.
2 papers · 0 benchmarks
DEplain-APA-sent: A German Parallel Corpus for Sentence Simplification on News Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
2 papers · 1 benchmark
DEplain-web-sent: A German Parallel Corpus for Sentence Simplification on Web Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
2 papers · 1 benchmark
DIALOCONAN is a dataset comprising over 3000 fictitious multi-turn dialogues between a hater and an NGO operator, covering 6 targets of hate.
2 papers · 0 benchmarks
DISC-Law-SFT comprises two subsets, DISC-Law-SFT-Pair and DISC-Law-SFT-Triplet.
2 papers · 0 benchmarks
DISL (Fueling Research with A Large Dataset of Solidity Smart Contracts)
DISL The full dataset report is available at: https://arxiv.org/abs/2403.16861 The DISL dataset features a collection of 514, 506 unique Solidity files that have been deployed to Ethereum mainnet.
2 papers · 0 benchmarks
DRI Corpus (Dr. Inventor Multi-layer Scientific Corpus)
The Dr.
2 papers · 2 benchmarks
DTGB (Dynamic Text-attributed Graph Benchmark)
We introduce Dynamic Text-attributed Graph Benchmark (DTGB), a collection of large-scale, time-evolving graphs from diverse domains, with nodes and edges enriched by dynamically changing text attributes and categories.
2 papers · 0 benchmarks
DUC 2007 (Document Understanding Conferences)
There is currently much interest and activity aimed at building powerful multi-purpose information systems.
2 papers · 0 benchmarks
Dataset of Legal Documents consists of court decisions from 2017 and 2018 were selected for the dataset, published online by the Federal Ministry of Justice and Consumer Protection.
2 papers · 0 benchmarks
DeCOCO is a bilingual (English-German) corpus of image descriptions, where the English part is extracted from the COCO dataset, and the German part are translations by a native German speaker.
2 papers · 0 benchmarks
Includes gold-standard labels for identifying statements of desire, textual evidence for desire fulfillment, and annotations for whether the stated desire is fulfilled given the evidence in the narrative context.
2 papers · 0 benchmarks
DialogUSR dataset covers 23 domains with a multi-step crowd-sourcing procedure.
2 papers · 0 benchmarks
The Dialogue Fairness dataset is used to evaluate and understand fairness in dialogue models, focusing on gender and racial biases.
2 papers · 0 benchmarks
DisKnE (Disease Knowledge Evaluation)
DisKnE is a benchmark for Disease Knowledge Evaluation built from MedNLI and MEDIQA-NLI.
2 papers · 0 benchmarks
Dusha (Dusha Crowd, Dusha Podcast)
Dusha is a dataset for speech emotion recognition (SER) tasks.
2 papers · 2 benchmarks
E-ReDial (Explainable Recommendation Dialogues)
E-ReDial is a conversational recommender system dataset with high-quality explanations.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
EDGAR-CORPUS is a novel corpus comprising annual reports from all the publicly traded companies in the US spanning a period of more than 25 years.
2 papers · 0 benchmarks
EDNA-Covid is a multilingual, large-scale dataset of coronavirus-related tweets collected since January 25, 2020.
2 papers · 0 benchmarks
EFO-1-QA is a new dataset to benchmark the combinatorial generalizability of Complex Query Answering (CQA) models by including 301 different queries types, which is 20 times larger than existing datasets.
2 papers · 0 benchmarks
ENTIGEN (Ethical NaTural Language Interventions in Text-to-Image GENeration)
ENTIGEN is a benchmark dataset to evaluate the change in image generations conditional on ethical interventions across three social axes -- gender, skin color, and culture.
2 papers · 0 benchmarks
To automatically generate Python and assembly programs used for security exploits, we curated a large dataset for feeding NMT techniques.
2 papers · 0 benchmarks
The Eedi dataset contains from two school years (September 2018 to May 2020) of students’ answers to mathematics questions from Eedi, a leading educational platform which millions of students interact with daily around the globe.
2 papers · 0 benchmarks
An open corpus of Scientific Research papers which has a representative sample from across scientific disciplines.
2 papers · 0 benchmarks
EmoCause is a dataset of annotated emotion cause words in emotional situations from the EmpatheticDialogues valid and test set.
2 papers · 1 benchmark
EmoPars is a dataset of 30,000 Persian Tweets labeled with Ekman’s six basic emotions (Anger, Fear, Happiness, Sadness, Hatred, and Wonder).
2 papers · 0 benchmarks
This repository contains essays written by high school Brazilian students.
2 papers · 0 benchmarks
Ethics (per ethics) dataset is created to test the knowledge of the basic concepts of morality.
2 papers · 1 benchmark
A multilingual etymological database extracted from the Wiktionary (described in Methodological Aspects of Developing and Managing an Etymological Lexical Resource: Introducing EtymDB-2.0)
2 papers · 0 benchmarks
The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments.
2 papers · 0 benchmarks
ExHVV is a novel dataset that offers natural language explanations of connotative roles for three types of entities -- heroes, villains, and victims, encompassing 4,680 entities present in 3K memes.
2 papers · 0 benchmarks
ExpMRC is a benchmark for the Explainability evaluation of Machine Reading Comprehension.
2 papers · 0 benchmarks
The Freebase Annotations of TREC KBA 2014 Stream Corpus with Timestamps (FAKBAT) is an extension of the FAKBA1 dataset that contains entity age and entity timestamp.
2 papers · 0 benchmarks
FIB (Factual Inconsistency Benchmark)
Factual Inconsistency Benchmark (FIB) is a benchmark that focuses on the task of summarization.
2 papers · 0 benchmarks
FIJO (French Insurance Job Offer dataset)
This dataset was collected as part of the multidisciplinary project Femmes face aux défis de la transformation numérique : une étude de cas dans le secteur des assurances (Women Facing the Challenges of Digital Transformation: A Case Study…
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.