Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 47 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 2209–2256 of 3,130
Conic10K is an open-ended math problem dataset on conic sections in Chinese senior high school education.
1 paper · 0 benchmarks
Description - Repository: Code, Page, Data - Paper: arxiv.org/abs/2411.17440 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a star and citation.
1 paper · 0 benchmarks
Enumerate–Conjecture–Prove: Formally Solving Answer-Construction Problem in Math Competitions We release the ConstructiveBench dataset as part of our Enumerate–Conjecture–Prove (ECP) paper.
1 paper · 0 benchmarks
This is a revised and extended second version of a Contextualised Polyseme Word Sense Dataset.
1 paper · 0 benchmarks
Corpus of controversial news articles extracted from Twitter.
1 paper · 0 benchmarks
ConvSumX is a cross-lingual conversation summarization benchmark, through a new annotation schema that explicitly considers source input context.
1 paper · 0 benchmarks
CoreSearch is a dataset for Cross-Document Event Coreference Search.
1 paper · 0 benchmarks
CoronaVis is a dataset of tweets related to coronavirus.
1 paper · 0 benchmarks
Using Council Data Project infrastructures (https://councildataproject.org), we assemble longitudinal municipal council meeting transcript data.
1 paper · 0 benchmarks
Probing cross-modal capabilities of Vision & Language models with a counting task.
1 paper · 0 benchmarks
CoverageEval is a dataset specifically designed for evaluating LLMs on this task.
1 paper · 0 benchmarks
The Creative Visual Storytelling Anthology is a collection of 100 author responses to an improved creative visual storytelling exercise over a sequence of three images.
1 paper · 0 benchmarks
This dataset focuses on 50 articles about climate science, which were annotated completely by 49 students, 26 Upwork workers, 3 science and 3 journalism experts.
1 paper · 0 benchmarks
Official dataset of Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP.
1 paper · 0 benchmarks
The dataset contains 30 million cryptocurrency-related tweets from 10.10.2020 to 3.3.2021.
1 paper · 0 benchmarks
Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language.
1 paper · 0 benchmarks
D-OCC (Dynamic-OneCommon Corpus)
D-OCC is a large-scale dataset of 5,617 dialogues to enable fine-grained evaluation and analysis of various dialogue systems.
1 paper · 0 benchmarks
A large benchmark dataset containing 50K human judgments for 5K distinct sentence pairs in the English dative alternation.
1 paper · 0 benchmarks
DAPFAM (A Domain‑Aware Patent Retrieval Dataset Aggregated at the Family Level)
Dataset DAPFAM See the accompanying paper: Ayaou et al., 2025 — “DAPFAM: A Domain‑Aware Patent Retrieval Dataset Aggregated at the Family Level” (arXiv:2506.22141).
1 paper · 0 benchmarks
DAVIS-Edit is a curated testing benchmark for video editing.
1 paper · 0 benchmarks
The dataset provides the content of all articles for 128 Wikipedia languages.
1 paper · 0 benchmarks
DEplain-APA-doc: A German Parallel Corpus for Document Simplification on News Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
1 paper · 1 benchmark
DEplain-web-doc: A German Parallel Corpus for Document Simplification on Web Texts DEplain is a new dataset of parallel, professionally written and manually aligned simplifications in plain German “plain DE” (or in German: “Einfache…
1 paper · 1 benchmark
Medical VQA dataset built from the IDRiD and eOphta datasets.
1 paper · 0 benchmarks
DODa (Darija Open Dataset)
Darija Open Dataset (DODa) is an open-source project for the Moroccan dialect.
1 paper · 0 benchmarks
The dataset was collected from DOTA 2 using OpenDota API via Python.
1 paper · 0 benchmarks
DUC 2006 (Document Understanding Conferences)
There is currently much interest and activity aimed at building powerful multi-purpose information systems.
1 paper · 0 benchmarks
DailyMoth-70h is a fully self-contained ASL-to-English sign language dataset containing over 70h of video (48K clips) with aligned English captions of a single native ASL signer (white, male, and early middle-aged) from the ASL news…
1 paper · 0 benchmarks
A dataset of images obtained from DALL-E 3 for 67 countries and 10 concept classes, similar to DollarStreet images.
1 paper · 0 benchmarks
About the study This study was exploring the landscape of interpersonal conflicts during code review in following areas: - how these conflicts look like - what role do they play in software development - what are their consequences - what…
1 paper · 0 benchmarks
Source: Linking Datasets on Organizations Using Half-a-Billion Open-Collaborated Records (Description (Markdown and LATEX enabled)) High-Level Explanation of the Dataset - Scale and Composition: This repository provides millions of…
1 paper · 0 benchmarks
DataCLUE is the first Data-Centric benchmark applied in NLP field.
1 paper · 0 benchmarks
This dataset contains around 218K sentences, with 1.5 million words, from 30 different books designed for Post-OCR text correction.
1 paper · 0 benchmarks
This data is for the Mis2-KDD 2021 under review paper: Dataset of Propaganda Techniques of the State-Sponsored Information Operation of the People’s Republic of China We present our dataset that focuses on propaganda techniques in Mandarin…
1 paper · 1 benchmark
This is the dataset used for classifying Gene-Disease relationship types from sentences.
1 paper · 1 benchmark
DeepParliament is a legal domain Benchmark Dataset that gathers bill documents and metadata and performs various bill status classification tasks.
1 paper · 0 benchmarks
Corpus for argument mining in legal documents, composed of 40 decisions of the Court of Justice of the European Union on matters of fiscal state aid
1 paper · 0 benchmarks
A curated dataset of 221 question-answer-rationale triples capturing visualization design decisions and the reasoning behind them, derived from real-world student-authored narratives.
1 paper · 0 benchmarks
Dhoroni (Dhoroni: A Multi-Perspective Bengali Climate Change and Environmental News Dataset)
Climate change poses critical challenges globally, disproportionately affecting low-income countries that often lack resources and linguistic representation on the international stage.
1 paper · 1 benchmark
DiaKG is a high-quality Chinese dataset for Diabetes knowledge graph.
1 paper · 0 benchmarks
Dialog-based Language Learning dataset is designed to measure how well models can perform at learning as a student given a teacher’s textual responses to the student’s answer (as well as potentially receiving an external real-valued reward…
1 paper · 0 benchmarks
DiscoSense is a benchmark sourced from datasets that contain two sentences connected through a discourse connective.
1 paper · 0 benchmarks
Dissonance Twitter Dataset is a dataset collected from annotating tweets for dissonance.
1 paper · 0 benchmarks
This dataset is named as the DistNLI dataset, which is a synthesized benchmark aiming to probe neural network models from the aspect of conjunctions on distributivity in NLI task in American English.
1 paper · 0 benchmarks
This dataset consisting 500 set of caption, table and coresponding paper page, processed from DocBank.
1 paper · 0 benchmarks
We manually annotate 800 sentences from 80 documents in two domains (Healthcare and Transportation) to form a DocOIE dataset for evaluation.
1 paper · 2 benchmarks
The DocRED Information Extraction (DocRED-IE) dataset extends the DocRED dataset for the Document-level Closed Information Extraction (DocIE) task.
1 paper · 6 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.