Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 22 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1009–1056 of 3,130
Peyma is a Persian NER dataset to train and test NER systems.
8 papers · 0 benchmarks
QA2D (Question to Declarative Sentence (QA2D) Dataset)
The Question to Declarative Sentence (QA2D) Dataset contains 86k question-answer pairs and their manual transformation into declarative sentences.
8 papers · 0 benchmarks
QUASAR-S (QUestion Answering by Search And Reading – Stack Overflow)
QUASAR-S is a large-scale dataset aimed at evaluating systems designed to comprehend a natural language query and extract its answer from a large corpus of text.
8 papers · 0 benchmarks
Consists of multiple sentences whose clues are arranged by difficulty (from obscure to obvious) and uniquely identify a well-known entity such as those found on Wikipedia.
8 papers · 1 benchmark
RFUND (Revised FUNSD and XFUND)
RFUND is a relabeled version of FUNSD and XFUND datasets, tackling the following issues in their original annotations: 1.
8 papers · 0 benchmarks
RIMES (Reconnaissance & Indexation de données Manuscrites et de fac similÉS / Recognition & Indexing of handwritten documents & faxes)
The RIMES database (Reconnaissance et Indexation de données Manuscrites et de fac similÉS / Recognition and Indexing of handwritten documents and faxes) was created to evaluate automatic systems of recognition and indexing of handwritten…
8 papers · 0 benchmarks
RRS (Restoration-200k for Response Selection)
| | Train | Validation | Test | Ranking Test | | --------- | ----- | ---------- | ------- | ------------ | | size | 0.4M | 50K | 5K | 800 | | pos:neg | 1:1 | 1:9 | 1.2:8.8 | - | | avg turns | 5.0 | 5.0 | 5.0 | 5.0 | Ranking test set…
8 papers · 1 benchmark
RUSSE (Russian Words in Context (based on RUSSE))
WiC: The Word-in-Context Dataset A reliable benchmark for the evaluation of context-sensitive word embeddings.
8 papers · 1 benchmark
SciGraphQA is a large-scale, open-domain dataset focused on generating multi-turn conversational question-answering dialogues centered around understanding and describing scientific graphs and figures.
8 papers · 0 benchmarks
NLPContributionGraph was introduced as Task 11 at SemEval 2021 for the first time.
8 papers · 0 benchmarks
We introduce a new audio dataset called SoundDescs that can be used for tasks such as text to audio retrieval, audio captioning etc.
8 papers · 1 benchmark
Spider 2.0 is a comprehensive code generation agent task that includes 632 examples.
8 papers · 1 benchmark
SubjQA is a question answering dataset that focuses on subjective (as opposed to factual) questions and answers.
8 papers · 0 benchmarks
T³Bench is the first comprehensive text-to-3D benchmark containing diverse text prompts of three increasing complexity levels that are specially designed for 3D generation (300 prompts in total).
8 papers · 1 benchmark
TalkDown is a labelled dataset for condescension detection in context.
8 papers · 0 benchmarks
UKP (UKP Argument Annotated Essays)
The UKP Argument Annotated Essays corpus consists of argument annotated persuasive essays including annotations of argument components and argumentative relations.
8 papers · 0 benchmarks
The Video-based Multimodal Summarization with Multimodal Output (VMSMO) corpus consists of 184,920 document-summary pairs, with 180,000 training pairs, 2,460 validation and test pairs.
8 papers · 0 benchmarks
A new dataset describing textual stories for events.
8 papers · 0 benchmarks
VideoXum is an enriched large-scale dataset for cross-modal video summarization.
8 papers · 1 benchmark
WANDS (Wayfair ANnotation Dataset)
The dataset contains: 42,994 candidate products with data comprising product class, title, description, attributes, category hierarchy, average rating, and number of reviews 480 search query strings with predicted product class 233,448…
8 papers · 0 benchmarks
News translation is a recurring WMT task.
8 papers · 0 benchmarks
The WebUI dataset contains 400K web UIs captured over a period of 3 months and cost about $500 to crawl.
8 papers · 0 benchmarks
Who's Waldo is a dataset of 270K image–caption pairs, depicting interactions of people, that is automatically mined from Wikimedia Commons.
8 papers · 1 benchmark
We manually performed the task of Open Information Extraction on 5 short documents, elaborating tentative guidelines for the task, and resulting in a ground truth reference of 347 tuples.
8 papers · 1 benchmark
WikiTableT contains Wikipedia article sections and their corresponding tabular data and various metadata.
8 papers · 0 benchmarks
WildReceipt is a collection of receipts.
8 papers · 0 benchmarks
ACES (A Translation Accuracy Challenge Set)
ACES a dataset consisting of 68 phenomena ranging from simple perturbations at the word/character level to more complex errors based on discourse and real-world knowledge.
7 papers · 1 benchmark
BiSECT is a dataset for sentence simplification, which is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary.
7 papers · 0 benchmarks
CEFR-SP contains 17k English sentences annotated with the levels based on the Common European Framework of Reference for Languages assigned by English-education professionals.
7 papers · 0 benchmarks
In this work we create a question answering dataset over the DBLP scholarly knowledge graph (KG).
7 papers · 0 benchmarks
DaNetQA (Yes/no Question Answering Dataset for the Russian)
DaNetQA is a question answering dataset for yes/no questions.
7 papers · 1 benchmark
Evaluate a natural language code generation model on real data science pedagogical notebooks!
7 papers · 0 benchmarks
Demetr is a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological…
7 papers · 0 benchmarks
GHOSTS is the first natural-language dataset made and curated by working researchers in mathematics that (1) aims to cover graduate-level mathematics and (2) provides a holistic overview of the mathematical capabilities of language models.
7 papers · 0 benchmarks
GINC (Generative IN-Context learning Dataset)
GINC (Generative In-Context learning Dataset) is a small-scale synthetic dataset for studying in-context learning.
7 papers · 0 benchmarks
A large-scale corpus of Gulf Arabic consisting of 110 million words from 1,200 forum novels.
7 papers · 0 benchmarks
HBW (Human Bodies in the Wild)
Human Bodies in the Wild (HBW) is a validation and test set for body shape estimation.
7 papers · 0 benchmarks
HiREST (HIerarchical REtrieval and STep-captioning)
HiREST (HIerarchical REtrieval and STep-captioning) dataset is a benchmark that covers hierarchical information retrieval and visual/textual stepwise summarization from an instructional video corpus.
7 papers · 0 benchmarks
Hindi Visual Genome is a multimodal dataset consisting of text and images suitable for English-Hindi multimodal machine translation task and multimodal research.
7 papers · 0 benchmarks
HyperRED (Hyper-Relational Extraction Dataset)
HyperRED is a dataset for the new task of hyper-relational extraction, which extracts relation triplets together with qualifier information such as time, quantity or location.
7 papers · 1 benchmark
IQUAD (Interactive Question Answering Dataset)
IQUAD is a dataset for Visual Question Answering in interactive environments.
7 papers · 0 benchmarks
KUAKE Query-Query Relevance, a dataset used to evaluate the relevance of the content expressed in two queries, is used for the KUAKE-QQR task.
7 papers · 1 benchmark
KUAKE Query Title Relevance, a dataset used to estimate the relevance of the title of a query document, is used for the KUAKE-QTR task.
7 papers · 1 benchmark
Language-molecule models have emerged as an exciting direction for molecular discovery and understanding.
7 papers · 1 benchmark
LEVEN (Legal Event Detection Dataset)
Overview LEVEN is the largest Legal Event Detection dataset as well as the largest Chinese Event Detection dataset.
7 papers · 0 benchmarks
Large Scale Composed Image Retrieval (LaSCo) is a new dataset for Composed Image Retrieval (CoIR), x10 times larger than current ones.
7 papers · 1 benchmark
This work proposes Long-RVOS, a large-scale benchmark for long-term video object segmentation.
7 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.