Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 31 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1441–1488 of 3,130
VANiLLa is a dataset for Question Answering over Knowledge Graphs (KGQA) offering answers in natural language sentences.
4 papers · 0 benchmarks
Visuelle 2.0 is a dataset containing real data for 5355 clothing products of the retail fast-fashion Italian company, Nuna Lie.
4 papers · 2 benchmarks
A Large Vision-Language Model Knowledge Editing Benchmark
4 papers · 0 benchmarks
ViMMRC (Vietnamese Multiple-choice Machine Reading Comprehension Corpus)
A challenging machine comprehension corpus with multiple-choice questions, intended for research on the machine comprehension of Vietnamese text.
4 papers · 0 benchmarks
The VideoNavQA dataset contains pairs of questions and videos generated in the House3D environment.
4 papers · 0 benchmarks
VietMed (VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain)
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
4 papers · 2 benchmarks
The VizWiz-VQA-Grounding dataset is a dataset that visually grounds answers to visual questions asked by people with visual impairments.
4 papers · 0 benchmarks
WEC-eng is a cross-document event coreference resolution dataset extracted from English Wikipedia.
4 papers · 0 benchmarks
WIKIPerson is a high-quality human-annotated visual person linking dataset based on Wikipedia.
4 papers · 0 benchmarks
WikiCLIR is a large-scale (German-English) retrieval data set for Cross-Language Information Retrieval (CLIR).
4 papers · 0 benchmarks
WikiGraphs is a dataset of Wikipedia articles each paired with a knowledge graph, to facilitate the research in conditional text generation, graph generation and graph representation learning.
4 papers · 1 benchmark
WikiNLDB is a novel dataset for training Natural Language Databases (NLDBs) which is generated by transforming structured data from Wikidata into natural language facts and queries.
4 papers · 0 benchmarks
The WikiSem500 dataset contains around 500 per-language cluster groups for English, Spanish, German, Chinese, and Japanese (a total of 13,314 test cases).
4 papers · 0 benchmarks
SRL is the task of extracting semantic predicate-argument structures from sentences.
4 papers · 0 benchmarks
A large-scale dataset built on questions from TyDi QA lacking same-language answers.
4 papers · 0 benchmarks
XTD10 is a dataset for cross-lingual image retrieval and tagging consisting of the MSCOCO2014 caption test dataset annotated in 7 languages that were collected using a crowdsourcing platform.
4 papers · 0 benchmarks
XWINO is a multilingual collection of Winograd Schemas in six languages that can be used for evaluation of cross-lingual commonsense reasoning capabilities.
4 papers · 1 benchmark
ZhihuRec dataset is collected from a knowledge-sharing platform (Zhihu), which is composed of around 100M interactions collected within 10 days, 798K users, 165K questions, 554K answers, 240K authors, 70K topics, and more than 501K user…
4 papers · 0 benchmarks
This dataset contains 1304 de-identified longitudinal medical records describing 296 patients.
4 papers · 1 benchmark
This is a dataset for disentangling conversations on IRC, which is the task of identifying separate conversations in a single stream of messages.
4 papers · 3 benchmarks
Pn-summary is a dataset for Persian abstractive text summarization.
4 papers · 0 benchmarks
This dataset gathers 10,874 title and abstract pairs from the ACL Anthology Network (until 2016).
3 papers · 1 benchmark
ACL-Fig is a large-scale automatically annotated corpus consisting of 112,052 scientific figures extracted from 56K research papers in the ACL Anthology.
3 papers · 0 benchmarks
ADE-Affordance is a new dataset that builds upon ADE20k, which contains annotations enabling such rich visual reasoning.
3 papers · 0 benchmarks
The AROT-COV23 (ARabic Original Tweets on COVID-19 as of 2023) dataset is a large-scale collection of original Arabic tweets related to COVID-19, spanning from January 2020 to January 2023, and the period for which we collected the data…
3 papers · 0 benchmarks
AS-V2 (The All-Seeing Dataset v2)
We propose a novel task, termed Relation Conversation (ReC), which unifies the formulation of text generation, object localization, and relation comprehension.
3 papers · 0 benchmarks
All Words Open IE (AW-OIE) is an open information extraction dataset derived from Question-Answer Meaning Representation (QAMR) dataset.
3 papers · 0 benchmarks
AfriQA is a cross-lingual QA dataset with a focus on African languages.
3 papers · 0 benchmarks
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
3 papers · 1 benchmark
Is a new open-domain question answering task which involves predicting a set of question-answer pairs, where every plausible answer is paired with a disambiguated rewrite of the original question.
3 papers · 0 benchmarks
ArzEn (Corpus of Egyptian Arabic-English Code-switching)
Corpus of Egyptian Arabic-English Code-switching (ArzEn) is a spontaneous conversational speech corpus, obtained through informal interviews held at the German University in Cairo.
3 papers · 0 benchmarks
BCOPA-CE (A Balanced COPA Test Set with cause-effect as alternatives)
We provide the BCOPA-CE test set, which has balanced token distribution in the correct and wrong alternatives and increases the difficulty of being aware of cause and effect.
3 papers · 0 benchmarks
BUSTER (BUSiness Transaction Entity Recognition dataset.)
BUSiness Transaction Entity Recognition dataset.
3 papers · 0 benchmarks
BenchIE: a benchmark and evaluation framework for comprehensive evaluation of OIE systems for English, Chinese and German.
3 papers · 1 benchmark
Contains 3,689,229 English news articles on politics, gathered from 11 United States (US) media outlets covering a broad ideological spectrum.
3 papers · 0 benchmarks
BnB is a large-scale and diverse in-domain VLN (Vision and Language Navigation) dataset.
3 papers · 0 benchmarks
Bongard-OpenWorld is a new benchmark for evaluating real-world few-shot reasoning for machine vision.
3 papers · 1 benchmark
The BuzzFeed-Webis Fake News Corpus 16 comprises the output of 9 publishers in a week close to the US elections.
3 papers · 0 benchmarks
CANNOT (Compilation of ANnotated, Negation-Oriented Text-pairs)
Dataset Summary CANNOT is a dataset that focuses on negated textual pairs.
3 papers · 0 benchmarks
CATT (CATT Arabic Diacritization Benchmark Dataset)
The CATT benchmark dataset comprises 742 sentences, which were scraped from an internet news source in 2023.
3 papers · 1 benchmark
CHQ-Summ (Consumer Healthcare Question Summarization)
Contains 1507 domain-expert annotated consumer health questions and corresponding summaries.
3 papers · 0 benchmarks
CICEROv2 (Contextualized Commonsense Inference in Dialogues (V2))
The CICEROv2 dataset can be found in the data directory.
3 papers · 1 benchmark
This dataset contains plot summaries for 16,559 books extracted from Wikipedia, along with aligned metadata from Freebase, including book author, title, and genre.
3 papers · 0 benchmarks
CNewSum is a large-scale Chinese news summarization dataset which consists of 304,307 documents and human-written summaries for the news feed.
3 papers · 0 benchmarks
Semi-Structured Explanations for COPA (COPA-SSE) is a new crowdsourced dataset of 9,747 semi-structured, English common sense explanations for COPA questions.
3 papers · 0 benchmarks
We introduce FUNSD-r and CORD-r in Token Path Prediction, the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.
3 papers · 1 benchmark
COVID-CQ is a stance data set of user-generated content on Twitter in the context of COVID-19.
3 papers · 0 benchmarks
COVID-Q consists of COVID-19 questions which have been annotated into a broad category (e.g.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.