Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 17 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 769–816 of 3,130

A new large scale plane geometry problem solving dataset called PGPS9K, labeled both fine-grained diagram annotation and interpretable solution program.
14 papers · 1 benchmark
SIBR (SIBR Dataset for VIE in the Wild)
SIBR是面向自然场景视觉信息抽取的数据集。 1)SIBR总的有1000张图片,400张测试,600张训练,包括中文、英文两种语言。 2)包含images.zip、label.zip、train.txt、test.txt四个文件,images.zip、label.zip中包含所有图片和标签,通过train.txt和test.txt区分训练和测试。…
14 papers · 1 benchmark
SOREL-20M (Sophos/ReversingLabs-20 Million)
SOREL-20M is a large-scale dataset consisting of nearly 20 million files with pre-extracted features and metadata, high-quality labels derived from multiple sources, information about vendor detections of the malware samples at the time of…
14 papers · 0 benchmarks
ANTIQUE is a collection of 2,626 open-domain non-factoid questions from a diverse set of categories.
13 papers · 0 benchmarks
Amazon Baby (Amazon Baby 5-core)
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
13 papers · 2 benchmarks
Created by Smith et al.
13 papers · 2 benchmarks
A large-scale cloze-style biomedical MRC dataset.
13 papers · 1 benchmark
DialFact is a testing benchmark dataset of 22,245 annotated conversational claims, paired with pieces of evidence from Wikipedia.
13 papers · 0 benchmarks
GUE (Genome Understanding Evaluation)
A collection of $28$ datasets across $7$ tasks constructed for genome language model evaluation.
13 papers · 7 benchmarks
GUM (Georgetown University Multilayer corpus)
GUM is an open source multilayer English corpus of richly annotated texts from twelve text types.
13 papers · 1 benchmark
The Ghostbusters dataset leverages the GPT-3.5-turbo model for generating texts in the domains of creative writing, news, and student essays, providing 2,000 texts in the first two domains and 1,994 in the latter.
13 papers · 1 benchmark
GooAQ is a large-scale dataset with a variety of answer types.
13 papers · 0 benchmarks
HINT3 is a dataset for intent detection.
13 papers · 0 benchmarks
JEC-QA is a LQA (Legal Question Answering) dataset collected from the National Judicial Examination of China.
13 papers · 0 benchmarks
The KLEJ benchmark (Kompleksowa Lista Ewaluacji Językowych) is a set of nine evaluation tasks for the Polish language understanding task.
13 papers · 0 benchmarks
KorSTS is a dataset for semantic textural similarity (STS) in Korean.
13 papers · 0 benchmarks
MEDIQA-AnS (MEDIQA-Answer Summarization)
The first summarization collection containing question-driven summaries of answers to consumer health questions.
13 papers · 0 benchmarks
MedConceptsQA - Open Source Medical Concepts QA Benchmark The benchmark can be found here: https://huggingface.co/datasets/ofir408/MedConceptsQA
13 papers · 2 benchmarks
Multilingual Reuters (Multilingual Reuters Collection)
The Multilingual Reuters Collection dataset comprises over 11,000 articles from six classes in five languages, i.e., English (E), French (F), German (G), Italian (I), and Spanish (S).
13 papers · 0 benchmarks
The NaturalProofs Dataset is a large-scale dataset for studying mathematical reasoning in natural language.
13 papers · 0 benchmarks
It contains 15K triplets of essay problem statements, student-written, and LLM-generated essays.
13 papers · 0 benchmarks
PARANMT-50M is a dataset for training paraphrastic sentence embeddings.
13 papers · 0 benchmarks
PersonalDialog is a large-scale multi-turn dialogue dataset containing various traits from a large number of speakers.
13 papers · 0 benchmarks
Most existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them.
13 papers · 1 benchmark
ROSCOE is a suite of interpretable, unsupervised automatic scores that improve and extend previous text generation evaluation metrics.
13 papers · 0 benchmarks
The goal of the Robust track is to improve the consistency of retrieval technology by focusing on poorly performing topics.
13 papers · 1 benchmark
Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment.
13 papers · 2 benchmarks
A new multi-target dataset of 5.4K TLDRs over 3.2K papers.
13 papers · 0 benchmarks
The COLOSSEUM (The COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation)
To realize effective large-scale, real-world robotic applications, we must evaluate how well our robot policies adapt to changes in environmental conditions.
13 papers · 1 benchmark
UPFD (User Preference-aware Fake News Detection)
For benchmarking, please refer to its variant UPFD-POL and UPFD-GOS.
13 papers · 0 benchmarks
The VQA-CP dataset was constructed by reorganizing VQA v2 such that the correlation between the question type and correct answer differs in the training and test splits.
13 papers · 1 benchmark
ViQuAE is a dataset for KVQAE (Knowledge-based Visual Question Answering about named Entities), a task which consists in answering questions about named entities grounded in a visual context using a Knowledge Base.
13 papers · 0 benchmarks
Visual Madlibs is a dataset consisting of 360,001 focused natural language descriptions for 10,738 images.
13 papers · 0 benchmarks
Who-did-What (Who did What)
Who-did-What collects its corpus from news and provides options for questions similar to CBT.
13 papers · 0 benchmarks
This is a dataset for evaluating summarisation methods for research papers.
13 papers · 3 benchmarks
This dataset contains 98k 2-hop explanations for questions in the QASC dataset, with annotations indicating if they are valid (~25k) or invalid (~73k) explanations.
13 papers · 0 benchmarks
AE-110k (AliExpress - 110k)
The dataset contains product information from AliExpress Sports & Entertainment category.
12 papers · 2 benchmarks
BenchLMM (BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models)
Large Multimodal Models (LMMs) such as GPT-4V and LLaVA have shown remarkable capabilities in visual reasoning with common image styles.
12 papers · 1 benchmark
Bongard-HOI testifies to which extent your few-shot visual learner can quickly induce the true HOI concept from a handful of images and perform reasoning with it.
12 papers · 1 benchmark
This paper introduces the Broad Twitter Corpus (BTC), which is not only significantly bigger, but sampled across different regions, temporal periods, and types of Twitter users.
12 papers · 2 benchmarks
CDCP (Cornell eRulemaking Corpus)
The Cornell eRulemaking Corpus – CDCP is an argument mining corpus annotated with argumentative structure information capturing the evaluability of arguments.
12 papers · 3 benchmarks
CDR (BioCreative V CDR Task Corpus)
The BioCreative V CDR task corpus is manually annotated for chemicals, diseases and chemical-induced disease (CID) relations.
12 papers · 2 benchmarks
CICERO (Contextualized Commonsense Inference in Dialogues)
CICERO contains 53,000 inferences for five commonsense dimensions -- cause, subsequent event, prerequisite, motivation, and emotional reaction -- collected from 5600 dialogues.
12 papers · 4 benchmarks
COVID-19 Fake News Dataset (COVID19 Fake News Detection in English)
Along with COVID-19 pandemic we are also fighting an infodemic'.
12 papers · 1 benchmark
ConditionalQA is a Question Answering (QA) dataset that contains complex questions with conditional answers, i.e.
12 papers · 1 benchmark
DSTC7 Task 1 (Dialog System Technology Challenges Task 1)
The DSTC7 Task 1 dataset is a dataset and task for goal-oriented dialogue.
12 papers · 1 benchmark
ECTSum is a dataset with transcripts of earnings calls (ECTs), hosted by public companies, as documents, and short experts-written telegram-style bullet point summaries derived from corresponding Reuters articles.
12 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.