Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 18 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 817–864 of 3,130

FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
FM2 (FoolMeTwice)
FoolMeTwice (FM2 for short) is a large dataset of challenging entailment pairs collected through a fun multi-player game.
12 papers · 0 benchmarks
Chinese Few-shot Learning Evaluation Benchmark (FewCLUE) is a comprehensive small sample evaluation benchmark in Chinese.
12 papers · 5 benchmarks
GLGE (General Language Generation Evaluation)
GLGE is a general language generation evaluation benchmark which is composed of 8 language generation tasks, including Abstractive Text Summarization (CNN/DailyMail, Gigaword, XSUM, MSNews), Answer-aware Question Generation (SQuAD 1.1,…
12 papers · 0 benchmarks
The Hutter Prize Wikipedia dataset, also known as enwiki8, is a byte-level dataset consisting of the first 100 million bytes of a Wikipedia XML dump.
12 papers · 1 benchmark
100 tasks from LIBERO-100 suite.
12 papers · 1 benchmark
LLaMEA (algorithms and experiments from the paper)
3500+ Generated evolutionary algorithms by the LLaMEA framework.
12 papers · 0 benchmarks
LANI is a 3D navigation environment and corpus, where an agent navigates between landmarks.
12 papers · 0 benchmarks
Lila is a unified mathematical reasoning benchmark consisting of 23 diverse tasks along four dimensions: (i) mathematical abilities e.g., arithmetic, calculus (ii) language format e.g., question-answering, fill-in-the-blanks (iii) language…
12 papers · 0 benchmarks
Math-Vision (Math-V) dataset is a meticulously curated collection of 3,040 high-quality mathematical problems with visual contexts sourced from real math competitions.
12 papers · 1 benchmark
MM-COVID (Multilingual and Multidimensional COVID-19 Fake News Data Repository)
MM-COVID is a dataset for fake news detection related to COVID-19.
12 papers · 0 benchmarks
MMNeedle (Multimodal Needle in a Haystack)
We introduce the MultiModal Needle-in-a-haystack (MMNeedle) benchmark, specifically designed to assess the long-context capabilities of MLLMs.
12 papers · 1 benchmark
MasakhaNEWS is a benchmark dataset for news topic classification covering 16 languages widely spoken in Africa.
12 papers · 0 benchmarks
Neptune (Neptune Long Video Understanding Benchmark)
Neptune is a dataset consisting of challenging question-answer-decoy (QAD) sets for long videos (up to 15 minutes).
12 papers · 0 benchmarks
The Overruling dataset is a law dataset corresponding to the task of determining when a sentence is overruling a prior decision.
12 papers · 1 benchmark
PROST (Physical Reasoning about Objects Through Space and Time)
The PROST (Physical Reasoning about Objects Through Space and Time) dataset contains 18,736 multiple-choice questions made from 14 manually curated templates, covering 10 physical reasoning concepts.
12 papers · 0 benchmarks
We present PeerQA, a real-world, scientific, document-level Question Answering (QA) dataset.
12 papers · 3 benchmarks
Co-speech gestures are everywhere.
12 papers · 1 benchmark
The Dataset is part of the KELM corpus This is the Wikipedia text--Wikidata KG aligned corpus used to train the data-to-text generation model.
12 papers · 1 benchmark
Consists of over 39,000 images originating from people who are blind that are each paired with five captions.
12 papers · 0 benchmarks
AKCES-GEC is a new dataset on grammatical error correction for Czech.
11 papers · 0 benchmarks
AMPS (Auxiliary Mathematics Problems and Solutions)
AMPS contains over 100,000 problems pulled from Khan Academy and approximately 5 million problems generated from manually designed Mathematica scripts.
11 papers · 0 benchmarks
ASAP-AES (Automated Student Assessment Prize)
There are eight essay sets.
11 papers · 1 benchmark
BCNB (Early Breast Cancer Core-Needle Biopsy WSI)
Breast cancer (BC) has become the greatest threat to women’s health worldwide.
11 papers · 0 benchmarks
COST (COCO Segmentation Text)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
11 papers · 0 benchmarks
This is the home of a collaborative data collection effort by U.
11 papers · 1 benchmark
This is the home of a collaborative data collection effort by U.
11 papers · 1 benchmark
ClariQ is an extension of the Qulac dataset with additional new topics, questions, and answers in the training set.
11 papers · 0 benchmarks
Com2Sense (Complementary Commonsense)
Complementary Commonsense (Com2Sense) is a dataset for benchmarking commonsense reasoning ability of NLP models.
11 papers · 0 benchmarks
DocILE is a large dataset of business documents for the tasks of Key Information Localization and Extraction and Line Item Recognition.
11 papers · 0 benchmarks
ImageCoDe (Image Retrieval from Contextual Descriptions)
Given 10 minimally contrastive (highly similar) images and a complex description for one of them, the task is to retrieve the correct image.
11 papers · 1 benchmark
Kobest is a benchmark for Korean language reasoning.
11 papers · 0 benchmarks
The MAGE dataset provides a large set of generated texts using 27 LLMs from seven different groups: OpenAI GPT, LLaMA, GLM130B, FLAN-T5, OPT, BigScience, and EleutherAI.
11 papers · 1 benchmark
MAVEN-ERE is a dataset designed for event relation extraction tasks containing 103,193 event coreference chains, 1,216,217 temporal relations, 57,992 causal relations, and 15,841 subevent relations.
11 papers · 0 benchmarks
MultI-Modal In-Context Instruction Tuning (MIMIC-IT) is a dataset for instruction tuning into multi-modal models, motivated by the Flamingo model's upstream interleaved format pretraining dataset.
11 papers · 0 benchmarks
MRR-Benchmark (Multi-Modal Reading Benchmark)
Multi-Modal Reading (MMR) Benchmark includes 550 annotated question-answer pairs across 11 distinct tasks involving texts, fonts, visual elements, bounding boxes, spatial relations, and grounding, with carefully designed evaluation metrics.
11 papers · 1 benchmark
We scrape data from GooBix, which contains 156 games of 5 × 5 mini crosswords.
11 papers · 0 benchmarks
MultiEURLEX is a multilingual dataset for topic classification of legal documents.
11 papers · 0 benchmarks
Ohsumed includes medical abstracts from the MeSH categories of the year 1991.
11 papers · 2 benchmarks
Open PI is the first dataset for tracking state changes in procedural text from arbitrary domains by using an unrestricted (open) vocabulary.
11 papers · 0 benchmarks
PATS (Pose Audio Transcript Style)
PATS dataset consists of a diverse and large amount of aligned pose, audio and transcripts.
11 papers · 0 benchmarks
ParCorFull (Parallel Corpus Annotated with Full Coreference)
ParCorFull is a parallel corpus annotated with full coreference chains that has been created to address an important problem that machine translation and other multilingual natural language processing (NLP) technologies face -- translation…
11 papers · 0 benchmarks
A large-scale collection of visually-grounded, task-oriented dialogues in English designed to investigate shared dialogue history accumulating during conversation.
11 papers · 0 benchmarks
ProtoQA is a question answering dataset for training and evaluating common sense reasoning capabilities of artificial intelligence systems in such prototypical situations.
11 papers · 0 benchmarks
ReCAM (SemEval-2021 Task 4: Reading Comprehension of Abstract Meaning)
Tasks Our shared task has three subtasks.
11 papers · 1 benchmark
SEDE (Stack Exchange Data Explorer)
SEDE is a dataset comprised of 12,023 complex and diverse SQL queries and their natural language titles and descriptions, written by real users of the Stack Exchange Data Explorer out of a natural interaction.
11 papers · 1 benchmark
StylePTB is a fine-grained text style transfer benchmark.
11 papers · 0 benchmarks
TRIPOD (TuRnIng POint Dataset)
TRIPOD contains screenplays and plot synopses with turning point (TP) annotations for 99 movies.
11 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.