Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 21 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 961–1008 of 3,130
SPARTQA (SPAtial Reasoning on Textual Question Answering)
SpartQA is a textual question answering benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior datasets and that is challenging for state-of-the-art language models…
9 papers · 0 benchmarks
SPARTQA - (SPAtial Reasoning on Textual Question Answering.)
We take advantage of the ground truth of NLVR images, design CFGs to generate stories, and use spatial reasoning rules to ask and answer spatial reasoning questions.
9 papers · 0 benchmarks
SciRepEval is a comprehensive benchmark for training and evaluating scientific document representations.
9 papers · 0 benchmarks
The SciTail dataset is an entailment dataset created from multiple-choice science exams and web sentences.
9 papers · 1 benchmark
SelQA is a dataset that consists of questions generated through crowdsourcing and sentence length answers that are drawn from the ten most prevalent topics in the English Wikipedia.
9 papers · 0 benchmarks
The Standardized Project Gutenberg Corpus (SPGC) is an open science approach to a curated version of the complete PG data containing more than 50,000 books and more than 3×109 word-tokens.
9 papers · 0 benchmarks
TEMPO (Localizing Moments in Video with Temporal Language)
TEMPOral reasoning in video and language (TEMPO) is a dataset that consists of two parts: a dataset with real videos and template sentences (TEMPO - Template Language) which allows for controlled studies on temporal language, and a human…
9 papers · 0 benchmarks
Social media are interactive platforms that facilitate the creation or sharing of information, ideas or other forms of expression among people.
9 papers · 1 benchmark
UA-GEC (UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language)
UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language
9 papers · 1 benchmark
UDIVA is a new non-acted dataset of face-to-face dyadic interactions, where interlocutors perform competitive and collaborative tasks with different behavior elicitation and cognitive workload.
9 papers · 0 benchmarks
UIIS (General Underwater Image Instance Segmentation dataset)
This is the first general Underwater Image Instance Segmentation (UIIS) dataset containing 4,628 images for 7 categories with pixel-level annotations for underwater instance segmentation task
9 papers · 1 benchmark
VISUELLE is a repository build upon the data of a real fast fashion company, Nunalie, and is composed of 5577 new products and about 45M sales related to fashion seasons from 2016-2019.
9 papers · 1 benchmark
VLM²-Bench: Benchmarking Vision-Language Models on Visual Cue Matching Description VLM²-Bench is the first comprehensive benchmark designed to evaluate vision-language models' (VLMs) ability to visually link matching cues across…
9 papers · 1 benchmark
We present a new large-scale human value dataset called ValueNet, which contains human attitudes on 21,374 text scenarios.
9 papers · 0 benchmarks
WebCPM is a Chinese LFQA dataset.
9 papers · 0 benchmarks
WiC-TSV (Words-in-Context: Target Sense Verification)
WiC-TSV is a new multi-domain evaluation benchmark for Word Sense Disambiguation.
9 papers · 2 benchmarks
The iWildCam2020-WILDS dataset is a variant of the iWildCam 2020 dataset.
9 papers · 1 benchmark
A scholarly data set with publications’ full-text, annotated in-text citations, and links to metadata.
9 papers · 0 benchmarks
These are 10 synthetic genomics datasets generated with NEAT v3 (based on TP53 gene of Homo Sapiens) for the use case of benchmarking somatic variant callers.
8 papers · 1 benchmark
AmaSum is the largest abstractive opinion summarization dataset, consisting of more than 33,000 human-written summaries for Amazon products.
8 papers · 0 benchmarks
This dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).
8 papers · 1 benchmark
The Arabic Sentiment Twitter Dataset for the Levantine dialect (ArSenTD-LEV) is a dataset of 4,000 tweets with the following annotations: the overall sentiment of the tweet, the target to which the sentiment was expressed, how the…
8 papers · 0 benchmarks
The Arena-Hard-Auto benchmark is an automatic evaluation tool for instruction-tuned Language Learning Models (LLMs)¹.
8 papers · 0 benchmarks
The Bacteria Biotope (BB) Task is part of the BioNLP Open Shared Tasks and meets the BioNLP-OST standards of quality, originality and data formats.
8 papers · 0 benchmarks
CAD (Contextual Abuse Dataset)
Dataset of primarily English Reddit entries which addresses several limitations of prior work.
8 papers · 1 benchmark
CHIP-CDN (Clinical Diagnosis Normalization Dataset)
CHIP Clinical Diagnosis Normalization, a dataset that aims to standardize the terms from the final diagnoses of Chinese electronic medical records, is used for the CHIP-CDN task.
8 papers · 0 benchmarks
CHIP-STS (Semantic Textual Similarity Dataset)
CHIP Semantic Textual Similarity, a dataset for sentence similarity in the non-i.i.d.
8 papers · 1 benchmark
CLEVR-Math is a multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction, represented partly by a textual description and partly by an image illustrating the scenario.
8 papers · 0 benchmarks
CMRC 2019 (Chinese Machine Reading Comprehension 2019)
CMRC 2019 is a Chinese Machine Reading Comprehension dataset that was used in The Third Evaluation Workshop on Chinese Machine Reading Comprehension.
8 papers · 0 benchmarks
Dataset [46 M] and readme: 42,306 movie plot summaries extracted from Wikipedia + aligned metadata extracted from Freebase, including: Movie box office revenue, genre, release date, runtime, and language Character names and aligned…
8 papers · 0 benchmarks
The COCO-MIG benchmark (Common Objects in Context Multi-Instance Generation) is a benchmark used to evaluate the generation capability of generators on text containing multiple attributes of multi-instance objects.
8 papers · 1 benchmark
CoNLL-2000 is a dataset for dividing text into syntactically related non-overlapping groups of words, so-called text chunking.
8 papers · 0 benchmarks
DocNLI is a large-scale dataset for document-level NLI.
8 papers · 0 benchmarks
EUR-Lex-Sum is a dataset for cross-lingual summarization.
8 papers · 0 benchmarks
A GQA-based dataset with 1,040,830 multi-modal explanations of visual reasoning processes.
8 papers · 1 benchmark
GTA (A Benchmark for General Tool Agents)
A benchmark to evaluate the tool-use capabilities of LLM-based agents in real-world scenarios.
8 papers · 0 benchmarks
GVFC (Gun Violence Frame Corpus)
This is a new dataset of news headlines and their frames related to the issue of gun violence in the United States.
8 papers · 0 benchmarks
GermanQuAD is a Question Answering (QA) dataset of 13,722 extractive question/answer pairs in German.
8 papers · 1 benchmark
The official HOList benchmark for automated theorem proving consists of all theorem statements in the core, complex, and flyspeck corpora.
8 papers · 1 benchmark
HumAID (Human-Annotated Disaster Incidents Data)
Social networks are widely used for information consumption and dissemination, especially during time-critical events such as natural disasters.
8 papers · 0 benchmarks
L3CubeMahaSent is a large publicly available Marathi Sentiment Analysis dataset.
8 papers · 0 benchmarks
Logic2Text is a large-scale dataset with 10,753 descriptions involving common logic types paired with the underlying logical forms.
8 papers · 0 benchmarks
LongForm dataset is created by leveraging English corpus examples with augmented instructions.
8 papers · 0 benchmarks
ManyTypes4Py is a large Python dataset for machine learning (ML)-based type inference.
8 papers · 0 benchmarks
OPIEC (Open Information Extraction Corpus)
OPIEC is an Open Information Extraction (OIE) corpus, constructed from the entire English Wikipedia.
8 papers · 0 benchmarks
OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data.
8 papers · 1 benchmark
OpenViDial is a large-scale open-domain dialogue dataset with visual contexts.
8 papers · 0 benchmarks
Source: BARThez: a Skilled Pretrained French Sequence-to-Sequence Model OrangeSum is a single-document extreme summarization dataset with two tasks: title and abstract.
8 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.