Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 25 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1153–1200 of 3,130

RELX is a benchmark dataset for cross-lingual relation classification in English, French, German, Spanish and Turkish.
6 papers · 0 benchmarks
Rad-ReStruct is a fine-grained structured reporting dataset for Chest X-Ray images.
6 papers · 0 benchmarks
10,000 news collected from a social network in Vietnam.
6 papers · 0 benchmarks
RealCQA Scientific Chart Question Answering as a Test-bed for First-Order Logic check on huggingface : https://huggingface.co/datasets/sal4ahm/RealCQA
6 papers · 1 benchmark
Reddit Conversation Corpus (RCC) consists of conversations, scraped from Reddit, for a 20 month period from November 2016 until August 2018.
6 papers · 0 benchmarks
SMART-101 (Simple Multimodal Algorithmic Reasoning Task Dataset)
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, ChatGPT, etc.
6 papers · 0 benchmarks
SWSR (Sina Weibo Sexism Review)
The Sina Weibo Sexism Review (SWSR) dataset is a dataset to research online sexism in Chinese.
6 papers · 0 benchmarks
Dataset with 625,000 ethical judgments over 32,000 real-life anecdotes.
6 papers · 0 benchmarks
SemEval 2014 is a collection of datasets used for the Semantic Evaluation (SemEval) workshop, an annual event that focuses on the evaluation and comparison of systems that can analyze diverse semantic phenomena in text.
6 papers · 0 benchmarks
A multimodal dataset for sentiment analysis on internet memes.
6 papers · 0 benchmarks
ShapeTalk (The ShapeTalk Dataset)
ShapeTalk contains over half a million discriminative utterances produced by contrasting the shapes of common 3D objects for a variety of object classes and degrees of similarity.
6 papers · 0 benchmarks
SherLIiC is a testbed for lexical inference in context (LIiC), consisting of 3985 manually annotated inference rule candidates (InfCands), accompanied by (i) ~960k unlabeled InfCands, and (ii) ~190k typed textual relations between Freebase…
6 papers · 0 benchmarks
SuperCLUE is a Chinese language model evaluation benchmark named after another popular Chinese LLM benchmark CLUE.
6 papers · 0 benchmarks
Swords (Stanford Word Substitution benchmark)
Swords (Standford Word Substitution) is a benchmark for lexical substitution, the task of finding appropriate substitutes for a target word in a context.
6 papers · 0 benchmarks
TREC-10 (TREC-10 Question Classification)
A question type classification dataset with 6 classes for questions about a person, location, numeric information, etc.
6 papers · 1 benchmark
TURL (Twitter News URL Corpus)
Twitter News URL Corpus is a human-labeled paraphrase corpus to date of 51,524 sentence pairs and the first cross-domain benchmarking for automatic paraphrase identification.
6 papers · 1 benchmark
The TalkSumm dataset contains 1705 automatically-generated summaries of scientific papers from ACL, NAACL, EMNLP, SIGDIAL (2015-2018), and ICML (2017-2018).
6 papers · 0 benchmarks
TutorialBank is a publicly available dataset which aims to facilitate NLP education and research.
6 papers · 0 benchmarks
We present a further analysis of visual modality incompleteness, benchmarking latest MMEA models on our proposed dataset MMEA-UMVM.
6 papers · 3 benchmarks
V2C (Video-to-Commonsense)
6 papers · 0 benchmarks
VisPro dataset contains coreference annotation of 29,722 pronouns from 5,000 dialogues.
6 papers · 0 benchmarks
WDC Products is an entity matching benchmark which provides for the systematic evaluation of matching systems along combinations of three dimensions while relying on real-word data.
6 papers · 4 benchmarks
WNLaMPro (WordNet Language Model Probing)
The WordNet Language Model Probing (WNLaMPro) dataset consists of relations between keywords and words.
6 papers · 0 benchmarks
Multi-level Benchmark of Watermarks for Large Language Models
6 papers · 0 benchmarks
WebLINX (Real-World Website Navigation with Multi-Turn)
WebLINX is a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation.
6 papers · 1 benchmark
An unsupervised dataset for co-reference resolution.
6 papers · 0 benchmarks
WikiNEuRal is a high-quality automatically-generated dataset for Multilingual Named Entity Recognition.
6 papers · 0 benchmarks
WikiNews Dataset (WikiNews Arabic Diacritization Benchmark Dataset)
The WikiNews Arabic Diacritization dataset is a test set composed of 70 WikiNews articles (majority are from 2013 and 2014) that cover a variety of themes, namely: politics, economics, health, science and technology, sports, arts, and…
6 papers · 0 benchmarks
Wild-Time is a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including patient prognosis and news classification.
6 papers · 0 benchmarks
XL-BEL is a benchmark for cross-lingual biomedical entity linking (XL-BEL).
6 papers · 0 benchmarks
XQA is a data which consists of a total amount of 90k question-answer pairs in nine languages for cross-lingual open-domain question answering.
6 papers · 0 benchmarks
The ZS-F-VQA dataset is a new split of the F-VQA dataset for zero-shot problem.
6 papers · 1 benchmark
Benchmark dataset for abstracts and titles of 100,000 ArXiv scientific papers.
6 papers · 1 benchmark
legalNER is a corpus of 46545 annotated legal named entities mapped to 14 legal entity types.
6 papers · 0 benchmarks
ADVErsarial Table perturbAtion (ADVETA) is a robustness evaluation benchmark featuring natural and realistic ATPs.
5 papers · 0 benchmarks
AMALGUM (A Machine Annotated Lookalike of GUM)
AMALGUM is a machine annotated multilayer corpus following the same design and annotation layers as GUM, but substantially larger (around 4M tokens).
5 papers · 0 benchmarks
BB-norm-habitat (Bacteria Biotope - entity normalization - bacterial habitat)
In the BB-norm modality of this task, participant systems had to normalize textual entity mentions according to the OntoBiotope ontology for habitats.
5 papers · 0 benchmarks
BB-norm-phenotype (Bacteria Biotope - entity normalization - phenotype)
In the BB-norm modality of this task, participant systems had to normalize textual entity mentions according to the OntoBiotope ontology for phenotypes.
5 papers · 0 benchmarks
This dataset consists of images and annotations in Bengali.
5 papers · 1 benchmark
BioCoder is a benchmark developed to evaluate existing pre-trained models in generating bioinformatics code.
5 papers · 0 benchmarks
BioNLI (Biomedical Natural Language Inference)
BioNLI is a dataset in biomedical natural language inference.
5 papers · 1 benchmark
BIWI 3D corpus comprises a total of 1109 sentences uttered by 14 native English speakers (6 males and 8 females).
5 papers · 1 benchmark
The dataset offers tag and mask annotations for image-text pairs from the CC3M validation set.
5 papers · 2 benchmarks
CCPM (Chinese Classical Poetry Matching)
Introduction CCPM is a large Chinese classical poetry matching dataset that can be used for poetry matching, understanding and translation.
5 papers · 0 benchmarks
CLIP (CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge Notes)
We created a dataset of clinical action items annotated over MIMIC-III.
5 papers · 0 benchmarks
CQADupStack is a benchmark dataset for community question-answering research.
5 papers · 1 benchmark
CSFCube is an expert annotated test collection to evaluate models trained to perform faceted Query by Example.
5 papers · 0 benchmarks
ChronoMagic with 2265 metamorphic time-lapse videos, each accompanied by a detailed caption.
5 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.