Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 34 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1585–1632 of 3,130

Natural Hazards is a natural disaster dataset with sentiment labels, which contains nearly 50,00 Twitter data about different natural disasters in the United States (e.g., a tornado in 2011, a hurricane named Sandy in 2012, a series of…
3 papers · 0 benchmarks
NusaCrowd is a collaborative initiative to collect and unite existing resources for Indonesian languages, including opening access to previously non-public resources.
3 papers · 0 benchmarks
The dataset contains Amazon products from 10 product categories with full human annotations.
3 papers · 2 benchmarks
ODSQA (Open-Domain Spoken Question Answering)
The ODSQA dataset is a spoken dataset for question answering in Chinese.
3 papers · 0 benchmarks
OGTD (Offensive Greek Tweet Dataset)
A manually annotated dataset containing 4,779 posts from Twitter annotated as offensive and not offensive.
3 papers · 0 benchmarks
OpenCHAIR is a benchmark for evaluating open-vocabulary hallucinations in image captioning models.
3 papers · 0 benchmarks
A benchmark designed to evaluate MLLMs’ proficiency in understanding inter-object relationships and textual content.
3 papers · 0 benchmarks
P3 (Psychophysical Patterns Dataset)
A set of patterns used in psychophysical research to evaluate the ability of saliency algorithms to find targets distinct from distractors in orientation, color and size.
3 papers · 0 benchmarks
PDFVQA: A New Dataset for Real-World VQA on PDF Documents
3 papers · 0 benchmarks
PNT (Parsing Time Normalizations)
The Parsing Time Normalizations (PNT) corpus in SCATE format allows the representation of a wider variety of time expressions than previous approaches.
3 papers · 1 benchmark
An open, broad-coverage corpus for informal Persian named entity recognition was collected from Twitter.
3 papers · 0 benchmarks
PatTR (Patent Translation Resource)
PatTR is a sentence-parallel corpus extracted from the MAREC patent collection.
3 papers · 0 benchmarks
Patzig contains handwritten texts written in modern German.
3 papers · 0 benchmarks
A corpus of 553k news articles from six Persian news websites and agencies with relatively high quality author extracted keyphrases, which is then filtered and cleaned to achieve higher quality keyphrases.
3 papers · 0 benchmarks
Modeling what makes an advertisement persuasive, i.e., eliciting the desired response from consumer, is critical to the study of propaganda, social psychology, and marketing.
3 papers · 0 benchmarks
PhoNERCOVID19 is a dataset for recognising COVID-19 related named entities in Vietnamese, consisting of 35K entities over 10K sentences.
3 papers · 1 benchmark
PubMedCite is a domain-specific dataset with about 192K biomedical scientific papers and a large citation graph preserving 917K citation relationships between them.
3 papers · 0 benchmarks
Large multimodal models extend the impressive capabilities of large language models by integrating multimodal understanding abilities.
3 papers · 0 benchmarks
Q-Pain, a dataset for assessing bias in medical QA in the context of pain management, one of the most challenging forms of clinical decision-making.
3 papers · 0 benchmarks
QALD-9-Plus Dataset Description QALD-9-Plus is the dataset for Knowledge Graph Question Answering (KGQA) based on well-known QALD-9.
3 papers · 1 benchmark
RESD (Russian Emotional Speech Dialogs with annotated text)
Russian dataset of emotional speech dialogues.
3 papers · 1 benchmark
RETWEET is a dataset of tweets and overall predominant sentiment of their replies.
3 papers · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
ROOR is a reading order prediction (ROP) benchmark which annotates layout reading order as ordering relations.
3 papers · 1 benchmark
The RareDis corpus contains more than 5,000 rare diseases and almost 6,000 clinical manifestations are annotated.
3 papers · 0 benchmarks
Ricordi contains handwritten texts written in Italian.
3 papers · 0 benchmarks
RoFT (Real or Fake Text)
RoFT is a dataset of 21,000 human annotations of generated text.
3 papers · 1 benchmark
SCapRepo (Google Play Screenshot Caption)
A screenshot-caption dataset containing 135k pairs of screenshots and captions extracted from Google Play.
3 papers · 0 benchmarks
SDN (Situated Dialogue Navigation)
Situated Dialogue Navigation (SDN) is a navigation benchmark of 183 trials with a total of 8415 utterances, around 18.7 hours of control streams, and 2.9 hours of trimmed audio.
3 papers · 0 benchmarks
A first-of-its-kind large dataset of sarcastic/non-sarcastic tweets with high-quality labels and extra features: (1) sarcasm perspective labels (2) new contextual features.
3 papers · 0 benchmarks
SPOT (Sentiment Polarity Annotations Dataset)
The SPOT dataset contains 197 reviews originating from the Yelp'13 and IMDB collections ([1][2]), annotated with segment-level polarity labels (positive/neutral/negative).
3 papers · 0 benchmarks
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models.
3 papers · 0 benchmarks
SYNTH-PEDES is a large-scale person dataset with image-text pairs by far, which contains 312,321 identities, 4,791,711 images, and 12,138,157 textual descriptions.
3 papers · 0 benchmarks
SaRoCo is a dataset for detecting satire in Romanian news containing 55,608 news articles from multiple real and satirical news sources, of which 27,980 are regular and 27,628 satirical news reports.
3 papers · 0 benchmarks
Schiller (Shiller)
Schiller contains handwritten texts written in modern German.
3 papers · 0 benchmarks
Schwerin contains handwritten texts written in medieval German.
3 papers · 0 benchmarks
Secim2023 is a comprehensive dataset for social media researchers to study the upcoming election, develop tools to prevent online manipulation, and gather novel information to inform the public.
3 papers · 0 benchmarks
Shmoop Corpus is a dataset of 231 stories that are paired with detailed multi-paragraph summaries for each individual chapter (7,234 chapters), where the summary is chronologically aligned with respect to the story chapter.
3 papers · 0 benchmarks
A short clip of video may contain progression of multiple events and an interesting story line.
3 papers · 3 benchmarks
Spanish TimeBank 1.0 was developed by researchers at Barcelona Media and consists of Spanish texts in the AnCora corpus annotated with temporal and event information according to the TimeML specification language.
3 papers · 1 benchmark
Stanceosaurus is a corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.
3 papers · 0 benchmarks
This repository contains a financial-domain-focused dataset for financial sentiment/emotion classification and stock market time series prediction.
3 papers · 0 benchmarks
Presents a new dataset of code snippets with short descriptions, created using data gathered from Stackoverflow, a popular programming help website.
3 papers · 0 benchmarks
TCAB (Text Classification Attack Benchmark)
Text Classification Attack Benchmark (TCAB) is a dataset for analyzing, understanding, detecting, and labeling adversarial attacks against text classifiers.
3 papers · 0 benchmarks
We present the development of a Named Entity Recognition (NER) dataset for Tagalog.
3 papers · 0 benchmarks
The TREC News Track features modern search tasks in the news domain.
3 papers · 1 benchmark
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
3 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.