Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 66 of 66: 10 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 3121–3130 of 3,130

The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
0 papers · 0 benchmarks
TS-TR (Turkish Scene Text Recognition Dataset)
The Turkish Scene Text Recognition (TS-TR) dataset was primarily developed to fill the gap in non-English text recognition resources, specifically addressing the unique challenges presented by the Turkish language, such as special…
0 papers · 0 benchmarks
The Reddit COVID Dataset is a dataset of 4.51M Reddit posts and 17.8M comments - all mentions of COVID until 2021-10-25 across the entire Reddit social network.
0 papers · 0 benchmarks
TiMoS (Tropes in Movie Synopses)
Tropes in Movie Synopses (TiMoS) is a dataset of movie tropes collected from a Wikipedia-style website, TVTropes3 with 5623 movie synopses associated with 95 most occurred tropes.
0 papers · 0 benchmarks
The University of Massachusetts Amherst citation field extraction dataset contains labels and segments for extracted citations from articles found on arXiv.
0 papers · 0 benchmarks
Video Dataset (Storytelling Video Dataset (Russian, Emotion, Gesture, Speech))
The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers.
0 papers · 0 benchmarks
Vript (🎬 Vript: A Video Is Worth Thousands of Words)
We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips).
0 papers · 0 benchmarks
WMT 2021 Ge'ez-Amharic is a Ge'ez-Amharic dataset prepared for NMT tasks of the 6th Workshop on NLP at Debre Berhan University, Ethiopia.
0 papers · 0 benchmarks
X-Wines (A Wine Dataset for Recommender Systems and Machine Learning)
X-Wines is a consistent wine dataset containing 100,646 instances and 21 million real evaluations carried out by users.
0 papers · 0 benchmarks
The prachathai-67k dataset was scraped from the news site Prachathai excluding articles with less than 500 characters of body text (mostly images and cartoons).
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.