Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 60 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2833–2880 of 3,130

TVPReid (Text-to-Video Person Re-identification)
The TVPReid dataset contains 6559 pedestrian videos, each of which is annotated with two text descriptions, for a total of 13118 descriptions.
1 paper · 0 benchmarks
TVRecap a story generation dataset that requires generating detailed TV show episode recaps from a brief summary and a set of documents describing the characters involved.
1 paper · 4 benchmarks
The TWT16 dataset contains ~30k conversations in Twitter, collected from January to June 2016.
1 paper · 0 benchmarks
Dataset Card for "tamil-alpaca" This repository includes a Tamil-translated version of the Alpaca dataset.
1 paper · 0 benchmarks
Dataset Card for "tamil-alpaca" This repository includes a Tamil-translated versions of the Alpaca dataset and a subset of OpenOrca dataset.
1 paper · 0 benchmarks
Taskography (PDDLGym Taskography)
PDDL dataset of Rearrangement tasks in large-scale 3D scene graphs.
1 paper · 0 benchmarks
TempWikiBio is a new data-to-text generation dataset containing more than 4 millions of chronologically ordered revisions of biographical articles from English Wikipedia, each paired with structured personal profiles.
1 paper · 0 benchmarks
We introduce TextAtlas5M, a dataset specifically designed for training and evaluating multimodal generation models on dense-text image generation.
1 paper · 0 benchmarks
TextWorld KG is a dynamic Knowledge Graph (KG) extraction dataset.
1 paper · 0 benchmarks
Content This dataset contains all utterances of two episodes of South Park (Latin American voices) and two episodes of Archer (Spanish voices).
1 paper · 0 benchmarks
The ComMA Dataset v0.2 is a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in which they occur.
1 paper · 0 benchmarks
Using the Experience-Sampling Method (ESM), participants are asked to report TV consumption multiple times each day for a five week period.
1 paper · 0 benchmarks
The EMBO SourceData-NLP dataset (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process.
1 paper · 1 benchmark
The Mafia Dataset was created to model the behavior of deceptive actors in the context of the Mafia game, as described in the paper “Putting the Con in Context: Identifying Deceptive Actors in the Game of Mafia”.
1 paper · 0 benchmarks
The Reddit Climate Change Dataset is a dataset of 620K Reddit posts and 4.6M comments - all mentions of the terms "climate" and "change" until 2022-09-01 across the entire Reddit social network.
1 paper · 0 benchmarks
We present a new annotated corpus of written learner English, derived from essays submitted to the learning platform Write & Improve (W&I).
1 paper · 0 benchmarks
This is not a Dataset (This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models)
We introduce a large semi-automatically generated dataset of ~400,000 descriptive sentences about commonsense knowledge that can be true or false in which negation is present in about 2/3 of the corpus in different forms that we use to…
1 paper · 1 benchmark
Thunder-NUBench (Negation Understanding Benchmark) is a benchmark specifically designed to evaluate large language models’ (LLMs) sentence-level understanding of negation.
1 paper · 0 benchmarks
Question answering over temporal knowledge graphs (TKGs) is crucial for understanding evolving facts and relationships, yet its development is hindered by limited datasets and difficulties in generating custom QA pairs.
1 paper · 0 benchmarks
TinySocial is a dataset to enable research on Social Visual Question Answering.
1 paper · 0 benchmarks
The data consists of a set of 3 task types and 4 question types, creating 12 total scenarios.
1 paper · 0 benchmarks
A prevalent use case of topic models is that of topic discovery.
1 paper · 1 benchmark
TraVLR is a synthetic dataset comprising four visio-linguistic reasoning tasks.
1 paper · 0 benchmarks
This data set is being released to support the spam and context-specific spam detection tasks on Twitter data.
1 paper · 3 benchmarks
Trailers12k is a movie trailer dataset comprised of 12,000 titles associated to ten genres.
1 paper · 0 benchmarks
Translated SNLI Dataset in Marathi A translated version of the SNLI dataset in Marathi, designed for Semantic Textual Similarity (STS) tasks.
1 paper · 1 benchmark
Trilemma Dataset (The Trilemma of Truth in Large Language Models)
The Trilemma of Truth is a multiclass probing dataset for evaluating the veracity-tracking mechanism of large language models.
1 paper · 0 benchmarks
TruthGen is a dataset of generated true and false statements, intended for research on truthfulness in reward models and language models, specifically in contexts where political bias is undesirable.
1 paper · 0 benchmarks
TuGebic (A Turkish-German Bilingual Code-Switching Corpus)
TuGebic is a corpus of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGebic.
1 paper · 0 benchmarks
TuPyE-Dataset (Portuguese Hate Speech Expanded Dataset)
TuPyE, an enhanced iteration of TuPy, encompasses a compilation of 43,668 meticulously annotated documents specifically selected for the purpose of hate speech detection within diverse social network contexts.
1 paper · 0 benchmarks
The TupleInf Open IE dataset contains Open IE tuples extracted from 263K sentences that were used by the solver in “Answering Complex Questions Using Open Information Extraction” (referred as Tuple KB, T).
1 paper · 0 benchmarks
TurkQA consists of a selection of sentences from English Wikipedia articles, with questions and answers crowdsourced from workers on Amazon Mechanical Turk.
1 paper · 0 benchmarks
Tweet Sentiment Extraction (Sentiment Analysis: Emotion in Text tweets with existing sentiment labels)
"My ridiculous dog is amazing." [sentiment: positive] With all of the tweets circulating every second it is hard to tell whether the sentiment behind a specific tweet will impact a company, or a person's, brand for being viral (positive),…
1 paper · 0 benchmarks
This dataset for abusive content detection in Twitter consists of two sets of annotations for the same set of tweets, one where the human annotators had access to the tweet's content and one where they didn't know the context.
1 paper · 0 benchmarks
Twitter Cyberthreat Detection Dataset is a dataset that contains tweets from two sets of accounts related to cybersecurity.
1 paper · 0 benchmarks
This dataset contains two subsets of flood images from Twitter: The Harz17 dataset comprises images from tweets containing flood-related keywords during the occurrence of a flood in the Harz region in Germany in July 2017.
1 paper · 0 benchmarks
Twitter MediaEval (MediaEval Benchmarking Initiative for Multimedia Evaluation)
The task addresses the problem of the appearance and propagation of posts that share misleading multimedia content (images or video).
1 paper · 0 benchmarks
Twitter PoS VCB (Twitter part-of-speech vote-constrained-bootstrapping)
The data is about 1.5 million English tweets annotated for part-of-speech using Ritter's extension of the PTB tagset.
1 paper · 0 benchmarks
We introduce a dataset consisting of 1314 samples, including users’ tweets and bios.
1 paper · 0 benchmarks
This task aims to extract named entities and entity types while further predicting segmentation masks of visual objects.
1 paper · 1 benchmark
UHGEvalDataset contains over 5000 news items.
1 paper · 0 benchmarks
UICaption is a dataset of 114k UI images paired with descriptions of their functionality.
1 paper · 0 benchmarks
This dataset comprises over 26,000 full names annotated with genders.
1 paper · 0 benchmarks
Definitions of jargon/terms in computer science, mathematics, and physics
1 paper · 0 benchmarks
UK Key Stage Readability (UK Key Stage Readability for English Texts)
Education is increasingly data-driven, and the ability to analyse and adapt educational materials quickly and effectively is important for keeping materials contemporary and interesting.
1 paper · 1 benchmark
Bangladesh's legal system struggles with major challenges like delays, complexity, high costs, and millions of unresolved cases, which deter many from pursuing legal action due to lack of knowledge or financial constraints.
1 paper · 0 benchmarks
UNER v1 (Universal NER v1)
UNER v1 adds an NER annotation layer to 18 datasets (primarily treebanks from UD) and covers 12 geneologically and ty- pologically diverse languages: Cebuano, Danish, German, English, Croatian, Portuguese, Russian, Slovak, Serbian,…
1 paper · 31 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.