Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 50 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2353–2400 of 3,130

This dataset was curated for Search Engine Optimization (SEO) analysis tasks, including categorization and spam detection.
1 paper · 0 benchmarks
GroundCap is a novel grounded image captioning dataset derived from MovieNet, containing 52,350 movie frames with detailed grounded captions.
1 paper · 0 benchmarks
GuardRails Dataset (GuardRails Dataset of Problems with Known Ambiguities)
For each problem, we provide 4 variants of prompts: 1.
1 paper · 0 benchmarks
This is a high-quality dataset consisting of 14.8M utterances in English, extracted from processed dialogues from publicly available online books.
1 paper · 0 benchmarks
Gutenberg Poem Dataset is used for the next verse prediction component.
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
HALvest-Geometric is a subset of HALvest: an academic citation network with 238,397 disambiguated authors and 18,662,037 scholarly papers.
1 paper · 0 benchmarks
HAVOC (Harmful Abstractions and Violations in Open Completions Benchmark)
measure the toxicity generated by language models across input severity and harm categories, by creating a new benchmark of open ended prefixes.
1 paper · 0 benchmarks
HS-BAN is a binary class hate speech (HS) dataset in Bangla language consisting of more than 50,000 labeled comments, including 40.17% hate and rest are non hate speech.
1 paper · 0 benchmarks
HarmfulTasks (Harmful and Malicious Tasks for LLMs in Jailbreaking Prompts)
This dataset consists of 225 malicious tasks, which were integrated into ten distinct jailbreaking prompts.
1 paper · 0 benchmarks
Multi-Modal Hate Speech Detection with Graph Context.
1 paper · 0 benchmarks
Hawk Annotation Dataset includes language descriptions specifically for anomaly scenes in seven existing video anomaly datasets.
1 paper · 0 benchmarks
The Helsinki Prosody Corpus is a dataset for predicting prosodic prominence from written text.
1 paper · 1 benchmark
HiNER-collapsed (HiNER: A Large Hindi Named Entity Recognition Dataset)
This dataset releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 3 collapsed tags (PER, LOC, ORG).
1 paper · 1 benchmark
HiNER-original (HiNER: A Large Hindi Named Entity Recognition Dataset)
This dataset releases a significantly sized standard-abiding Hindi NER dataset containing 109,146 sentences and 2,220,856 tokens, annotated with 11 tags.
1 paper · 1 benchmark
HiXSTest (Hindi XSTest)
For testing refusal behavior in a language-specific setting, we introduce HiXSTest — a set of manually curated prompts in the Hindi language designed to measure exaggerated safety.
1 paper · 0 benchmarks
Hinglish-TOP is a human annotated code-switched semantic parsing dataset containing 10k human annotations for Hindi-English (HINGLISH) code switched utterances, and over 170K CST5 generated code-switched utterances from the TOPv2 dataset.
1 paper · 0 benchmarks
This dataset is composed of 7,753 pairs of whole slide images and their corresponding diagnostic reports, extracted from the TCGA platform and refined with large language models.
1 paper · 1 benchmark
Here the dataset described in Hitchhiking Rides Dataset: Two decades of crowd-sourced records on stochastic traveling(https://arxiv.org/abs/2506.21946) is published.
1 paper · 0 benchmarks
HoaxItaly consists of over 1 million tweets shared during 2019 and containing links to thousands of news articles published on two classes of Italian outlets: (1) disinformation websites, i.e.
1 paper · 0 benchmarks
This dataset is used for predicting house prices from both images and textual information.
1 paper · 0 benchmarks
HowSumm is a large-scale query-focused multi-document summarization dataset.
1 paper · 2 benchmarks
HpVaxFrames includes 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines.
1 paper · 0 benchmarks
A dataset including texts by humans (labeled 0) and then rephrased by ChatGPT (labeled 1), created to train models for machine-generated text detection.
1 paper · 0 benchmarks
HumanMT is a collection of human ratings and corrections of machine translations.
1 paper · 0 benchmarks
IAPR TC-12 (IAPR TC-12 Benchmark)
The image collection of the IAPR TC-12 Benchmark consists of 20,000 still natural images taken from locations around the world and comprising an assorted cross-section of still natural images.
1 paper · 0 benchmarks
This dataset contains general and named entities annotations on both clean written text and on noisy speech data.
1 paper · 0 benchmarks
ICLR Database (ICLR Database (with Textual Covariates))
A maintained database tracks ICLR submissions and reviews, augmented with author profiles and higher-level textual features.
1 paper · 0 benchmarks
Intelligent vehicle systems require a deep understanding of the interplay between road conditions, surrounding entities, and the ego vehicle's driving behavior for explainable driving decision-making and safe and efficient navigation.
1 paper · 0 benchmarks
IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.
1 paper · 0 benchmarks
IEE is a financial-domain dataset of the Insurance-entity extraction task.
1 paper · 0 benchmarks
ILIAS (ILIAS: Instance-Level Image retrieval At Scale)
ILIAS is a large-scale test dataset for evaluation on Instance-Level Image retrieval At Scale.
1 paper · 0 benchmarks
A collection of test sets for evaluating base and chat LLMs (incl.
1 paper · 0 benchmarks
IMPACT Patent (A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents)
It is a large-scale multimodal patent dataset with detailed captions for design patent figures.
1 paper · 1 benchmark
IMaSC (ICFOSS Malayalam Speech Corpus)
IMaSC is a Malayalam text and speech corpus made available by ICFOSS for the purpose of developing speech technology for Malayalam, particularly text-to-speech.
1 paper · 0 benchmarks
IRLCov19 is a multilingual Twitter dataset related to Covid-19 collected in the period between February 2020 to July 2020 specifically for regional languages in India.
1 paper · 0 benchmarks
ITCPR dataset (Image-Text Composed Person Retrieval dataset)
The ITCPR dataset is a comprehensive collection specifically designed for the Zero-Shot Composed Person Retrieval (ZS-CPR) task.
1 paper · 1 benchmark
IVM-Mix-1M provide over 1M image-instruction pairs with corresponding instruction-relevant mask labels.
1 paper · 0 benchmarks
Ice Hockey News Dataset is a corpus of Finnish ice hockey news, edited to be suitable for training of end-to-end news generation methods, as well as demonstrate generation of text, which was judged by journalists to be relatively close to…
1 paper · 0 benchmarks
IgboNLP is a standard machine translation benchmark dataset for Igbo.
1 paper · 0 benchmarks
IllusionAnimalstest Dataset Characteristics IllusionAnimalstest is a generated dataset based on a synthetic collection of animal images, including 10 animal classes: cat, dog, pigeon, butterfly, elephant, horse, deer, snake, fish, and…
1 paper · 0 benchmarks
IllusionChartest Dataset Characteristics IllusionChartest is a generated dataset containing 3,300 samples of images that feature sequences of 3 to 5 random characters.
1 paper · 0 benchmarks
IllusionFashionMNISTtest Dataset Characteristics IllusionFashionMNISTtest is a generated dataset derived from the FashionMNIST dataset.
1 paper · 0 benchmarks
IllusionMNISTtest Dataset Characteristics IllusionMNISTtest is a generated dataset derived from the MNIST dataset.
1 paper · 0 benchmarks
This dataset contains 6,387 ChatGPT prompts collected from four platforms (Reddit, Discord, websites, and open-source datasets) during Dec 2022 to May 2023.
1 paper · 0 benchmarks
Simulates unanticipated user needs in the deployment stage.
1 paper · 0 benchmarks
We present two multi-modal datasets, one for Main Board IPOs, and the other for Small and Medium Enterprises (SME) IPOs.
1 paper · 0 benchmarks
A collection of large languge model responses to tasks of propositional logic.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.