Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 42 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1969–2016 of 3,130
UCC (Unhealthy Comments Corpus)
The Unhealthy Comments Corpus (UCC) is corpus of 44355 comments intended to assist in research on identifying subtle attributes which contribute to unhealthy conversations online.
2 papers · 0 benchmarks
UK Biobank participants have generously provided a very wide range of information about their health and well-being since recruitment began in 2006.
2 papers · 1 benchmark
UPFD-POL (User Preference-aware Fake News Detection)
The PolitiFact variant of the UPFD dataset for benchmarking.
2 papers · 1 benchmark
UQuAD (Urdu Question Answering Dataset)
Large scale machine reading comprehension dataset in Urdu language.
2 papers · 1 benchmark
Underwater Trash Detection Dataset Overview The Underwater Trash Detection Dataset is a custom-annotated dataset designed to address the challenges of underwater trash detection caused by varying environmental features.
2 papers · 0 benchmarks
V2VBench is a comprehensive benchmark designed to evaluate video editing methods.
2 papers · 0 benchmarks
VGaokao is a verification style reading comprehension dataset designed for native speakers' evaluation.
2 papers · 0 benchmarks
The dataset, VIST-Edit, includes 14,905 human-edited versions of 2,981 machine-generated visual stories.
2 papers · 0 benchmarks
VQA 360° is a dataset for visual question answering on 360° images containing around 17,000 real-world image-question-answer triplets for a variety of question types.
2 papers · 0 benchmarks
VQDv1 (Visual Query Detection v1)
In Visual Query Detection (VQD), a system is given a query (prompt) natural language and an image, and then the system must produce 0 - N boxes that satisfy that query.
2 papers · 0 benchmarks
The dataset contains traffic traces collected from 3 different VR applications.
2 papers · 0 benchmarks
VTC (Videos, Titles and Comments)
VTC is a large-scale multimodal dataset containing video-caption pairs (~300k) alongside comments that can be used for multimodal representation learning.
2 papers · 0 benchmarks
Binary labels for Validity and Novelty respectively are given for each Conclusion.
2 papers · 1 benchmark
VerilogEval Dataset The VerilogEval Dataset is a benchmark specifically designed to assess the ability of large language models (LLMs) to generate syntactically correct and functionally accurate Verilog code.
2 papers · 1 benchmark
ViText2SQL is a dataset for the Vietnamese Text-to-SQL semantic parsing task, consisting of about 10K question and SQL query pairs.
2 papers · 0 benchmarks
WDC-PAVE (Web Data Commones - Product Attribute Value Extraction)
The datasets contains 1,420 human annotated product offers, systematically selected from the Web Data Commons Product Matching Corpus, featuring 24,582 annotated attribute-value pairs, making it a valuable resource for both product…
2 papers · 1 benchmark
WT-WT (Will-They-Won't-They)
Will-They-Won't-They (WT-WT) is a large dataset of English tweets targeted at stance detection for the rumor verification task.
2 papers · 0 benchmarks
This paper is a condensed report on the second year of the Touché shared task on argument retrieval held at CLEF 2021.
2 papers · 0 benchmarks
Weibo-COV is a large-scale COVID-19 social media dataset from Weibo, covering more than 30 million posts from 1 November 2019 to 30 April 2020.
2 papers · 0 benchmarks
WikiCaps is a large-scale multilingual but non-parallel data set for multimodal machine translation and retrieval.
2 papers · 0 benchmarks
WikiSRS is a novel dataset of similarity and relatedness judgments of paired Wikipedia entities (people, places, and organizations), as assigned by Amazon Mechanical Turk workers.
2 papers · 0 benchmarks
Wikidata-14M is a recommender system dataset for recommending items to Wikidata editors.
2 papers · 0 benchmarks
The Wikidata-Disamb dataset is intended to allow a clean and scalable evaluation of NED with Wikidata entries, and to be used as a reference in future research.
2 papers · 0 benchmarks
WildDESED (Wild Domestic Environment Sound Event Detection)
WildDESED is an extension of the original DESED dataset, created to reflect various domestic scenarios by incorporating complex and unpredictable background noises.
2 papers · 1 benchmark
X-WikiRE is a new, large-scale multilingual relation extraction dataset in which relation extraction is framed as a problem of reading comprehension to allow for generalization to unseen relations.
2 papers · 0 benchmarks
It consists of an extensive collection of a high quality cross-lingual fact-to-text dataset in 11 languages: Assamese (as), Bengali (bn), Gujarati (gu), Hindi (hi), Kannada (kn), Malayalam (ml), Marathi (mr), Oriya (or), Punjabi (pa),…
2 papers · 1 benchmark
XL-R2R (Cross-lingual Room-to-Room)
The XL-R2R dataset is built upon the R2R dataset and extends it with Chinese instructions.
2 papers · 0 benchmarks
Given a question and passage in an Indic language, generate a short answer span from the passage as the answer.
2 papers · 0 benchmarks
Given a question in an Indic language and a passage in English, generate a short answer span.
2 papers · 0 benchmarks
YASO is a crowd-sourced TSA evaluation dataset, collected using a new annotation scheme for labeling targets and their sentiments.
2 papers · 0 benchmarks
We present YTSeg, a topically and structurally diverse benchmark for the text segmentation task based on YouTube transcriptions.
2 papers · 1 benchmark
The YouTube8M-MusicTextClips dataset consists of over 4k high-quality human text descriptions of music found in video clips from the YouTube8M dataset.
2 papers · 0 benchmarks
arXivEdits an annotated corpus of 751 full papers from arXiv with gold sentence alignment across their multiple versions of revision, as well as fine-grained span-level edits and their underlying intentions for 1,000 sentence pairs.
2 papers · 0 benchmarks
kickstarter (Funding Successful Projects on Kickstarter)
Kickstarter is a community of more than 10 million people comprising of creative, tech enthusiasts who help in bringing creative project to life.
2 papers · 1 benchmark
Large language models such as ChatGPT and GPT-4 have recently achieved astonishing performance on a variety of natural language processing tasks.
2 papers · 0 benchmarks
A data set of Sudoku grids with more than one solution.
2 papers · 0 benchmarks
The pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language.
2 papers · 0 benchmarks
robo-vln (Robotics Vision-and-Language Navigation)
The Robo-VLN dataset is a continuous control formulation of the VLN-CE dataset by Krantz et al ported over from Room-to-Room (R2R) dataset created by Anderson et al.
2 papers · 1 benchmark
A set of 180,000 Sudoku grids with a variable number of hints from the minimal number of 17 (extremely hard instances) to 34 (easy instances), with 10,000 instances per level of hardness.
2 papers · 0 benchmarks
A set of easy Sudoku instances used in the SATNet paper for training SatNet on how to learn to play Sudoku.
2 papers · 0 benchmarks
satp-zsm-stage1 (Replication Data for: Crossing the Linguistic Causeway: A Binational Approach for Translating Soundscape Attributes to zsm)
This is the replication data for the paper: "Crossing the Linguistic Causeway: A Binational Approach for Translating Soundscape Attributes to Bahasa Melayu".
2 papers · 0 benchmarks
Clean version of UDHR (Universal Declaration of Human Rights), at the long sentence level.
2 papers · 0 benchmarks
chinahate dataset contains a total of 2,172,333 tweets hashtagged #china posted during the time it was collected.
1 paper · 0 benchmarks
Official dataset for Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models.
1 paper · 0 benchmarks
We provide here a new multi-view text dataset, collected from three well-known online news sources: BBC, Reuters, and The Guardian.
1 paper · 0 benchmarks
The dataset contains 60,000 Stack Overflow questions from 2016-2020, classified into three categories: 1.
1 paper · 1 benchmark
This paper constructs 7-digit product Supply-Use Tables (SUTs) and symmetric Input-Output Tables (IOTs) for the Indian economy using microdata from the Annual Survey of Industries (ASI) for the period 2016-2021.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.