Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 62 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 2929–2976 of 3,130
In doctor-patient conversations, identifying medically relevant information is crucial, posing the need for conversation summarization.
1 paper · 0 benchmarks
Molecules represent tokens of the language of chemistry, which underlies not only chemistry itself, but also scientific fields that use chemical information such as pharmacy, material science, and molecular biology.
1 paper · 0 benchmarks
VisArgs is a densely annotated benchmark for visual argument understanding.
1 paper · 0 benchmarks
VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents.
1 paper · 0 benchmarks
VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents.
1 paper · 0 benchmarks
Dataset for testing the ability of Vision Language Models (LVM) to recognize and match 3D objects of the exact same 3D shapes but with different orientation/materials/textures/ environments and light conditions.
1 paper · 0 benchmarks
Visual Haystacks (VHs) is a "visual-centric" Needle-In-A-Haystack (NIAH) benchmark specifically designed to evaluate the capabilities of Large Multimodal Models (LMMs) in visual retrieval and reasoning over sets of unrelated images.
1 paper · 0 benchmarks
We introduce a first Vietnamese Spelling Correction dataset containing manual labelling mistakes and corresponding correct words.
1 paper · 0 benchmarks
VlogQA (Vietnamese Spoken-Based Machine Reading Comprehension)
The VlogQA consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube - an extensive source of user-uploaded content, covering the topics of food and travel in the Vietnamese language.
1 paper · 0 benchmarks
Votranh DREAMLOG is a poetic, philosophical dataset generated by the self-evolving AI system Votranh V8.
1 paper · 0 benchmarks
VulScribeR (VulScriber: 22K+ unfiltered vul samples generated with ChatGPT via Injection)
Datasets are listed in the repository's readme file.
1 paper · 1 benchmark
The dataset consists of two versions: X1 with P3 and X1 without P3, where P3 represents a set of random unchanged functions from vulnerability fixing commits.
1 paper · 1 benchmark
Vulnerable Verified Smart Contracts is a dataset of real vulnerable Ethereum smart contracts.
1 paper · 0 benchmarks
This Sanskrit speech corpus has more than 78 hours of audio data and contains recordings of 45,953 sentences with a sampling rate of 22KHz.
1 paper · 0 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks
The WEB-FORUM-52 gold standard comprises (i) 13 web forums from the health domain, (ii) 15 forums obtained from a Wikipedia list of popular forums (https://en.wikipedia.org/wiki/ListofInternetforums), (iii) 13 forums mentioned on a list of…
1 paper · 0 benchmarks
WIKIOG is a public collection which consists of over 1.75 million document-outline pairs for research on the OG task.
1 paper · 0 benchmarks
The Medical Translation Task of WMT 2014 addresses the problem of domain-specific and genre-specific machine translation.
1 paper · 0 benchmarks
News translation is a recurring WMT task.
1 paper · 0 benchmarks
The Biomedical Translation Shared Task was first introduced at the First Conference of Machine Translation.
1 paper · 0 benchmarks
The IT Translation Task is a shared task introduced in the First Conference on Machine Translation.
1 paper · 0 benchmarks
We provide separate training, development and test data.
1 paper · 0 benchmarks
The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure.
1 paper · 0 benchmarks
The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering.
1 paper · 0 benchmarks
It contains data from two different realities: Food.com, a well-known American recipe site, and Planeat, an Italian site that allows you to plan recipes to save food waste.
1 paper · 0 benchmarks
Test-driven benchmark to challenge LLMs to write long JavaScript React application GitHub Script
1 paper · 1 benchmark
Fact-based Text Editing dataset based on WebNLG dataset.
1 paper · 1 benchmark
WebGen-Bench WebGen-Bench is created to benchmark LLM-based agent's ability to generate websites from scratch.
1 paper · 0 benchmarks
A dataset automatically generated using question generation neural models and alt-text video captions from the WebVid dataset, with 3M video-question-answer triplets.
1 paper · 0 benchmarks
Webis-ConcluGen-21 is a large-scale corpus of 136,996 samples of argumentative texts and their conclusions used for the task of generating informative conclusions.
1 paper · 0 benchmarks
This corpus contains preprocessed posts from the Reddit dataset, suitable for abstractive summarization using deep learning.
1 paper · 0 benchmarks
This dataset is used for user identity linkage across two online social networks in Chinese.
1 paper · 0 benchmarks
WhenAct (Temporal Human Action Localization in Lifestyle Vlogs)
We consider the task of temporal human action localization in lifestyle vlogs.
1 paper · 0 benchmarks
The Wiki-Flick Event dataset for cross-modal event retrieval is a well-labelled but weakly-aligned dataset collected for cross-modality event retrieval.
1 paper · 0 benchmarks
Wiki-Reliability is the first dataset of English Wikipedia articles annotated with a wide set of content reliability issues.
1 paper · 0 benchmarks
Wiki-en is an annotated English dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
Wiki-zh is an annotated Chinese dataset for domain detection extracted from Wikipedia.
1 paper · 0 benchmarks
A dataset comprising 8,551 ban evasion pairs on Wikipedia, where each pair comprises a parent account and the child account.
1 paper · 0 benchmarks
WikiBioCTE is a dataset for controllable text edition based on the existing dataset WikiBio (originally created for table-to-text generation).
1 paper · 0 benchmarks
WikiContradiction is a novel wiki dataset for self-contradiction Wikipedia article detection.
1 paper · 0 benchmarks
WikiFactDiff is a dataset designed as a resource to perform atomic factual knowledge updates on language models, with the goal of aligning them with current knowledge.
1 paper · 0 benchmarks
a high-level explanation of the dataset characteristics We introduce WikiOFGraph, a novel large-scale, domain-diverse dataset synthesized by LLMs, ensuring superior graph-text consistency to advance general-domain graph-to-text generation.
1 paper · 1 benchmark
WikiPII, an automatically labeled dataset composed of Wikipedia biography pages, annotated for personal information extraction.
1 paper · 0 benchmarks
To collect WikiSuggest, Google Suggest API is used to harvest natural language questions and submit them to Google Search.
1 paper · 0 benchmarks
Wikipedia Webpage 2M (WikiWeb2M) is a multimodal open source dataset consisting of over 2 million English Wikipedia articles.
1 paper · 0 benchmarks
WildQA is a video understanding dataset of videos recorded in outside settings.
1 paper · 1 benchmark
Test set of sentences in Hindi with complex coreference involving two entities inspired by WinoBias format of sentences in English.
1 paper · 0 benchmarks
WinoPron is a novel dataset of Winogender-like template pairs in English, which fixes inconsistencies in Winogender Schemas and contains balanced template pairs for pronoun forms in 3 grammatical cases, which we find impacts performance…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.