Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 61 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2881–2928 of 3,130

USCOCO (Unexpected Situations of Common Objects in Context)
A test set of grammatically correct sentences and layouts (visual “imagined” situations), called Unexpected Situations of Common Objects in Context (USCOCO) describing compositions of entities and relations that are unlikely to be found in…
1 paper · 0 benchmarks
The UTRSet-Real dataset is a comprehensive, manually annotated dataset specifically curated for Printed Urdu OCR research.
1 paper · 0 benchmarks
The UTRSet-Synth dataset is introduced as a complementary training resource to the UTRSet-Real Dataset, specifically designed to enhance the effectiveness of Urdu OCR models.
1 paper · 0 benchmarks
The Ubuntu Chat Corpus (UCC) is composed of archived chat logs from Ubuntu's Internet Relay Chat technical support channels.
1 paper · 0 benchmarks
Urban Dict spelling variant is a variant spelling dataset for use of NLP research in the informal domain.
1 paper · 0 benchmarks
This dataset is the translation of the MS-marco dataset, marking it the first large-scale urdu IR dataset.
1 paper · 0 benchmarks
Urdu News Headlines Dataset with VOA and BBC An Urdu news headlines dataset is a collection of news headlines in the Urdu language, typically scraped from news websites and social media platforms.
1 paper · 1 benchmark
The UrduDoc Dataset is a benchmark dataset for Urdu text line detection in scanned documents.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This dataset includes User Story (or Issue) text descriptions, User Story titles, and Story Points from 33 software development projects, comprising a total of 20,479 User Stories (or issues) extracted from GitLab repositories, amounting…
1 paper · 0 benchmarks
V3C1 (the Vimeo Creative Commons Collection 1)
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, and will serve as evaluation basis for the Video Browser Showdown 2019-2021 and TREC Video Retrieval…
1 paper · 0 benchmarks
VCG+112K (Video Instruction Dataset 112K)
Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs.
1 paper · 0 benchmarks
This task stems from the observation that text embedded in images is intrinsically different from common visual elements and natural language due to the need to align the modalities of vision, text, and text embedded in images.
1 paper · 0 benchmarks
VD-Ref is a dataset with ground-truth mappings from both noun phrases and pronouns to image regions.
1 paper · 0 benchmarks
VDQG (Visual Discriminative Question Generation)
The Visual Discriminative Question Generation (VDQG) dataset contains 11202 ambiguous image pairs collected from Visual Genome.
1 paper · 0 benchmarks
VMD (Virtual Moderation Dataset)
This dataset contains synthetically generated discussions and annotations using exclusively Large Language Model (LLM) agents.
1 paper · 0 benchmarks
The benchmark for VPData, the largest video inpainting dataset, which comprises over 390K clips (> 866.7 hours) and features precise masks and detailed video captions.
1 paper · 0 benchmarks
The largest video inpainting dataset comprises over 390K clips (> 866.7 hours), featuring precise masks and detailed video captions.
1 paper · 0 benchmarks
VQA-MHUG is a 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker.
1 paper · 0 benchmarks
VSTaR-1M is a 1M instruction tuning dataset, created using Video-STaR, with the source datasets: Kinetics700 STAR-benchmark FineDiving The videos for VSTaR-1M can be found in the links above.
1 paper · 0 benchmarks
VTQA (Visual Text Question Answering)
VTQA is a dataset containing open-ended questions about image-text pairs.
1 paper · 0 benchmarks
Combines CoVaxFrames and HpVaxFrames into a unified dataset of 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines and 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines.
1 paper · 0 benchmarks
A Natural Language Resource for Learning to Recognize Misinformation about the COVID-19 and HPV Vaccines.
1 paper · 0 benchmarks
Validity and Novelty are determined in a comparative setting between two conclusions at a time.
1 paper · 1 benchmark
ValiMath is a high-quality benchmark consisting of 2,147 carefully curated mathematical questions designed to evaluate an LLM's ability to verify the correctness of math questions based on multiple logic-based and structural criteria.
1 paper · 0 benchmarks
The Vashantor dataset consists of 32,500 sentences from different regions, including Chittagong, Noakhali, Sylhet, Barishal, and Mymensingh.
1 paper · 0 benchmarks
VedantaNY-10M is a curated dataset of over 750 hours of transcripts from public discourses on the Indian philosophy of Advaita Vedanta.
1 paper · 0 benchmarks
VerbCL is a dataset that consists of the citation graph of court opinions, which cite previously published court opinions in support of their arguments.
1 paper · 0 benchmarks
Verifee is a dataset of news articles with fine-grained trustworthiness annotations.
1 paper · 0 benchmarks
Verified Smart Contracts Code Comments is a dataset of real Ethereum smart contract functions, containing "code, comment" pairs of both Solidity and Vyper source code.
1 paper · 1 benchmark
Verified Smart Contracts is a dataset of real Ethereum smart contracts, containing both Solidity and Vyper source code.
1 paper · 0 benchmarks
Verireason-RTL-Coder_7b_reasoning_tb (VeriReason Verilog Dataset with Reasoning, Testbench, and Simulation Results)
Verireason-RTL-Coder7breasoningtb For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update Log…
1 paper · 0 benchmarks
Verireason-RTL-Coder_7b_reasoning_tb_simple (Simple Problems of VeriReason Verilog Dataset with Reasoning, Testbench, and Simulation Results)
Verireason-RTL-Coder7breasoningtbsimple For implementation details, visit our GitHub repository: VeriReason Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation Update…
1 paper · 0 benchmarks
100 videos with varying danger levels (on a scale of 0-10) and different scenarios, annotated by 18 human annotators using our annotation pipeline to represent human perception and respective Vision Language model summaries for each of the…
1 paper · 0 benchmarks
ViLCo (ViLCo-Bench)
We propose the first standardized benchmark in multimodal continual learning for video data, defining protocols for training and metrics for evaluation.
1 paper · 0 benchmarks
ViMATH (Vietnamese MATH)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ViSP (Vietnamese Sentence Paraphrases)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ViSR (Vietnamese Synthetic Reasoning)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ViTHSD (Vietnamese Targeted-Hate-Speech-Detection)
A Vietnamese dataset for hate speech detection by the specific target.
1 paper · 0 benchmarks
This is the reference headset microphone variant of the VibraVox dataset.
1 paper · 2 benchmarks
This is the in-ear rigid earpiece-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the in-ear comply foam-embedded microphone variant of the VibraVox dataset.
1 paper · 3 benchmarks
This is the throat microphone (laryngophone) variant of the VibraVox dataset.
1 paper · 3 benchmarks
Vid2RealHRI online video and results dataset (Community embedded robotics: Vid2RealHRI online video and perceived social intelligence in human-robot encounters dataset)
Introduction This dataset was gathered during the Vid2RealHRI study of humans’ perception of robots' intelligence in the context of an incidental Human-Robot encounter.
1 paper · 0 benchmarks
Dataset Introduction This dataset leverages VideoDB's Public Collection to offer a diverse range of videos featuring text-containing scenes.
1 paper · 1 benchmark
Spoken Named Entity Recognition (NER) aims to extracting named entities from speech and categorizing them into types like person, location, organization, etc.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.