Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 57 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 2689–2736 of 3,130
RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).
1 paper · 0 benchmarks
Multilingual explainable fact-checking dataset on Russia-Ukraine Conflict 2022
1 paper · 0 benchmarks
RUSS (Rapid Universal Support Service) is a dataset that consists of a collection of 741 real-world step-by-step natural language instructions (raw and annotated) from the open web, and for each: its corresponding webpage DOM, ground-truth…
1 paper · 0 benchmarks
RVL-CDIPMP is our first contribution to retrieve the original documents of the IIT-CDIP test collection which were used to create RVL-CDIP.
1 paper · 0 benchmarks
RVL-CDIPMP-N can serve its original goal as a covariate shift test set, now for multi-page document classification.
1 paper · 0 benchmarks
RaTE-NER dataset is a large-scale, radiological named entity recognition (NER) dataset, including 13,235 manually annotated sentences from 1,816 reports within the MIMIC-IV database, that spans 9 imaging modalities and 23 anatomical…
1 paper · 0 benchmarks
RadCases Dataset This HuggingFace (HF) dataset contains the raw case labels for input patient "one-liner" case summaries according to the ACR Appropriateness Criteria.
1 paper · 0 benchmarks
RadioTalk is a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019.
1 paper · 0 benchmarks
realfred is an embodied instruction following benchmark.
1 paper · 0 benchmarks
RealVul (RealVul-Vulnerability Dataset following realistic settings)
This is a C++ vulnerability detection dataset following realistic settings.
1 paper · 0 benchmarks
Reddit Engagement Dataset (RED), a distant-supervision set, with 80k single-turn conversations.
1 paper · 0 benchmarks
Dataset with articles posted in the r/Liberal and r/Conservative subreddits.
1 paper · 1 benchmark
This is a dataset of over 40K Reddit comments removed by moderators according to the specific type of macro norm being violated.
1 paper · 0 benchmarks
This dataset comprises 77,175 Reddit posts from 115 subreddit forums, annotated for the presence of 15 topics related to eating disorders and dieting.
1 paper · 0 benchmarks
https://arxiv.org/abs/2503.15222
1 paper · 0 benchmarks
Refer360° is a novel large-scale referring expression recognition dataset consisting of 17,137 instruction sequences and ground-truth actions for completing these instructions in 360° scenes.
1 paper · 0 benchmarks
Teaching assistants (TAs) are heavily used in computer science courses as a way to handle high enrollment and still being able to offer students individual tutoring and detailed assessments.
1 paper · 0 benchmarks
This is a dataset of regular expressions collected from regex101.com.
1 paper · 0 benchmarks
Dataset Details Total Labeled: 100% Labeled and Curated: 24,478 Pending: 0 Drafts: 0 Discarded: 696 High-Level Explanation This dataset includes labeled samples from the Colombian Aeronautical Regulations (RAC), covering all chapters…
1 paper · 0 benchmarks
This dataset is a collection of 5348 links from bug-introducing and bug-fixing commit sets extracted from Mozilla's Bugzilla with the use of bugbug.
1 paper · 0 benchmarks
The relational pattern similarity dataset is a new dataset upon the work of Zeichner et al.
1 paper · 0 benchmarks
RepLab 2013 dataset uses Twitter data in English and Spanish (more than 142,000 tweets).
1 paper · 0 benchmarks
This dataset contains the data used for all statistical analysis in our publication "Singapore Soundscape Site Selection Survey (S5): Identification of Characteristic Soundscapes of Singapore via Weighted k-means Clustering", summarised in…
1 paper · 0 benchmarks
This is the replication package for our systematic literature review and can be used for the reproducibility of the individual steps of our search and selection methodology.
1 paper · 0 benchmarks
ReviewRobot Dataset Overview This repository contains data for paper ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Ritter PoS (Ritter Twitter part-of-speech tagging)
PTB-tagged English Tweets
1 paper · 0 benchmarks
Robust Summarization Evaluation Benchmark is a large human evaluation dataset consisting of over 22k summary-level annotations over state-of-the-art systems on three datasets.
1 paper · 0 benchmarks
The Room environment - v0 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
The Room environment - v1 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
The Room environment - v2 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
RoomSpace: a new benchmark designed to evaluate language models on spatial reasoning tasks demanding spatial relation knowledge and multi-hop reasoning.
1 paper · 0 benchmarks
Fact-based Text Editing dataset based on RotoWire dataset
1 paper · 1 benchmark
The RotoWire-Modified dataset is a cleaned extension of the RotoWire dataset, with writer information about each document.
1 paper · 0 benchmarks
The work provides a comprehensive overview of the corpus for the Russian language for the commonsense inference task.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SAD-Instruct (Situational Awareness Database for Instruct-Tuning)
The Situational Awareness Database for Instruct-Tuning (SAD-Instruct) is a dataset for dynamic task guidance.
1 paper · 0 benchmarks
SAIL 2017 (Sentiment Analysis for Indian Languages)
India is a linguistic area with one of the longest histories of contact, influence, use, teaching and learning of English-in-diaspora in the world (Kachru and Nelson, 2006).
1 paper · 1 benchmark
SBU-WSD-Corpus is a corpus for Persian Word Sense Disambiguation (WSD).
1 paper · 0 benchmarks
SCI (Self-Contradictory Instructions)
Large multimodal models (LMMs) excel in adhering to human instructions.
1 paper · 0 benchmarks
SCIMAT is a large question-answer dataset for mathematics and science problems; such dataset can have impact on online education, intelligent tutoring and automated grading.
1 paper · 0 benchmarks
A Chinese sign language dataset that includes dialogue information.
1 paper · 0 benchmarks
SE-PEF (Stack Exchange - Personalized Expert Finding)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brisef description of the datdaset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SES (Spanish Emotional Speech)
Currently, an essential point in speech synthesis is the addressing of the variability of human speech.
1 paper · 0 benchmarks
Short-Films 20K (SF20K) is the largest publicly available movie dataset.
1 paper · 0 benchmarks
SG-NLG (Schema-Guided Natural Language Generation)
The SG-NLG dataset is a pre-processed version of the DSTC8 Schema-Guided Dialogue SGD dataset, designed specifically for data-to-text Natural Language Generation (NLG).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.