Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 57 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2689–2736 of 3,130

RRT (RuRSTreebank)
RST corpus for Russian.
1 paper · 0 benchmarks
RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).
1 paper · 0 benchmarks
Multilingual explainable fact-checking dataset on Russia-Ukraine Conflict 2022
1 paper · 0 benchmarks
RUSS (Rapid Universal Support Service) is a dataset that consists of a collection of 741 real-world step-by-step natural language instructions (raw and annotated) from the open web, and for each: its corresponding webpage DOM, ground-truth…
1 paper · 0 benchmarks
RVL-CDIP_MP (RVL-CDIP multi-page)
RVL-CDIPMP is our first contribution to retrieve the original documents of the IIT-CDIP test collection which were used to create RVL-CDIP.
1 paper · 0 benchmarks
RVL-CDIP_N_MP (RVL-CDIP-N multi-page)
RVL-CDIPMP-N can serve its original goal as a covariate shift test set, now for multi-page document classification.
1 paper · 0 benchmarks
RaTE-NER dataset is a large-scale, radiological named entity recognition (NER) dataset, including 13,235 manually annotated sentences from 1,816 reports within the MIMIC-IV database, that spans 9 imaging modalities and 23 anatomical…
1 paper · 0 benchmarks
RadCases Dataset This HuggingFace (HF) dataset contains the raw case labels for input patient "one-liner" case summaries according to the ACR Appropriateness Criteria.
1 paper · 0 benchmarks
RadioTalk is a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019.
1 paper · 0 benchmarks
realfred is an embodied instruction following benchmark.
1 paper · 0 benchmarks
RealVul (RealVul-Vulnerability Dataset following realistic settings)
This is a C++ vulnerability detection dataset following realistic settings.
1 paper · 0 benchmarks
Reddit Engagement Dataset (RED), a distant-supervision set, with 80k single-turn conversations.
1 paper · 0 benchmarks
Dataset with articles posted in the r/Liberal and r/Conservative subreddits.
1 paper · 1 benchmark
This is a dataset of over 40K Reddit comments removed by moderators according to the specific type of macro norm being violated.
1 paper · 0 benchmarks
Reddit Posts Related To Eating Disorders and Dieting (Topic Annotations on Reddit Posts from Eating Disorders and Dieting Forums by Human and LLMs)
This dataset comprises 77,175 Reddit posts from 115 subreddit forums, annotated for the presence of 15 topics related to eating disorders and dieting.
1 paper · 0 benchmarks
https://arxiv.org/abs/2503.15222
1 paper · 0 benchmarks
Refer360° is a novel large-scale referring expression recognition dataset consisting of 17,137 instruction sequences and ground-truth actions for completing these instructions in 360° scenes.
1 paper · 0 benchmarks
Teaching assistants (TAs) are heavily used in computer science courses as a way to handle high enrollment and still being able to offer students individual tutoring and detailed assessments.
1 paper · 0 benchmarks
This is a dataset of regular expressions collected from regex101.com.
1 paper · 0 benchmarks
Dataset Details Total Labeled: 100% Labeled and Curated: 24,478 Pending: 0 Drafts: 0 Discarded: 696 High-Level Explanation This dataset includes labeled samples from the Colombian Aeronautical Regulations (RAC), covering all chapters…
1 paper · 0 benchmarks
This dataset is a collection of 5348 links from bug-introducing and bug-fixing commit sets extracted from Mozilla's Bugzilla with the use of bugbug.
1 paper · 0 benchmarks
The relational pattern similarity dataset is a new dataset upon the work of Zeichner et al.
1 paper · 0 benchmarks
RepLab 2013 dataset uses Twitter data in English and Spanish (more than 142,000 tweets).
1 paper · 0 benchmarks
Replication Data for: Singapore Soundscape Site Selection Survey (S5) (Identification of Characteristic Soundscapes of Singapore via Weighted k-means Clustering)
This dataset contains the data used for all statistical analysis in our publication "Singapore Soundscape Site Selection Survey (S5): Identification of Characteristic Soundscapes of Singapore via Weighted k-means Clustering", summarised in…
1 paper · 0 benchmarks
This is the replication package for our systematic literature review and can be used for the reproducibility of the individual steps of our search and selection methodology.
1 paper · 0 benchmarks
ReviewRobot Dataset Overview This repository contains data for paper ReviewRobot: Explainable Paper Review Generation based on Knowledge Synthesis.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Ritter PoS (Ritter Twitter part-of-speech tagging)
PTB-tagged English Tweets
1 paper · 0 benchmarks
Robust Summarization Evaluation Benchmark is a large human evaluation dataset consisting of over 22k summary-level annotations over state-of-the-art systems on three datasets.
1 paper · 0 benchmarks
RoomEnv-v0 (The Room environment - v0)
The Room environment - v0 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
RoomEnv-v1 (The Room environment - v1)
The Room environment - v1 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
RoomEnv-v2 (The Room environment - v2)
The Room environment - v2 We have released a challenging Gymnasium compatible environment.
1 paper · 1 benchmark
RoomSpace: a new benchmark designed to evaluate language models on spatial reasoning tasks demanding spatial relation knowledge and multi-hop reasoning.
1 paper · 0 benchmarks
Fact-based Text Editing dataset based on RotoWire dataset
1 paper · 1 benchmark
The RotoWire-Modified dataset is a cleaned extension of the RotoWire dataset, with writer information about each document.
1 paper · 0 benchmarks
The work provides a comprehensive overview of the corpus for the Russian language for the commonsense inference task.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SAD-Instruct (Situational Awareness Database for Instruct-Tuning)
The Situational Awareness Database for Instruct-Tuning (SAD-Instruct) is a dataset for dynamic task guidance.
1 paper · 0 benchmarks
SAIL 2017 (Sentiment Analysis for Indian Languages)
India is a linguistic area with one of the longest histories of contact, influence, use, teaching and learning of English-in-diaspora in the world (Kachru and Nelson, 2006).
1 paper · 1 benchmark
SBU-WSD-Corpus is a corpus for Persian Word Sense Disambiguation (WSD).
1 paper · 0 benchmarks
SCI (Self-Contradictory Instructions)
Large multimodal models (LMMs) excel in adhering to human instructions.
1 paper · 0 benchmarks
SCIMAT is a large question-answer dataset for mathematics and science problems; such dataset can have impact on online education, intelligent tutoring and automated grading.
1 paper · 0 benchmarks
A Chinese sign language dataset that includes dialogue information.
1 paper · 0 benchmarks
SE-PEF (Stack Exchange - Personalized Expert Finding)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Click to add a brisef description of the datdaset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SES (Spanish Emotional Speech)
Currently, an essential point in speech synthesis is the addressing of the variability of human speech.
1 paper · 0 benchmarks
SF20K (Short-Films 20K)
Short-Films 20K (SF20K) is the largest publicly available movie dataset.
1 paper · 0 benchmarks
SG-NLG (Schema-Guided Natural Language Generation)
The SG-NLG dataset is a pre-processed version of the DSTC8 Schema-Guided Dialogue SGD dataset, designed specifically for data-to-text Natural Language Generation (NLG).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.