Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 40 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1873–1920 of 3,130

OCR-IDL (OCR Annotations for Industry Document Library Dataset)
The OCR-IDL dataset comprises the OCR annotations for a subset of 26M pages of the large-scale IDL document library.
2 papers · 0 benchmarks
OIR is a financial-domain dataset of the outbound intent recognition task.
2 papers · 0 benchmarks
ORCAS-I (Queries Annotated with Intent using Weak Supervision)
A labelled version of the ORCAS click-based dataset of Web queries, which provides 18 million connections to 10 million distinct queries.
2 papers · 1 benchmark
This is a transnational data set which contains all the transactions occurring between 01/12/2010 and 09/12/2011 for a UK-based and registered non-store online retail.
2 papers · 0 benchmarks
OpenAsp Dataset OpenAsp is an Open Aspect-based Multi-Document Summarization dataset derived from DUC and MultiNews summarization datasets.
2 papers · 0 benchmarks
OpenSLR (Open Speech and Language Resources)
OpenSLR is a repository of open speech and language resources, including large-scale transcribed audio corpora and related software.
2 papers · 1 benchmark
OpenViDial 2.0 is a larger-scale open-domain multi-modal dialogue dataset compared to the previous version OpenViDial 1.0.
2 papers · 1 benchmark
The dataset contains 45 documents containing narrative description of business process and their annotations.
2 papers · 0 benchmarks
PETCI (PETCI: A Parallel English Translation Dataset of Chinese Idioms)
PETCI is a Parallel English Translation dataset of Chinese Idioms, collected from an idiom dictionary and Google and DeepL translation.
2 papers · 0 benchmarks
PIPPA (Personal Interaction Pairs between People and AI) is a partially-synthetic dataset.
2 papers · 0 benchmarks
The dataset used in the experiments on the paper "Modeling citation worthiness by using attention‑based bidirectional long short‑term memory networks and interpretable models" There are one million sentences in total, and further splitted…
2 papers · 0 benchmarks
POINTREC is a test collection for point of interest (POI) recommendation, comprising of (i) a set of information needs, (ii) a dataset of POIs, and (iii) graded relevance assessments for information need and POI pairs.
2 papers · 0 benchmarks
Dataset Card for The Cancer Genome Atlas (TCGA) Multimodal Dataset The Cancer Genome Atlas (TCGA) Multimodal Dataset is a comprehensive collection of clinical data, pathology reports, molecular, and slide images for cancer patients.
2 papers · 0 benchmarks
Paper2Fig100k is a dataset with over 100k images of figures and texts from research papers.
2 papers · 0 benchmarks
ParaMAWPS (Paraphrased Math Word Problem Solving Repository)
This repository contains the code, data, and models of the paper titled "Math Word Problem Solving by Generating Linguistic Variants of Problem Statements" published in the Proceedings of the 61st Annual Meeting of the Association for…
2 papers · 1 benchmark
PcMSP is a dataset annotated from 305 open access scientific articles for material science information extraction that simultaneously contains the synthesis sentences extracted from the experimental paragraphs, as well as the entity…
2 papers · 0 benchmarks
PeerSum is a new MDS dataset using peer reviews of scientific publications.
2 papers · 0 benchmarks
PerCQA is the first Persian dataset for CQA (Community Question Answering).
2 papers · 0 benchmarks
The PATIS is a Persian language dataset for intent detection and slot filling.
2 papers · 2 benchmarks
An annotated dataset of 38,800 phishing and benign websites.
2 papers · 0 benchmarks
145k natural language and PDDL problem pairs from the Blocks World, Gripper, and Floor Tile domains.
2 papers · 0 benchmarks
PoKi is a corpus of 61,330 poems written by children from grades 1 to 12.
2 papers · 0 benchmarks
PoliteRewrite (the politerewrite dataset)
https://huggingface.co/datasets/jdustinwind/Polite
2 papers · 0 benchmarks
ProSLU (Profile-based Spoken Language Understanding)
In the paper, to bridge the research gap, we propose a new and important task, Profile-based Spoken Language Understanding (ProSLU), which requires a model not only depends on the text but also on the given supporting profile information.
2 papers · 2 benchmarks
PropSegmEnt is a corpus of over 35K propositions annotated by expert human raters.
2 papers · 0 benchmarks
A collection of 385,705 scientific abstracts about Cognitive Control and their GPT-3 embeddings.
2 papers · 1 benchmark
PICO is a framework to formulate a well-defined focused clinical question.
2 papers · 0 benchmarks
The Pump and dump dataset is an annotated set of messages to detect cryptocurrency market manipulations.
2 papers · 0 benchmarks
Python Programming Puzzles (P3) is an open-source dataset where each puzzle is defined by a short Python program , and the goal is to find an input which makes output "True".
2 papers · 0 benchmarks
The QTUNA dataset is the result of a series of elicitation experiments in which human speakers were asked to perform a linguistic task that invites the use of quantified expressions in order to inform possible Natural Language Generation…
2 papers · 0 benchmarks
QUITE (Quantifying Uncertainty in natural language Text)
QUITE (Quantifying Uncertainty in natural language Text) is an entirely new benchmark that allows for assessing the capabilities of neural language model-based systems w.r.t.
2 papers · 0 benchmarks
Introduction Generalized quantifiers (e.g., few, most) are used to indicate the proportions predicates are satisfied.
2 papers · 0 benchmarks
R2VQ (Recipe-to-Video Questions)
R2VQ is a dataset designed for testing competence-based comprehension of machines over a multimodal recipe collection, which contains text-video aligned recipes.
2 papers · 0 benchmarks
The first benchmark comprising 473 prompts designed to assess the ability of LLMs to resist malicious code generation.
2 papers · 0 benchmarks
ROPE (Recognition-based Object Probing Evaluation)
We introduce Recognition-based Object Probing Evaluation (ROPE), an automated evaluation protocol that considers the distribution of object classes within a single image during testing and uses visual referring prompts to eliminate…
2 papers · 0 benchmarks
RTC (Reddit Time Corpus)
RTC is a benchmark corpus of social media comments sampled over three years.
2 papers · 0 benchmarks
RUFF is a large-scale dataset to measure pronoun fidelity in English.
2 papers · 0 benchmarks
RUSLAN is a Russian spoken language corpus for text-to-speech task.
2 papers · 0 benchmarks
Rare Diseases Mentions in MIMIC-III (Rare disease mention annotations from a sample of MIMIC-III clinical notes)
Data annotation The 1,073 full rare disease mention annotations (from 312 MIMIC-III discharge summaries) are in fullsetRDannMIMICIIIdisch.csv.
2 papers · 1 benchmark
ReactionGIF is an affective dataset of 30K tweets which can be used for tasks like induced sentiment prediction and multilabel classification of induced emotions.
2 papers · 0 benchmarks
Articles originating from subreddits with explicitly stated ideologies are categorized into three groups: 72,488 articles in the Liberal class, 79,573 articles in the Conservative class, and 225,083 articles in the Restricted class.
2 papers · 1 benchmark
The Rendered SST2 dataset is a dataset released by OpenAI, that measures the optical character recognition capability of visual representations.
2 papers · 1 benchmark
ReviewQA is a question-answering dataset based on hotel reviews.
2 papers · 0 benchmarks
We collect a dataset of Rich Human Feedback on 18K images (RichHF-18K), which contains (i) point annotations on the image that highlight regions of implausibility/artifacts, and text-image misalignment; (ii) labeled words on the prompts…
2 papers · 0 benchmarks
RoMQA is a benchmark for robust, multi-evidence, and multi-answer question answering (QA).
2 papers · 0 benchmarks
RuMedBench is a benchmark dataset for Russian medical language understanding.
2 papers · 0 benchmarks
RuOpenBookQA is a QA dataset with multiple-choice elementary-level science questions which probe the understanding of core science facts.
2 papers · 1 benchmark
RuSentNE (RuSentNE-2023)
https://github.com/dialogue-evaluation/RuSentNE-evaluation
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.