Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 56 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 2641–2688 of 3,130
PheMT is a phenomenon-wise dataset designed for evaluating the robustness of Japanese-English machine translation systems.
1 paper · 0 benchmarks
Phrase in Context is a curated benchmark for phrase understanding and semantic search, consisting of three tasks of increasing difficulty: Phrase Similarity (PS), Phrase Retrieval (PR) and Phrase Sense Disambiguation (PSD).
1 paper · 0 benchmarks
PhysNLU is a collection of 4 core datasets related to sentence classification, ordering, and coherence of physics explanations based on related tasks.
1 paper · 0 benchmarks
Physical concept understanding benchmark.
1 paper · 0 benchmarks
Pirá (Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean)
A large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English.
1 paper · 0 benchmarks
PlainFact is a high-quality human-annotated dataset with fine-grained explanation (i.e., added information) annotations.
1 paper · 0 benchmarks
An evaluation dataset for planning with LLM agents
1 paper · 0 benchmarks
Poker Hand Histories A collection of poker hand histories, covering 11 poker variants, in the poker hand history (PHH) format.
1 paper · 0 benchmarks
PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
1 paper · 0 benchmarks
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
1 paper · 0 benchmarks
PolyNews is a multilingual parallel dataset containing news titles 833 language pairs, spanning in 64 languages and 17 scripts.
1 paper · 0 benchmarks
A dataset for studying situated goal-directed human communication.
1 paper · 0 benchmarks
In this Pre-Contest Workshop Slidedeck.pdf: Instructional materials delivered for the seven pre-contest workshops
1 paper · 0 benchmarks
We release various types of word embeddings for multiple Indian languages.
1 paper · 0 benchmarks
PreRAID (Prescreening Rheumatoid Arthritis Information Database (PreRAID))
PreRAID is a structured dataset designed to evaluate the diagnostic capabilities of Large Language Models (LLMs) in Rheumatoid Arthritis (RA) diagnosis.
1 paper · 0 benchmarks
Dataset Description High-level explanation of dataset characteristics: This dataset includes electromyographic (EMG) signals captured using the BiTalino device.
1 paper · 0 benchmarks
Press Briefing Claim Dataset The dataset contains a total of 53 press briefings from a time span of over four years (2017-2021).
1 paper · 0 benchmarks
A large-scale dataset for proactive document retrieval that consists of over 2.8 million conversations from Reddit.
1 paper · 0 benchmarks
ProNCI consists of 22.5K proper noun compounds along with their free-form semantic interpretations.
1 paper · 0 benchmarks
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP), e.g.
1 paper · 1 benchmark
Processed Twitter is a dataset that is used for Twitter topic recognition.
1 paper · 0 benchmarks
Product Page is a large-scale and realistic dataset of webpages.
1 paper · 0 benchmarks
The corpus contains review sentences mostly of products in electronics domain, annotated and segregated into 4 comparison categories.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ProofNetVerif is an evaluation benchmark comprising 3,752 entries, each including an informal mathematical statement, its reference formalization, a predicted formalization, and a binary label indicating semantic equivalence.
1 paper · 0 benchmarks
OOD split of the Mol-Instructions Dataset about Protein Annotation.
1 paper · 0 benchmarks
This is a large-scale court judgment dataset, where each judgment is a summary of the case description with a patternized style.
1 paper · 0 benchmarks
PsOCR (Pashto OCR Dataset)
PsOCR is a large-scale synthetic dataset for Optical Character Recognition in low-resource Pashto language.
1 paper · 0 benchmarks
A.2.1 AN OPEN, LARGE-SCALE DATASET FOR ZERO-SHOT DRUG DISCOVERY DERIVED FROM PUBCHEM We constructed a large public dataset extracted from PubChem (Kim et al., 2019; Preuer et al., 2018), an open chemistry database, and the largest…
1 paper · 0 benchmarks
This dataset gathers 14,857 entities, 133 relations, and entities corresponding tokenized text from PubMed.
1 paper · 0 benchmarks
This dataset gathers three types of pairs: Title-to-Abstract (Training: 22,811/Development: 2095/Test: 2095), Abstract-to-Conclusion and Future work (Training: 22,811/Development: 2095/Test: 2095), Conclusion and Future work-to-Title…
1 paper · 0 benchmarks
PubMedQA-MetaGen: Metadata-Enriched PubMedQA Corpus Dataset Summary PubMedQA-MetaGen is a metadata-enriched version of the PubMedQA biomedical question-answering dataset, created using the MetaGenBlendedRAG enrichment pipeline.
1 paper · 2 benchmarks
QASiNa (Question Answering Sirah Nabawiyah)
Question Answering Sirah Nabawiyah (QASiNa) Dataset is a reading comprehension dataset consists of QA from Sirah Nabawiyah literature in Indonesian Language
1 paper · 0 benchmarks
QASports (A Question Answering Dataset about Sports)
Sport is one of the most popular and revenue-generating forms of entertainment.
1 paper · 0 benchmarks
A high-quality large-scale dataset consisting of 49,000+ data samples for the task of Chinese query-based document summarization.
1 paper · 0 benchmarks
The R1-Onevision dataset is a meticulously crafted resource designed to empower models with advanced multimodal reasoning capabilities.
1 paper · 0 benchmarks
REBUS (A Robust Evaluation Benchmark of Understanding Symbols)
Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input.
1 paper · 1 benchmark
REFCAT (Internet Archive Scholar Reference Dataset)
Internet Archive Scholar Reference Dataset.
1 paper · 0 benchmarks
Reader Emotion News 20k Dataset
1 paper · 0 benchmarks
RES-Q (RES-Q: Evaluating Code-Editing Large Language Model Systems at the Repository Scale)
RES-Q is a natural language instruction-based benchmark for evaluating Repository Editing Systems, which consists of 100 handcrafted repository editing tasks derived from real GitHub commits.
1 paper · 1 benchmark
The data used in - "Radio Galaxy Zoo EMU: Towards a Semantic Radio Galaxy Morphology Taxonomy" (Bowles et al.
1 paper · 0 benchmarks
The datasets of "Reinforcement Learning-enhanced Shared-account Cross-domain Sequential Recommendation" (TKDE 2022)
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
RLM25 (Research-Level Mathematics 2025)
RLM25 is an evaluation benchmark containing 619 paired examples of research-level natural language mathematical statements and their corresponding Lean formalizations.
1 paper · 0 benchmarks
ROAST (Review level Opinion Aspect Sentiment Target Joint Detection for ABSA)
This repository has a review-level multidomain multilingual dataset for Aspect-based Sentiment Analysis(ABSA) for the paper ROAST: Review-level Opinion Aspect Sentiment Target Joint Detection.
1 paper · 0 benchmarks
RPCD (Reddit Photo Critique Dataset)
The Reddit Photo Critique Dataset (RPCD) contains tuples of image and photo critiques.
1 paper · 0 benchmarks
RPEval (Role-Playing Evaluation Dataset)
Role-Playing Eval (RPEval) is a benchmark dataset designed to evaluate large language models' role-playing abilities across emotional understanding, decision-making, moral alignment, and in-character consistency.
1 paper · 0 benchmarks
RRG (Russian RST dataset from GUM v9.1 corpus)
Parallel version of annotations in GUM RST v9.1.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.