Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 55 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2593–2640 of 3,130

The dataset identifies the shortcomings of existing benchmarks in evaluating the problem of compositional generalization, which underscores the need for the development of datasets tailored to assess compositional generalization in open…
1 paper · 1 benchmark
Dataset OQRanD and OQGenD for paper "Asking the crowd: Asking the Crowd: Question Analysis, Evaluation and Generation for Open Discussion on Online Forums" by Zi Chai, Xinyu Xing, Xiaojun Wan and Bo Huang.
1 paper · 0 benchmarks
OTSC-Hindi (Occupation Test Set with Simple Context - Hindi)
Test set of sentences in Hindi with simple gender-specific context used to measure gender bias in NMT systems for Hindi-English.
1 paper · 0 benchmarks
This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents.
1 paper · 0 benchmarks
OllaBench v.0.2 (OllaBench for Interdependent Cybersecurity v.0.2)
Large Language Models (LLMs) have the potential to enhance Agent-Based Modeling by better representing complex interdependent cybersecurity systems, improving cybersecurity threat modeling and risk management.
1 paper · 0 benchmarks
Olympic 2024 is a human-annotated dataset that contains 220 high-quality instance.
1 paper · 0 benchmarks
To effectively evaluate OmniCount across open-vocabulary, supervised, and few-shot counting tasks, a dataset catering to a broad spectrum of visual categories and instances featuring various visual categories with multiple instances and…
1 paper · 2 benchmarks
The OnlySports Dataset is a comprehensive collection of sports-related text data, comprising approximately 600 billion tokens.
1 paper · 0 benchmarks
A human-refined dataset of OpenAPI definitions based on the APIs.guru OpenAPI directory.
1 paper · 1 benchmark
OpenD5 is a a meta-dataset which aggregates 675 open-ended problems ranging across business, social sciences, humanities, machine learning, and health, and uses a set of unified evaluation metrics: validity, relevance, novelty, and…
1 paper · 0 benchmarks
We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community.
1 paper · 0 benchmarks
We create the first open-source large-scale S2V generation dataset OpenS2V-5M, which consists of five million high-quality 720P subject-text-video triples.
1 paper · 1 benchmark
OrdinalDataset (Ordinal Encoding Data set)
It includes 10 data sets that consists of both raw data set and encoded data set where it is encoded through BERT-Sort Encoder with MLM initialization of .
1 paper · 1 benchmark
P4D prompts (P4D universal jailbreaking prompt for T2I models)
This dataset contains prompts designed to evaluate and challenge the safety mechanisms of generative text-to-image models, with a particular focus on identifying prompts that are likely to produce images containing nudity.
1 paper · 0 benchmarks
The PART-OF dataset is a dataset of relations extracted from a medical ontology.
1 paper · 0 benchmarks
PECC (PECC: Problem Extraction and Coding Challenges)
Recent advancements in large language models (LLMs) have showcased their exceptional abilities across various tasks, such as code generation, problem-solving and reasoning.
1 paper · 1 benchmark
PGDataset (Profile Generation Dataset)
PGDataset (Profile Generation Dataset) is a dataset created for the PGTask (Profile Generation Task), where the goal is to extract/generate a profile sentence given a dialogue utterance.
1 paper · 1 benchmark
PIAST (PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We assembled a benchmark of electronic component pinouts, PINS100, containing 100 common parts frequently used in circuits found on high-traffic electronic tutorial websites such as the ARDUINO PROJECT HUB and AUTODESK TINKERCAD CIRCUITS.
1 paper · 0 benchmarks
PIZZA is a dataset for parsing pizza and drink orders, whose semantics cannot be captured by flat slots and intents.
1 paper · 0 benchmarks
PLOD-filtered (PLOD: An Abbreviation Detection Dataset for Scientific Documents)
PLOD: An Abbreviation Detection Dataset This is the PLOD (filtered) Dataset published at LREC 2022.
1 paper · 0 benchmarks
PLOD-unfiltered (PLOD: An Abbreviation Detection Dataset for Scientific Documents)
PLOD: An Abbreviation Detection Dataset This is the PLOD (unfiltered) Dataset published at LREC 2022.
1 paper · 0 benchmarks
PMC-SA (PMC Structured Abstracts)
PMC-SA (PMC Structured Abstracts) is a dataset of academic publications, used for the task of structured summarization.
1 paper · 0 benchmarks
PMPC (Persona Match on Persona-Chat)
PMPC (Persona Match on Persona-Chat) is a dataset for Speaker Persona Detection (SPD) which aims to detect speaker personas based on the plain conversational text.
1 paper · 0 benchmarks
The LiT.RL POLIT-FALSE-n-LEGIT NEWS DB 2016-2017 contains a total of 274 news articles about U.S.
1 paper · 0 benchmarks
PQ-decaNLP (Paraphrase Questions - decaNLP)
Multitask learning has led to significant advances in Natural Language Processing, including the decaNLP benchmark where question answering is used to frame 10 natural language understanding tasks in a single model.
1 paper · 0 benchmarks
PQAref (Pubmed Question Answering with references)
The PQAref dataset is a dataset for fine-tuning large language models for referenced question-answering in biomedical domain.
1 paper · 0 benchmarks
This is the official dataset for PRMBench.
1 paper · 0 benchmarks
PSM is a financial-domain dataset of the pairwise search matching task.
1 paper · 0 benchmarks
PTVD is a plot-oriented multimodal dataset in the TV domain.
1 paper · 0 benchmarks
We introduce a framework for benchmarking multi-step retrosynthesis methods, i.e.
1 paper · 0 benchmarks
PaSa is a dataset to train Machine Learning algorithms to automate the highlighting of patent paragraphs with semantic annotations.
1 paper · 0 benchmarks
Pan+ChiPhoto dataset is a Chinese character dataset.
1 paper · 0 benchmarks
Paper Field is built from the Microsoft Academic Graph and maps paper titles to one of 7 fields of study.
1 paper · 1 benchmark
We have prepared a dataset, ParagraphOrdreing, which consists of around 300,000 paragraph pairs.
1 paper · 0 benchmarks
Despite recent advances in vision-and-language tasks, most progress is still focused on resource-rich languages such as English.
1 paper · 0 benchmarks
The Part-Whole Relations dataset is a dataset of semantic relations between entities.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We address the computer-assisted search for prior art by creating a training dataset for supervised machine learning called PatentMatch.
1 paper · 0 benchmarks
PatternCom is a composed image retrieval benchmark based on PatternNet.
1 paper · 1 benchmark
Peer to Peer Hate is a comprehensive hate speech dataset capturing various types of hate.
1 paper · 0 benchmarks
Pentachromatic Cultural Palette Dataset is characterized by unique cultural semantics and values.
1 paper · 0 benchmarks
PIE stands for Performance Improving Code Edits.
1 paper · 0 benchmarks
The Perfume Co-Preference Network dataset comprises comprehensive user reviews and ratings collected from the Persian retail platform Atrafshan.
1 paper · 0 benchmarks
Perla Dataset (Perla Depression Screening Dataset)
This dataset contains the results of a depression screening experiment using two instruments: The PHQ-9 depression screening questionnaire and the chabot Perla.
1 paper · 0 benchmarks
The Permuted bAbi dialog task is an adaptation of the "Dialog bAbI tasks data" dataset released by Facebook.
1 paper · 0 benchmarks
PersianQA (Persian Question Answering Dataset)
PersianQA: a dataset for Persian Question Answering Persian Question Answering (PersianQA) Dataset is a reading comprehension dataset on Persian Wikipedia.
1 paper · 0 benchmarks
The PEDC is a corpus of 14 episodes of This American Life podcast transcripts that have been annotated for events.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.