Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 30 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1393–1440 of 3,130

We generate epistemic reasoning problems using modal logic to target theory of mind (tom) in natural language processing models.
4 papers · 0 benchmarks
MixSet (Mixcase Dataset)
MIXSET comprises a total of 3.6k mixtext instances.
4 papers · 1 benchmark
MuMu is a new dataset of more than 31k albums classified into 250 genre classes.
4 papers · 0 benchmarks
For each dataset we provide a short description as well as some characterization metrics.
4 papers · 0 benchmarks
MultiSense is a dataset of 9,504 images annotated with an English verb and its translation in Spanish and German.
4 papers · 0 benchmarks
MultiSpider is a large multilingual text-to-SQL dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese).
4 papers · 0 benchmarks
MultiSubs (MultiSubs: A Large-scale Multimodal and Multilingual Dataset)
MultiSubs is a dataset of multilingual subtitles gathered from the OPUS OpenSubtitles dataset, which in turn was sourced from opensubtitles.org.
4 papers · 5 benchmarks
MuseASTE (MuSe-CarASTE: A comprehensive dataset for aspect sentiment triplet extraction in automotive review videos)
•A new benchmark dataset for Aspect Sentiment Triplet Extraction.
4 papers · 1 benchmark
NELA-GT-2020 is an updated version of the NELA-GT-2019 dataset.
4 papers · 0 benchmarks
OASum is a large-scale open-domain aspect-based summarization dataset which contains more than 3.7 million instances with around 1 million different aspects on 2 million Wikipedia pages.
4 papers · 0 benchmarks
OLPBENCH is a large Open Link Prediction benchmark, which was derived from the state-of-the-art Open Information Extraction corpus OPIEC (Gashteovski et al., 2019).
4 papers · 0 benchmarks
OVDEval includes 9 sub-tasks and introduces evaluations on commonsense knowledge, attribute understanding, position understanding, object relation comprehension, and more.
4 papers · 0 benchmarks
OVQA contains 19,020 medical visual question and answer pairs generated from 2,001 medical images collected from 2,212 EMRs in Orthopedics.
4 papers · 0 benchmarks
OpenViVQA (Open-domain Visual Question Answering in Vietnamese)
In recent years, visual question answering (VQA) has attracted attention from the research community because of its highly potential applications (such as virtual assistance on intelligent cars, assistant devices for blind people, or…
4 papers · 0 benchmarks
PASTEL is a parallelly annotated stylistic language dataset.
4 papers · 0 benchmarks
PCFG SET (Probabilistic Context Free Grammar String Edit Task)
The Probabilistic Context Free Grammar String Edit Task (PCFG SET) dataset is a dataset with sequence to sequence problems specifically designed to test different aspects of compositional generalisation.
4 papers · 0 benchmarks
PELD is a text-based emotional dialog dataset with personality traits for speakers.
4 papers · 0 benchmarks
ParaShoot is the first question answering dataset in modern Hebrew.
4 papers · 0 benchmarks
Polaris (Polaris dataset)
The Polaris dataset offers a large-scale, diverse benchmark for evaluating metrics for image captioning, surpassing existing datasets in terms of size, caption diversity, number of human judgments, and granularity of the evaluations.
4 papers · 0 benchmarks
ProofNet# is an evaluation benchmark derived from the original ProofNet, which contains 371 paired examples of informal undergraduate mathematical statements and their corresponding formalizations.
4 papers · 0 benchmarks
QC-Science contains 47832 question-answer pairs belonging to the science domain tagged with labels of the form subject - chapter - topic.
4 papers · 1 benchmark
Aims to help V-NLIs recognize analytic tasks from free-form natural language by training and evaluating cutting-edge multi-label classification models.
4 papers · 0 benchmarks
RISeC (Recipe Instruction Semantics Corpus)
We propose a newly annotated dataset for information extraction on recipes.
4 papers · 0 benchmarks
RONEC (Romanian Named Entity Corpus)
Romanian Named Entity Corpus is a named entity corpus for the Romanian language.
4 papers · 0 benchmarks
RR (Review-Rebuttal)
Review-Rebuttal (RR) dataset is introduced to facilitate the study of argument pair extraction in the peer review and rebuttal domain.
4 papers · 1 benchmark
RTASC (ROBIN Technical Acquisition Speech Corpus)
The ROBIN Technical Acquisition Speech Corpus (ROBINTASC) was developed within the ROBIN project.
4 papers · 0 benchmarks
RWWD (Real World Worry Dataset)
Real World Worry Dataset (RWWD) captures the emotional responses of UK residents to COVID-19 at a point in time where the impact of the COVID19 situation affected the lives of all individuals in the UK.
4 papers · 0 benchmarks
The Restaurant-ACOS dataset is constructed based on the SemEval 2016 Restaurant dataset (Pontiki et al., 2016) and its expansion datasets (Fan et al., 2019; Xu et al., 2020).
4 papers · 1 benchmark
RetVQA (Retrieval-Based Visual Question Answering)
The RetVQA dataset is a large-scale dataset designed for Retrieval-Based Visual Question Answering (RetVQA).
4 papers · 1 benchmark
Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness.
4 papers · 1 benchmark
SCDE is a human-created sentence cloze dataset, collected from public school English examinations in China.
4 papers · 1 benchmark
The Situated Corpus Of Understanding Transactions (SCOUT) is a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration.
4 papers · 0 benchmarks
SLING (Sino LINGuistics)
SLING consists of 38K minimal sentence pairs in Mandarin Chinese grouped into 9 high-level linguistic phenomena.
4 papers · 0 benchmarks
SPACE is a large-scale opinion summarization benchmark for the evaluation of unsupervised summarizers.
4 papers · 1 benchmark
STAIR Captions is a large-scale dataset containing 820,310 Japanese captions.
4 papers · 0 benchmarks
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
SYMON (Synopses of Movie Narratives)
Contains 5,193 video summaries of popular movies and TV series.
4 papers · 0 benchmarks
SciDuet is a dataset for training and benchmarking models for automating document-to-slides generation.
4 papers · 0 benchmarks
SecQA is a specialized dataset created for the evaluation of Large Language Models (LLMs) in the domain of computer security.
4 papers · 0 benchmarks
SoftAttributes (SoftAttributes: Relative movie attribute dataset for soft attributes)
The dataset consists of sets of movie titles, with each set annotated with a single English soft attribute (subjective descriptive property, such as 'confusing' or 'romantic') and a reference movie.
4 papers · 0 benchmarks
SpeechInstruct is a large-scale cross-modal speech instruction dataset.
4 papers · 0 benchmarks
SynthPAI (SynthPAI: A Synthetic Dataset for Personal Attribute Inference)
SynthPAI was created to provide a dataset that can be used to investigate the personal attribute inference (PAI) capabilities of LLM on online texts.
4 papers · 1 benchmark
The Szeged Treebank is the largest fully manually annotated treebank of the Hungarian language.
4 papers · 0 benchmarks
TAC 2010 is a dataset for summarization that consists of 44 topics, each of which is associated with a set of 10 documents.
4 papers · 1 benchmark
TemporalWiki is a lifelong benchmark for ever-evolving LMs that utilizes the difference between consecutive snapshots of English Wikipedia and English Wikidata for training and evaluation, respectively.
4 papers · 0 benchmarks
TimeBankPT (Portuguese TimeBank)
TimeBankPT is a corpus of Portuguese text with annotations about time.
4 papers · 1 benchmark
Twitter Sentiment Analysis (Entity-Level Twitter Sentiment Analysis Dataset)
This is an entity-level Twitter Sentiment Analysis dataset.
4 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.