Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 48 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2257–2304 of 3,130

DpgMedia2019 is a Dutch news dataset for partisanship detection.
1 paper · 0 benchmarks
Dubbing Test Set consists of two subsets extracted from the En→De test set of COVOST-2, a large-scale multilingual speech translation corpus based on Common Voice.
1 paper · 0 benchmarks
As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial.
1 paper · 0 benchmarks
E-Manual Corpus is a corpus of 307,957 E-manuals, used for pre-training models for Question Answering on e-manuals.
1 paper · 0 benchmarks
E2E Refined is a dataset for sentence classification.
1 paper · 0 benchmarks
EAPD (Expert-labeled Aesthetics Perception Database)
An expert benchmark aiming to comprehensively evaluate the aesthetic perception capacities of MLLMs.
1 paper · 0 benchmarks
ECTF (Early COVID-19 Twitter Fake news)
ECTF is a dataset for Twitter fake news detection in the Covid-19 domain.
1 paper · 0 benchmarks
EHT (The English Headline Treebank)
The English Headline Treebank (EHT) is an English headline treebank of 1,055 manually annotated and adjudicated universal dependency (UD) syntactic dependency trees to encourage research in improving NLP pipelines for English headlines.
1 paper · 0 benchmarks
For Emotion Interpretation task
1 paper · 2 benchmarks
The ELITR ECA corpus is a multilingual corpus derived from publications of the European Court of Auditors.
1 paper · 0 benchmarks
ELITR Minuting Corpus in JSON format.
1 paper · 0 benchmarks
first everyday task dataset featuring COT outputs, diverse task designs, detailed re-plan processes, along with SFT and DPO sub-datasets.
1 paper · 0 benchmarks
EPIC30M contains a subset of 26.2 millions tweets related to three general diseases, namely Ebola, Cholera and Swine Flu, and another subset of 4.7 millions tweets of six global epidemic outbreaks, including 2009 H1N1 Swine Flu, 2010 Haiti…
1 paper · 0 benchmarks
ERD (Educational Resource Discovery)
ERD (Educational Resource Discovery) is a corpus of 39,728 manually labeled web resources and 659 queries from NLP, Computer Vision (CV), and Statistics (STATS) for educational resource discovery.
1 paper · 0 benchmarks
Dataset Card for ESG/DLT Named Entity Recognition Dataset This dataset contains named entities related to Distributed Ledger Technology (DLT) and Environmental, Social, and Governance (ESG) topics created to support research in these areas…
1 paper · 0 benchmarks
ESG-FTSE (ESG-FTSE corpus)
We present ESG-FTSE, the first corpus comprised of news articles with Environmental, Social and Governance (ESG) relevance annotations.
1 paper · 0 benchmarks
ESP (Evaluation for Styled Prompt)
ESP dataset (Evaluation for Styled Prompt dataset) is a benchmark for zero-shot domain-conditional caption generation.
1 paper · 0 benchmarks
Dataset Description EUROPA is a dataset designed for training and evaluating multilingual keyphrase generation models in the legal domain.
1 paper · 0 benchmarks
The EVI dataset is a challenging, multilingual spoken-dialogue dataset with 5,506 dialogues in English, Polish, and French.
1 paper · 3 benchmarks
This is an assembly dataset built on top of ShellcodeIA32, a dataset for automatically generating assembly from natural language descriptions that consists of 3,200 assembly instructions, commented in the English language, which were…
1 paper · 0 benchmarks
This dataset contains samples to generate Python code for security exploits.
1 paper · 0 benchmarks
A large dataset of over 18,000,000 English tweets posted by ∼7K echo users was constructed in the following manner: 1.
1 paper · 0 benchmarks
Echo Corpus (Arviv et al, 2021) infused with information from KnowledJe (Halevy, 2023).
1 paper · 0 benchmarks
Educational Grade School Math (EGSM) contains 2,093 question/answer pairs generated by MATHWELL, a reference-free educational grade school math word problem generator that outputs a word problem and Program of Thought (PoT) solution based…
1 paper · 0 benchmarks
EpiK-Eval (Epistemic Knowledge Evaluation)
Benchmark to evaluate the capability of LMs to consolidate and recall information from multiple training documents.
1 paper · 0 benchmarks
The ConcoDisco Corpus is an English-French parallel corpus with discourse relations (DRs) and discourse connectives (DCs) annotations.
1 paper · 0 benchmarks
EventEA is an event-centric entity alignment dataset, harvested from EventKG, DBpedia and Wikidata.
1 paper · 0 benchmarks
Intermediate annotations from the FEVER dataset that describe original facts extracted from Wikipedia and the mutations that were applied, yielding the claims in FEVER.
1 paper · 0 benchmarks
ExBAN (ExBAN Corpus (Explanations for BAyesian Networks))
The ExBAN dataset: a corpus of NL explanations generated by crowd-sourced participants presented with the task of explaining simple Bayesian Network (BN) graphical representations.
1 paper · 0 benchmarks
ExPUNations is a humor dataset with such extensive and fine-grained annotations specifically for puns.
1 paper · 0 benchmarks
The ExaASC dataset is a dataset for Target-based Stance Detection in the Arabic Language that contains different types of targets like persons, entities and events.
1 paper · 0 benchmarks
Expository Prose (Expository-Prose-V1)
Expository-Prose-V1 is a collection of specially-curated corpora gathered from diverse sources, ranging from research papers (arXiv) to European Parliament proceedings (EuroParl).
1 paper · 0 benchmarks
Minecraft Corpus dataset with builder utterance annotations
1 paper · 0 benchmarks
The EyeInfo Dataset is an open-source eye-tracking dataset created by Fabricio Batista Narcizo, a research scientist at the IT University of Copenhagen (ITU) and GN Audio A/S (Jabra), Denmark.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
FABSA (An aspect-based sentiment analysis dataset of Customer Feedback reviews)
FABSA, An aspect-based sentiment analysis dataset in the Customer Feedback space (Trustpilot, Google Play and Apple Store reviews).
1 paper · 2 benchmarks
FCGEC (FCGEC: Fine-Grained Corpus for Chinese Grammatical Error Correction)
a fine-grained corpus to detect, identify and correct the chinese grammatical errors.
1 paper · 1 benchmark
FCoT (Foreground Chain-of-Thought)
FCoT (Chain‑of‑Thought Segmentation) is replicate the step-by-step reasoning process a human annotator follows when using SAM2 to generate masks.
1 paper · 0 benchmarks
FETA Car-Manuals (FETA Car-Manuals dataset, image-text retrieval for foundation models' expert data performance.)
FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures.
1 paper · 2 benchmarks
FETA benchmark focuses on text-to-image and image-to-text retrieval in public car manuals and sales catalogue brochures.
1 paper · 0 benchmarks
FICLE (Factual Inconsistency CLassification with Explanation)
The FICLE dataset is a derivative of the FEVER dataset, which is a collection of 185,445 claims generated by modifying sentences obtained from Wikipedia.
1 paper · 0 benchmarks
Optical images of printed circuit boards as well as detailed annotations of any text, logos, and surface-mount devices (SMDs).
1 paper · 0 benchmarks
FIG-Loneliness (FIne-Grained Loneliness) is a dataset collected by using Reddit posts in two young adult-focused forums and two loneliness related forums consisting of a diverse age group.
1 paper · 0 benchmarks
FMC-MWO2KG (The MWO2KG Failure Mode Classification Dataset)
The Failure Mode Classification dataset released in the paper "MWO2KG and Echidna: Constructing and exploring knowledge graphs from maintenance data" by Stewart et al.
1 paper · 1 benchmark
FSC-P2 (Fearless Steps Challenge Phase2)
The Fearless Steps Initiative by UTDallas-CRSS led to the digitization, recovery, and diarization of 19,000 hours of original analog audio data, as well as the development of algorithms to extract meaningful information from this…
1 paper · 0 benchmarks
FSOCO is a collaborative dataset for vision-based cone detection systems in Formula Student Driverless competitions.
1 paper · 0 benchmarks
FTR-18 is a multilingual rumour dataset on football transfer news.
1 paper · 0 benchmarks
The FairTranslate Dataset includes 2,418 sentence pairs, each centered around an occupation, designed to assess gender expression and translation in English-French contexts.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.