Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 65 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 3073–3120 of 3,130
ESP dataset (Evaluation for Styled Prompt dataset) is a new benchmark for zero-shot domain-conditional caption generation.
0 papers · 0 benchmarks
The EU-ADR corpus is a biomedical relation extraction dataset that contains 100 abstracts, with relations between drug, disorder, and targets.
0 papers · 0 benchmarks
Eduge (Eduge news classification dataset)
Eduge news classification dataset provided by Bolorsoft LLC.
0 papers · 0 benchmarks
The English-Pashto Language Dataset (EPLD) is a comprehensive resource aimed to provide linguistic insights into the Pashto language.
0 papers · 0 benchmarks
A dataset specifically tailored to the biotech news sector, aiming to transcend the limitations of existing benchmarks.
0 papers · 0 benchmarks
A high-quality dataset forms the foundation for machine learning-based predictions of structural load capacity.
0 papers · 0 benchmarks
We introduce FortisAVQA, a dataset designed to assess the robustness of AVQA models.
0 papers · 0 benchmarks
GenAI-Bench benchmark consists of 1,600 challenging real-world text prompts sourced from professional designers.
0 papers · 0 benchmarks
ICConv (A Large-scale Automated Intent-oriented and Context-aware Conversational Search Dataset)
The dataset contains 105,811 information-seeking conversations converted from MS MARCO.
0 papers · 0 benchmarks
This is a Dataset for Arabic/English text detection and optical character recognition.
0 papers · 0 benchmarks
The ISIBengaliCharacter dataset contains 158 classes of Bengali numerals, characters or their parts.
0 papers · 0 benchmarks
IndirectRequests is an LLM-generated dataset of user utterances in a task-oriented dialogue setting where the user does not directly specify their preferred slot value.
0 papers · 0 benchmarks
InpaintCOCO is a benchmark to understand fine-grained concepts in multimodal models (vision-language) similar to Winoground.
0 papers · 0 benchmarks
LMCQA (Legal Multiple Choice Question Answering)
This dataset contains a set of multiple-choice questions related to various legal topics.
0 papers · 0 benchmarks
LSARS (Large Scale Abstractive multi-Review Summarization)
In an active e-commerce environment, customers process a large number of reviews when deciding on whether to buy a product or not.
0 papers · 0 benchmarks
Court decisions from 2017 and 2018 were selected for the dataset, published online by the Federal Ministry of Justice and Consumer Protection.
0 papers · 0 benchmarks
This dataset contains 4606 articles from 1996 to 2024 that were presented in MIE (Medical Informatics Europe Conference) conferences.
0 papers · 0 benchmarks
MIMIC Meme Dataset (Misogyny Identification in Multimodal Internet Content in Hindi-English Code-Mix Language)
This dataset endeavors to fill the research void by presenting a meticulously curated collection of misogynistic memes in a code-mixed language of Hindi and English.
0 papers · 0 benchmarks
IQ testing has served as a foundational methodology for evaluating human cognitive capabilities, deliberately decoupling assessment from linguistic background, language proficiency, or domain-specific knowledge to isolate core competencies…
0 papers · 0 benchmarks
Mac-Morpho is a corpus of Brazilian Portuguese texts annotated with part-of-speech tags.
0 papers · 0 benchmarks
MoralChoice is a survey dataset to evaluate the moral beliefs encoded in LLMs.
0 papers · 0 benchmarks
A great number of situational comedies (sitcoms) are being regularly made and the task of adding laughter tracks to these is a critical task.
0 papers · 0 benchmarks
Multimodal Large Language Models (MLLMs) have shown significant promise in various applications, leading to broad interest from researchers and practitioners alike.
0 papers · 0 benchmarks
NERGRIT involves machine learning based NLP Tools and a corpus used for Indonesian Named Entity Recognition, Statement Extraction, and Sentiment Analysis.
0 papers · 1 benchmark
NHR-Edit (NoHumansRequired Edit Dataset)
NHR-Edit is a training dataset for instruction-based image editing.
0 papers · 0 benchmarks
NIAN (Needle in a Needlestack)
The Needle in a Needlestack (NIAN) is a new benchmark designed to measure how well Language Learning Models (LLMs) pay attention to the information in their context window¹.
0 papers · 0 benchmarks
NSMC (Naver Sentiment Movie Corpus)
This is a movie review dataset in the Korean language.
0 papers · 0 benchmarks
Numbers Station Text to SQL
0 papers · 0 benchmarks
POPCORN (POPCORN: Fictional and Synthetic Intelligence Reports for Named Entity Recognition and Relation Extraction Tasks)
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually.
0 papers · 0 benchmarks
The Parzival dataset consists of 47 pages by three writers.
0 papers · 0 benchmarks
There are about 208 000 jokes in this database scraped from three sources.
0 papers · 0 benchmarks
The PropBankPT (Branco et al., 2012) is a set of sentences annotated with their constituency structure and semantic role tags, composed of 3,406 sentences and 44,598 tokens taken from the Wall Street Journal translated.
0 papers · 0 benchmarks
QuAIL (Question Answering for Artificial Intelligence)
A new kind of question-answering dataset that combines commonsense, text-based, and unanswerable questions, balanced for different genres and reasoning types.
0 papers · 0 benchmarks
A review on raw subjective scores and data manipulation for before and after refining Mean opinion Scores
0 papers · 0 benchmarks
SEN (Sentiment analysis of Entities in News headlines)
SEN is a novel publicly available human-labelled dataset for training and testing machine learning algorithms for the problem of entity level sentiment analysis of political news headlines.
0 papers · 0 benchmarks
Dataset Card for SENTINEL: Mitigating Object Hallucinations via Sentence-Level Early Intervention For the details of this dataset, please refer to the documentation of the GitHub repo.
0 papers · 0 benchmarks
This dataset was taken from the SIGARRA information system at the University of Porto (UP).
0 papers · 0 benchmarks
SILD (Survey Item Linking Dataset)
This dataset contains a collection of texts from publications from a broad range of social science domains (e.g., economics, politics, psychology, etc.).
0 papers · 0 benchmarks
SMCOVID19-CT (Contact Tracing Data (from Italian SM-COVID-19 App))
We present a real data analysis of a CT experiment that was conducted in Italy for 8 months and involved more than 100,000 CT app users.
0 papers · 0 benchmarks
Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources The STEM ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity…
0 papers · 0 benchmarks
STVD-FC is the largest public dataset on the political content analysis and fact-checking tasks.
0 papers · 0 benchmarks
The Second HAREM was an evaluation exercise in Portuguese Named Entity Recognition.
0 papers · 0 benchmarks
SourceData-NLP (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)
Introduction: The scientific publishing landscape is expanding rapidly, creating challenges for researchers to stay up-to-date with the evolution of the literature.
0 papers · 0 benchmarks
SportsSum is a Chinese sports game summarization dataset that contains 5,428 soccer games of live commentaries and the corresponding news articles.
0 papers · 0 benchmarks
Dataset Introduction TFHAnnotatedDataset is an annotated patent dataset pertaining to thin film head technology in hard-disk.
0 papers · 0 benchmarks
The increase in religiously motivated hate on social media is clear and ongoing.
0 papers · 0 benchmarks
TRACT (Tweets Reporting Abuse Classification Task Corpus)
TRACT is a small scale manually annotated corpus for abuse classification problem.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.