Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 43 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2017–2064 of 3,130

A Dataset for Relation Extraction of Natural-Products (A curated evaluation dataset for end-to-end Relation Extraction of relationships between organisms and natural-products)
A curated evaluation dataset for end-to-end Relation Extraction of relationships between organisms and natural-products.
1 paper · 0 benchmarks
We present a dataset of dialogs in which journalists of The Guardian replied to reader comments and identify the reasons why.
1 paper · 0 benchmarks
This is a dataset of 583,437 tweets by 155,715 users that were censored between 2012-2020 July.
1 paper · 0 benchmarks
A Large Scale Fish Dataset (A Large-Scale Dataset for Fish Segmentation and Classification)
This dataset contains 9 different seafood types collected from a supermarket in Izmir, Turkey for a university-industry collaboration project at Izmir University of Economics, and this work was published in ASYU 2020.
1 paper · 0 benchmarks
A2Dre (Subset of A2D Sentences which are not trivial)
We obtain A2Dre by selecting only instances that were labeled as non-trivial, which are 433 REs from 190 videos.
1 paper · 1 benchmark
A2Dre+ (Extension of A2D sentences where trivial cases where filtered)
A2Dre is a subset from the A2D test set including $433$~\textit{non-trivial} REs.
1 paper · 0 benchmarks
AAVE/SAE Paired Dataset contains 2019 intent-equivalent AAVE/SAE pairs.
1 paper · 0 benchmarks
AC-Bench: A Benchmark for Actual Causality Reasoning Dataset Description AC-Bench is designed to evaluate the actual causality (AC) reasoning capabilities of large language models (LLMs).
1 paper · 0 benchmarks
A benchmark environment based on the datasets "Adult" and "Names", which allows researchers to test how well their language model can abide by pre-defined access rights rules.
1 paper · 0 benchmarks
ACCORD CSQA is an extension of the popular CommonsenseQA (CSQA) dataset using ACCORD, a scalable framework for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) through controlled, multi-hop…
1 paper · 0 benchmarks
This repository provides full-text and metadata to the ACL anthology collection (80k articles/posters as of September 2022) also including .pdf files and grobid extractions of the pdfs.
1 paper · 0 benchmarks
Dataset Card for the ACR Appropriateness Criteria Corpus This dataset contains chunked guidelines and narratives from the ACR Appropriateness Criteria, an set of societal guidelines from the American College of Radiology (ACR) to help…
1 paper · 0 benchmarks
AESI (Athens Emotional States Inventory)
The development of ecologically valid procedures for collecting reliable and unbiased emotional data towards computer interfaces with social and affective intelligence targeting patients with mental disorders.
1 paper · 0 benchmarks
AGB-DE is a legal NLP corpus for the automated detection of potentially void clauses in German standard form consumer contracts.
1 paper · 1 benchmark
Consists of 7.5k sentences with gapping (as well as 15k relevant negative sentences) and comprises data from various genres: news, fiction, social media and technical texts.
1 paper · 0 benchmarks
Replication Material This document contains the necessary materials and instructions to replicate the findings presented in our paper.
1 paper · 0 benchmarks
This project contains instructions and codes to reconstruct a dataset for the development and evaluation of forensic tools for detecting machine-generated text in social media.
1 paper · 0 benchmarks
AISECKG (AISecKG: Knowledge Graph Dataset for Cybersecurity Education)
Cybersecurity education is exceptionally challenging as it involves learning the complex attacks; tools and developing critical problem-solving skills to defend the systems.
1 paper · 0 benchmarks
AISIA-VN-Review-S (AISIA-VN-Review-F)
In AISIA-VN-Review-S and AISIA-VN-Review-F datasets, we first collect 450K customer reviewing comments from various e–commerce websites.
1 paper · 0 benchmarks
ANTILLES (ANTILLES: An Open French Linguistically Enriched Part-of-Speech Corpus)
ANTILLES is a part-of-speech tagging corpus based on UDFrench-GSD which was originally created in 2015 and is based on the universal dependency treebank v2.0.
1 paper · 1 benchmark
AP (Adversarial Paraphrase)
This is a paraphrasing dataset created using the adversarial paradigm.
1 paper · 1 benchmark
AQL-22 (Archive Query Log)
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years.
1 paper · 0 benchmarks
ARCT (Argument Reasoning Comprehension Task)
Freely licensed dataset with warrants for 2k authentic arguments from news comments.
1 paper · 0 benchmarks
ASLG-PC12 (English-ASL Gloss Parallel Corpus 2012)
An artificial corpus built using grammatical dependencies rules due to the lack of resources for Sign Language.
1 paper · 1 benchmark
A Rich Annotated Mandarin Conversational (RAMC) Speech Dataset, including 180 hours of Mandarin Chinese dialogue, 150, 10 and 20 hours for the training set, development set and test set respectively.
1 paper · 0 benchmarks
ASyMOB (ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark)
ASyMOB (pronounced Asimov, in tribute to the renowned author), is a novel assessment framework focused exclusively on symbolic manipulation, featuring 17,092 unique math challenges, organized by similarity and complexity.
1 paper · 0 benchmarks
Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers.
1 paper · 0 benchmarks
Science Journal for Kids Data This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers.
1 paper · 0 benchmarks
The AbstRCT dataset consists of randomized controlled trials retrieved from the MEDLINE database via PubMed search.
1 paper · 2 benchmarks
The dataset contains 7,601 Gab posts classified on three different aspects: abuse presence or not, abuse severity and abuse target.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
AdvSuffixes (Adversarial Suffixes)
AdvSuffixes - Information AdvSuffixes is a curated dataset of adversarial prompts and suffixes designed to evaluate and enhance the robustness of large language models (LLMs) against adversarial attacks.
1 paper · 0 benchmarks
The Advice-Seeking Questions (ASQ) dataset is a collection of personal narratives with advice-seeking questions.
1 paper · 0 benchmarks
We filter and match the landmarks in the Google Landmarks dataset with their OpenStreetMap polygons and filter for those located in the United States, resulting in 602 landmarks.
1 paper · 0 benchmarks
AesVQA is a dataset that contains 72168 high-quality images and 324756 pairs of aesthetic questions.
1 paper · 0 benchmarks
The Alexa Point of View dataset is point of view conversion dataset, a parallel corpus of messages spoken to a virtual assistant and the converted messages for delivery.
1 paper · 1 benchmark
We introduce the novel task of multimodal puzzle solving, framed within the context of visual question-answering.
1 paper · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
AlpacaEval in Thai.
1 paper · 0 benchmarks
The Ambiguous VQA dataset is a dataset of ambiguous questions about images.
1 paper · 0 benchmarks
Amharic Error Corpus is a manually annotated spelling error corpus for Amharic, lingua franca in Ethiopia.
1 paper · 0 benchmarks
Among Them (Among Them dialogs and persuasion labels)
The dataset contains dialogs of different LLMs from the discussion phase of a text-based Among Us-like game.
1 paper · 0 benchmarks
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable.
1 paper · 1 benchmark
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
1 paper · 0 benchmarks
Official repository for the AnnoMI dataset: the first public collection of expert-annotated MI transcripts.
1 paper · 0 benchmarks
Includes two datasets for this task, one for English-French (En-Fr) and another for English-German (En-De).
1 paper · 0 benchmarks
AnswerSumm is a dataset of 4,631 CQA threads for answer summarization, curated by professional linguists.
1 paper · 0 benchmarks
Antibody Watch is a dataset of text snippets extracted from over 2000 PubMed articles with annotations denoting specificity of antibodies.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.