Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 36 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1681–1728 of 3,130

AWARE (AWARE: Aspect-Based Sentiment Analysis Dataset of Apps Reviews for Requirements Elicitation)
The peer-reviewed paper of AWARE dataset is published in ASEW 2021, and can be accessed through: http://doi.org/10.1109/ASEW52652.2021.00049.
2 papers · 3 benchmarks
We present the AWS documentation corpus, an open-book QA dataset, which contains 25,175 documents along with 100 matched questions and answers.
2 papers · 0 benchmarks
Almawave-SLU is the first Italian dataset for Spoken Language Understanding (SLU).
2 papers · 0 benchmarks
Amazon-PQA is a product question-answer dataset.
2 papers · 0 benchmarks
Amharic - English Parallel Corpus for Machine Translation contains 33,955 sentence pairs extracted text from such news platforms as Ethiopian Press Agency1, Fana Broadcasting Corporate2, and Walta Information Center3.
2 papers · 0 benchmarks
Sentiment detection remains a pivotal task in natural language processing, yet its development in Arabic lags due to a scarcity of training materials compared to English.
2 papers · 0 benchmarks
AraCOVID19-MFH (AraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News and Hate Speech Detection Dataset)
AraCOVID19-MFH is a manually annotated multi-label Arabic COVID-19 fake news and hate speech detection dataset.
2 papers · 0 benchmarks
The ArxivPapers dataset is an unlabelled collection of over 104K papers related to machine learning and published on arXiv.org between 2007–2020.
2 papers · 0 benchmarks
Audio-alpaca: A preference dataset for aligning text-to-audio models Audio-alpaca is a pairwise preference dataset containing about 15k (prompt,chosen, rejected) triplets where given a textual prompt, chosen is the preferred generated…
2 papers · 0 benchmarks
BC7 NLM-Chem (BioCreative VII NLM-Chem)
Full-text chemical identification and indexing in PubMed articles.
2 papers · 3 benchmarks
BabySLM is a language-acquisition-friendly benchmark to probe speech-based LMs at the lexical and syntactic levels, both of which are compatible with the vocabulary typical of children's language experiences.
2 papers · 0 benchmarks
Named entities in Bavarian text Details: Siyao Peng, Zihang Sun, Huangyan Shan, Marie Kolm, Verena Blaschke, Ekaterina Artemova, and Barbara Plank.
2 papers · 0 benchmarks
A set of basque documents annotated with EusTimeML - a mark-up language for temporal information in Basque.
2 papers · 1 benchmark
Dataset of the Beacon3D benchmark: Unveiling the Mist over 3D Vision-Language Understanding: Object-centric Evaluation with Chain-of-Analysis.
2 papers · 0 benchmarks
BiGe (Bielefeld Gesture Corpus)
The BiGe corpus is comprised of 54.360 shots of interest extracted from TED and TEDx talks.
2 papers · 0 benchmarks
BiRdQA is a bilingual multiple-choice question answering dataset with 6614 English riddles and 8751 Chinese riddles.
2 papers · 0 benchmarks
BiasCorp is a dataset for racism detection containing 139,090 comments and news segment from three specific sources - Fox News, BreitbartNews and YouTube.
2 papers · 0 benchmarks
Biographical (Biographical: A Semi-Supervised Relation Extraction Dataset)
Biographical is a semi-supervised dataset for RE.
2 papers · 0 benchmarks
BoostCLIR is a bilingual (Japanese-English) corpus of patent abstracts, extracted from the MAREC patent data, and the data from the NTCIR PatentMT workshop collections, accompanied with relevance judgements for the task of patent prior-art…
2 papers · 0 benchmarks
BugRepo (Bug Reports)
BugRepo maintains a collection of bug reports that are publicly available for research purposes.
2 papers · 0 benchmarks
CA4P-483 is a dataset designed to facilitate the sequence labeling tasks and regulation compliance identification between privacy policies and software.
2 papers · 0 benchmarks
CAVES (A Dataset to facilitate Explainable Classification and Summarization of Concerns towards COVID Vaccines)
CAVES is the first large-scale dataset containing about 10k COVID-19 anti-vaccine tweets labelled into various specific anti-vaccine concerns in a multi-label setting.
2 papers · 0 benchmarks
CC-News (CommonCrawl News dataset)
CommonCrawl News is a dataset containing news articles from news sites all over the world.
2 papers · 0 benchmarks
CIC (Catalonia Independence Corpus)
The dataset is annotated with stance towards one topic, namely, the independence of Catalonia.
2 papers · 3 benchmarks
CII-Bench (Chinese Image Implication understanding Benchmark)
We introduce the Chinese Image Implication Understanding Benchmark CII-Bench, a new benchmark measuring the higher-order perceptual, reasoning and comprehension abilities of MLLMs when presented with complex Chinese implication images.
2 papers · 0 benchmarks
CLEAR-Bias (Corpus for Linguistic Evaluation of Adversarial Robustness against Bias)
CLEAR-Bias is a benchmark dataset designed to evaluate the robustness of large language models (LLMs) against bias elicitation, particularly under adversarial conditions.
2 papers · 0 benchmarks
CLSE (Corpus of Linguistically Significant Entities)
2 papers · 0 benchmarks
COFAR (Commonsense and Factual Reasoning in Image Search)
The COFAR (COmmonsense and FActual Reasoning) dataset is a collection of images and text queries specifically designed to challenge and evaluate image search models that aim to go beyond simple visual matching.
2 papers · 1 benchmark
CORE (Company Relation Extraction)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
CQR (Contextual Query Rewrite)
CQR is an extension to the Stanford Dialogue Corpus.
2 papers · 0 benchmarks
CREPE is QA dataset containing a natural distribution of presupposition failures from online information-seeking forums.
2 papers · 0 benchmarks
A fundamental characteristic common to both human vision and natural language is their compositional nature.
2 papers · 1 benchmark
Chinese Spelling Correction Dataset for errors generated by pinyin IME (CSCD-IME), a dataset containing 40,000 annotated sentences from real posts of official media on Sina Weibo.
2 papers · 0 benchmarks
CUP (Context-sitUated Pun) is a dataset containing 4.5k tuples of context words and pun pairs, each labelled with whether they are compatible for composing a pun.
2 papers · 0 benchmarks
CareCall (CareCall for Seniors)
carecall is a Korean dialogue dataset for role-satisfying dialogue systems.
2 papers · 0 benchmarks
Catalan TimeBank 1.0 was developed by researchers at Barcelona Media and consists of Catalan texts in the AnCora corpus annotated with temporal and event information according to the TimeML specification language.
2 papers · 1 benchmark
CheGeKa is a Jeopardy!-like Russian QA dataset collected from the official Russian quiz database ChGK.
2 papers · 1 benchmark
We introduce ChinaTravel, the first open-ended benchmark grounded in authentic Chinese travel requirements collected from 1,154 human participants.
2 papers · 0 benchmarks
Classifiers are function words that are used to express quantities in Chinese and are especially difficult for language learners.
2 papers · 0 benchmarks
Chinese Gigaword corpus consists of 2.2M of headline-document pairs of news stories covering over 284 months from two Chinese newspapers, namely the Xinhua News Agency of China (XIN) and the Central News Agency of Taiwan (CNA).
2 papers · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
2 papers · 0 benchmarks
The Circa (meaning ‘approximately’) dataset aims to help machine learning systems to solve the problem of interpreting indirect answers to polar questions.
2 papers · 0 benchmarks
CiteWorth is a a large, contextualized, rigorously cleaned labelled dataset for cite-worthiness detection built from a massive corpus of extracted plain-text scientific documents.
2 papers · 0 benchmarks
The Climate Change Claims dataset for generating fact checking summaries contains claims broadly related to climate change and global warming from climatefeedback.org.
2 papers · 0 benchmarks
The dataset was created to address the crucial need for effective Extreme Weather Events Detection (EWED), an increasingly urgent task due to the rising frequency of such events driven by global warming.
2 papers · 0 benchmarks
This dataset is created from MIMIC-III (Medical Information Mart for Intensive Care III) and contains simulated patient admission notes.
2 papers · 4 benchmarks
CoNLL-2020 (CoNLLpp)
A test dataset that annotated articles in 2020 following the CoNLL-2003 NER task.
2 papers · 1 benchmark
CoVaxFrames includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.