Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 45 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 2113–2160 of 3,130
BioFuelQR is a dataset consisting of complex reasoning questions related to catalyst discovery in biofuels.
1 paper · 0 benchmarks
BioLeaflets is a biomedical dataset for Data2Text generation.
1 paper · 0 benchmarks
📚 BlendNet The dataset contains $12k$ samples.
1 paper · 0 benchmarks
This dataset is parallel text for Bornholmsk and Danish.
1 paper · 0 benchmarks
With the remarkable capability to reach the public instantly, social media has become integral in sharing scholarly articles to measure public response.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Brazilian Protest is a dataset for event filtering that focuses on protests in multi-modal social media data, with most of the text in Portuguese.
1 paper · 0 benchmarks
Dataset of 5,591 labeled issue tickets.
1 paper · 0 benchmarks
This is the C++ dataset used in the TASTY research paper which was published at the ICLR DL4Code (Deep Learning for Code) workshop.
1 paper · 0 benchmarks
📚 CADBench CADBench is a comprehensive benchmark to evaluate the ability of LLMs to generate CAD scripts.
1 paper · 0 benchmarks
CAGUI (Chinese Android GUI Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data.
1 paper · 3 benchmarks
CAsT-answerability dataset contains binary answerability labels on three levels: sentence, passage, and ranking.
1 paper · 0 benchmarks
CC-Riddle is a Chinese character riddle dataset covering the majority of common simplified Chinese characters by crawling riddles from the Web and generating brand new ones.
1 paper · 0 benchmarks
To address the scarcity of high-quality safety datasets in the Chinese, we open-sourced the CCI (Chinese Corpora Internet) dataset on November 29, 2023.
1 paper · 0 benchmarks
CCPT (Conceptual Combination with Property Type)
CCPT is a dataset containing 12.3K triplets of noun phrases, properties, and property types for conceptual combination.
1 paper · 0 benchmarks
CCSE (Chinese Character Stroke Extraction)
Chinese Character Stroke Extraction (CCSE) is a benchmark containing two large-scale datasets: Kaiti CCSE (CCSE-Kai) and Handwritten CCSE (CCSE-HW).
1 paper · 0 benchmarks
CECW (Colorful Extended Cleanup World)
The CECW dataset is a color-extended version of the Cleanup World (CW) borrowed from the mobile-manipulation robot domain.
1 paper · 0 benchmarks
CEREC (Corpus for Entity Resolution in Email Conversations)
CEREC is a large scale corpus for entity resolution in email conversations.
1 paper · 0 benchmarks
- CFEVER is a Chinese Fact Extraction and VERification dataset published at AAAI 2024.
1 paper · 0 benchmarks
CHAMP (Concept and Hint-Annotated Math Problems)
The Concept and Hint-Annotated Math Problems (CHAMP) consists of high school math competition problems, annotated with concepts, or general math facts, and hints, or problem-specific tricks.
1 paper · 0 benchmarks
CHIP Clinical Trial Classification, a dataset aimed at classifying clinical trials eligibility criteria, which are fundamental guidelines of clinical trials defined to identify whether a subject meets a clinical trial or not, is used for…
1 paper · 1 benchmark
CHORD (CHOrus Recognition Dataset)
CHORD is the first chorus recognition dataset containing 627 songs for public use.
1 paper · 0 benchmarks
CI-ToD is a dataset for Consistency Identification in Task-oriented Dialog system.
1 paper · 0 benchmarks
CKBP v2 is a new CSKB Population benchmark, which addresses the two mentioned problems by using experts instead of crowd-sourced annotation and by adding diversified adversarial samples to make the evaluation set more representative.
1 paper · 0 benchmarks
CLEVR Mental Rotation Tests (CLEVR-MRT) is a new version of the CLEVR dataset.
1 paper · 0 benchmarks
CLIPS (Corpora e Lessici dell'Italiano Parlato e Scritto)
CLIPS, ovvero Corpora e Lessici dell'Italiano Parlato e Scritto, è uno degli otto progetti (Progetto n.
1 paper · 0 benchmarks
CLUES (Constrained Language Understanding Evaluation Standard)
CLUES (Constrained Language Understanding Evaluation Standard) is a benchmark for evaluating the few-shot learning capabilities of NLU models.
1 paper · 0 benchmarks
CLUES is a benchmark for Classifier Learning Using natural language ExplanationS, consisting of a range of classification tasks over structured data along with natural language supervision in the form of explanations.
1 paper · 0 benchmarks
CLaRO is a new dataset of 234 Competency Questions that had been processed automatically into 106 patterns.
1 paper · 0 benchmarks
CMeIE (Chinese Medical Information Extraction Dataset)
Chinese Medical Information Extraction, a dataset that is also released in CHIP2020, is used for CMeIE task.
1 paper · 1 benchmark
COAT (CommonSense Object Affordance Task)
Useful for checking the physical reasoning capabilities in household agents.
1 paper · 0 benchmarks
COFFE (COFFE: A Code Efficiency Benchmark for Code Generation)
COFFE COFFE is a Python benchmark for evaluating the time efficiency of LLM-generated code.
1 paper · 0 benchmarks
The causal reasoning dataset is generated using the Causal Reasoning in Closed Daily Activities (COLD) framework that helps evaluate large language models (LLMs) on their causal reasoning abilities within real-world, everyday activities.
1 paper · 0 benchmarks
COLLIE-v1 is a dataset with 2080 instances comprising 13 constraint structures designed for text generation under constraints.
1 paper · 0 benchmarks
COMFORT (Consistent Multilingual Frame of Reference Test)
COMFORT is an evaluation protocol to systematically assess the spatial reasoning capabilities of VLMs.
1 paper · 0 benchmarks
The COPA-HR dataset (Choice of plausible alternatives in Croatian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology.
1 paper · 0 benchmarks
COSTRA 1.0 is a dataset of complex sentence transformations.
1 paper · 0 benchmarks
These datasets were used in the paper 'Evaluation of Thematic Coherence in Microblogs' (ACL, 2021).
1 paper · 0 benchmarks
The dataset contains Tweet IDs along with the location and tweet timestamp.
1 paper · 0 benchmarks
The data contains CSV files with anonymized user names, tweet texts, vaccine stance, cumulative score for the vaccine stance, location, and topic information.
1 paper · 0 benchmarks
COVID-19-TweetIDs (Tracking Social Media Discourse About the COVID-19 Pandemic: Development of a Public Coronavirus Twitter Data Set)
Since the inception of our collection, we have actively maintained and updated our GitHub repository on a weekly basis.
1 paper · 0 benchmarks
The Covid19-CountryImage dataset is a Twitter dataset which contains COVID-19-related tweets.
1 paper · 0 benchmarks
COVMis-Stance is a stance detection dataset for COVID-19 misinformation.
1 paper · 0 benchmarks
CPMC (crawled persian medical corpus)
a 90 million token medical corpus crawled from medical websites
1 paper · 0 benchmarks
In this repository you can find all the elaborate results that were used for the simulated evaluation of an innovative, optimized for real-life use, STC-based, multi-robot Coverage Path Planning (mCPP) algorithm.
1 paper · 0 benchmarks
CRED (Crowd Reaction Estimation Dataset)
In the realm of social media, understanding and predicting post reach is a significant challenge.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.