Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 46 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 2161–2208 of 3,130
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CSAbstruct is a new dataset of annotated computer science abstracts with sentence labels according to their rhetorical roles.
1 paper · 0 benchmarks
The dataset contains gold-standard summary labels for 39 "CSI: Crime Scene Investigation" episodes from seasons 1-5.
1 paper · 0 benchmarks
We present CSL, a large-scale Chinese Scientific Literature dataset, which contains the titles, abstracts, keywords and academic fields of 396,209 papers.
1 paper · 0 benchmarks
CSPRD (Chinese Stock Policy Retrieval Dataset)
The Chinese Stock Policy Retrieval Dataset (CSPRD) contains a Chinese policy corpus of 10,002 articles and 709 prospectus examples from 545 companies listed on China’s Science and Technology Innovation Board (STAR Market).
1 paper · 0 benchmarks
CTFW is a large annotated procedural text dataset in the cybersecurity domain (3154 documents).
1 paper · 0 benchmarks
This dataset contains samples of CTI (Cyber Threat Intelligence) data in natural language, labeled with the corresponding adversarial techniques from the MITRE ATT&CK framework.
1 paper · 0 benchmarks
CUHK-QA is a dataset for natural language-based person search using iterative questioning.
1 paper · 0 benchmarks
CURE (A dataset for Clinical Understanding & Retrieval Evaluation)
CURE is a retrieval dataset with a monolingual and two cross-lingual conditions, with splits spanning ten medical domains.
1 paper · 0 benchmarks
CVE (Common Vulnerabilities and Exposures)
CVE stands for Common Vulnerabilities and Exposures.
1 paper · 0 benchmarks
In this dataset an uppertorso humanoid robot with 7-DOF arm explored 100 different objects belonging to 20 different categories using 10 behaviors: Look, Crush, Grasp, Hold, Lift, Drop, Poke, Push, Shake and Tap.
1 paper · 0 benchmarks
Capriccio is a sentiment classification dataset on tweets that simulates data drift.
1 paper · 0 benchmarks
In this dataset we added [Company Name, Car Model, Car Type, Fuel Type, Transmission, Engine (cc), Mileage, Kmsdriven, Buyers, Horsepower (kw), Year Price (Lakhs)]
1 paper · 1 benchmark
A synthetic dataset from an automobile manufacturer datasource.
1 paper · 0 benchmarks
A dataset of games played in the card game "Cards Against Humanity" (CAH), by human players, derived from the online CAH labs.
1 paper · 0 benchmarks
The CareerCoach 2022 gold standard is available for download in the NIF and JSON format, and draws upon documents from a corpus of over 99,000 education courses which have been retrieved from 488 different education providers.
1 paper · 0 benchmarks
Caselaw4 is a dataset of 350k common law judicial decisions from the U.S.
1 paper · 0 benchmarks
Casino Reviews (Online reviews of North American Casinos from Google Reviews)
This dataset contain online reviews gathered from google reviews written by north american casino users.
1 paper · 0 benchmarks
SyntaxGym, adapted for interventional interpretability.
1 paper · 1 benchmark
The dataset covers Hindi and Tamil, collected without the use of translation.
1 paper · 1 benchmark
Taking Advice from ChatGPT is a laboratory study of how student participants incorporate advice generated by ChatGPT.
1 paper · 0 benchmarks
Dataset Overview vanilla.csv: Represents the interactions without specific role-play instructions.
1 paper · 0 benchmarks
The dataset contains two few-shot chemical fine-grained entity extraction datasets, based on human-annotated ChemNER+ and CHEMET.
1 paper · 0 benchmarks
ChiQA is a dataset designed for visual question answering tasks that not only measures the relatedness but also measures the answerability, which demands more fine-grained vision and language reasoning.
1 paper · 0 benchmarks
ChinaOpen is a new video dataset targeted at open-world multimodal learning, with raw data gathered from Bilibili, a popular Chinese video-sharing website.
1 paper · 1 benchmark
Large-scale Chinese legal dataset for judgment prediction.
1 paper · 0 benchmarks
Chinese Literature NER RE is a Discourse-Level Named Entity Recognition and Relation Extraction Dataset for Chinese Literature Text.
1 paper · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
1 paper · 0 benchmarks
CinePile is a question-answering-based, long-form video understanding dataset.
1 paper · 1 benchmark
- An evaluation test bed for assessing the robustness of sentence embedding models against user-informed misinformation edits.
1 paper · 0 benchmarks
Lang-8 Preprocessed Dataset (for GED): - Dataset: Lang-8, a publicly available dataset containing user-generated content, primarily from second-language learners, focused on writing errors.
1 paper · 0 benchmarks
Clickbait PDFs (From Attachments to SEO: Click Here to Learn More about Clickbait PDFs!)
The paper presents a study of Clickbait PDFs, which are PDF documents leading to various attacks on the Web.
1 paper · 0 benchmarks
CoNECo (Complex Named Entity Corpus)
Complex Named Entity Corpus (CoNECo) is an annotated corpus for NER and NEN of protein-containing complexes.
1 paper · 0 benchmarks
Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts in 45 languages, generated by UDPipe (http://ufal.mff.cuni.cz/udpipe), together with word embeddings of dimension 100 computed from lowercased…
1 paper · 0 benchmarks
CoRAL dataset (CoRAL: a Context-aware Croatian Abusive Language Dataset)
CoRAL is a language and culturally aware Croatian Abusive dataset covering phenomena of implicitness and reliance on local and global context.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Code and Data for Replication of "Microsimulation Estimates of Decision Uncertainty and Value of Information Are Biased but Consistent" This is the full data set for replication of all results in the paper along with the R code for doing…
1 paper · 0 benchmarks
The dataset is specifically constructed for the library-oriented code generation task, which are constructed in the paper “CodeGen4Libs: A Two-Stage Approach for Library-Oriented Code Generation”.
1 paper · 0 benchmarks
InstructCoder is the first dataset designed to adapt LLMs for general code editing.
1 paper · 0 benchmarks
A diverse dataset of written code-switched productions, curated from topical threads of multiple bilingual communities on the Reddit discussion platform, and explore questions that were mainly addressed in the context of spoken language…
1 paper · 0 benchmarks
A large dataset of color names and their respective RGB values stores in CSV.
1 paper · 1 benchmark
ComSum is a data set of 7 million commit messages for text summarization.
1 paper · 0 benchmarks
Comet is a dataset which contains 11.5k user-assistant dialogs (totalling 103k utterances), grounded in simulated personal memory graphs.
1 paper · 0 benchmarks
CompMix-IR Dataset Overview: Characteristics: CompMix-IR is a heterogeneous knowledge retrieval benchmark dataset, featuring four knowledge types (text, knowledge graphs, tables, and infoboxes), 9,400+ QA pairs, and a corpus of 10 million…
1 paper · 0 benchmarks
The Composed Quora dataset consists of questions extracted from Quora that are grouped together if they are asking the same thing.
1 paper · 0 benchmarks
ConQA (Conceptual Query Answering)
ConQA is a dataset created using the intersection between VisualGenome and MS-COCO.
1 paper · 2 benchmarks
Concept-1K contains 1023 novel concepts from six domains, including economy, culture, science and technology, environment, education, and health and medical.
1 paper · 0 benchmarks
French sentences are sourced from Tatoeba repository and then translated into Congolese Swahili.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.