Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 51 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2401–2448 of 3,130

The goal of InfoLossQA is to generate a series of QA pairs that reveal to lay readers what information a simplified text lacks compared to its original.
1 paper · 0 benchmarks
InstructOpenWiki is a substantial instruction tuning dataset for Open-world IE enriched with a comprehensive corpus, extensive annotations, and diverse instructions.
1 paper · 0 benchmarks
A dataset for image editing containing >450k samples of: 1.
1 paper · 0 benchmarks
For the purpose of training and evaluating our intent classification model for electric automation, we curated a dataset consisting of intent-based user instructions.
1 paper · 0 benchmarks
A dataset of sentence pairs annotated following the formalization.
1 paper · 0 benchmarks
IoT-23 (IoT-23: A labeled dataset with malicious and benign IoT network traffic)
IoT-23 is a dataset of network traffic from Internet of Things (IoT) devices.
1 paper · 0 benchmarks
The dataset includes source code vulnerabilities in some of the most commonly used IoT frameworks.
1 paper · 0 benchmarks
Text from Irish Wikipedia, an online encyclopedia.
1 paper · 0 benchmarks
JDDC 2.0 is a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform JD.com, containing about 246 thousand dialogue sessions, 3 million utterances, and 507 thousand images, along with…
1 paper · 0 benchmarks
JECC (Jericho Environment Commonsense Comprehension)
Jericho Environment Commonsense Comprehension (JECC) is a dataset for commonsense reasoning.
1 paper · 0 benchmarks
The JPersonaChat dataset is built by NTT CS LAB for Japanese dialog transformers models.
1 paper · 0 benchmarks
JSS Dataset (Jejueo Single Speaker Speech)
The Jejueo Single Speaker Speech (JSS) dataset consists of 10k high-quality audio files recorded by a native Jejueo speaker and a transcript file.
1 paper · 0 benchmarks
The information contained in JUSThink Dialogue and Actions Corpus dataset includes dialogue transcripts, event logs, and test responses of children aged 9 through 12, as they participate in a robot-mediated human-human collaborative…
1 paper · 0 benchmarks
📊 Dataset Details - Name: JamendoMaxCaps - URL: https://huggingface.co/datasets/amaai-lab/JamendoMaxCaps - Content: 362,238 songs with captions generated by Qwen2-Audio Metadata Fields - genre - speed - variable tags 🎯 Rationale 1.
1 paper · 0 benchmarks
Korean Multi-label Hate Speech Dataset We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns.
1 paper · 0 benchmarks
K-QA(fa) (persian translation of K-QA dataset)
persian translation of K-QA dataset
1 paper · 0 benchmarks
KCIF (Knowledge Conditioned Instruction Following (KCIF))
KCIF is a benchmark for evaluating the instruction-following capabilities of Large Language Models (LLM).
1 paper · 0 benchmarks
KD-EmoR (Korean Drama Scene Transcript Dataset for Emotion Recognition in Conversations)
KD-EmoR is socio-behavioral emotion dataset for emotion recognition in realistic conversation scenarios.
1 paper · 1 benchmark
KGRED (Knowledge-graph-enhanced relation extraction datasets--)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Dataset contains CS/Math articles abstracts (in Russian) obtained from two online sources.
1 paper · 0 benchmarks
The Kite database is a multi-modal dataset for the control of unmanned aerial vehicles (UAVs).
1 paper · 0 benchmarks
We introduce KnowledJe, an English-language knowledge graph of antisemitic history and language from the 20th century to the present.
1 paper · 0 benchmarks
Kor-Lang8 (Lang-8 Korean Corpus)
Kor-Lang8 is a Korean grammatical error correction (GEC) dataset extracted from the NAIST Lang-8 Learner Corpora by the language label.
1 paper · 0 benchmarks
Kor-Learner (Korean Learner Corpus)
Kor-Learner is a Korean grammatical error correction (GEC) dataset made from the NIKL learner corpus containing essays written by Korean learners and their grammatical error correction annotations by their tutors in an morpheme-level XML…
1 paper · 0 benchmarks
Kor-Native (Native Korean Corpus)
Kor-Learner is a Korean grammatical error correction (GEC) dataset collected grammatically from two sources, and the correct sentences were read using Google Text-to-Speech(TTS) system.
1 paper · 0 benchmarks
APEACH is the first crowd-generated Korean evaluation dataset for hate speech detection.
1 paper · 0 benchmarks
Korean UnSmile Dataset (SmilegateAI Korean UnSmile Dataset)
1.9K Korean Online Hate Speech Comments for Multilabel Classification (Annotated by Three Independent Labelers per Data)
1 paper · 0 benchmarks
L3Cube-MahaCorpus is a Marathi monolingual data set scraped from different internet sources.
1 paper · 0 benchmarks
Raw negotiation transcripts generated for the paper "Evaluating Language Model Agency through Negotiations".
1 paper · 0 benchmarks
For our experiments, we collected a dataset of procedural knowledge of the LangChain Python library, unseen by many extant LLMs.
1 paper · 0 benchmarks
The datasets of "Towards Lightweight Cross-domain Sequential Recommendation via External Attention-enhanced Graph Convolution Network" (DASFAA 2023)
1 paper · 0 benchmarks
LEARNING STYLE IDENTIFICATION (Learning Style Identification Using Semi-Supervised Self-Taught Labeling)
The dataset was collected from two courses offered on the University of Jordan's E-learning Portal during the second semester of 2020, namely "Computer Skills for Humanities Students" (CSHS) and "Computer Skills for Medical Students"…
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
LIGHT-Quests is an extension of LIGHT, a large-scale crowd-sourced fantasy text-game, to generate a dataset of quests.
1 paper · 0 benchmarks
This dataset comprises high-quality, targeted spear-phishing emails created using a proprietary system that harnesses the power of LLMs and knowledge graphs.
1 paper · 0 benchmarks
LLM Health Benchmarks (LLM Health Benchmarks - Yesil Science)
LLM Health Benchmarks Dataset The Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties.
1 paper · 0 benchmarks
To evaluate our proposed strategy of asynchronous communication for LLMs, we run games of Mafia with human players, incorporating an LLM-based agent as an additional player, within an asynchronous chat environment.
1 paper · 0 benchmarks
Three tasks were addressed in the LLMs4OL paradigm.
1 paper · 0 benchmarks
LLaVA-Rad MIMIC-CXR features more accurate section extractions from MIMIC-CXR free-text radiology reports.
1 paper · 0 benchmarks
Are Large Pre-Trained Language Models Leaking Your Personal Information?
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LPSC (Planetary Science Data Set)
This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998 to 2020 of relevance to four Mars missions.
1 paper · 2 benchmarks
LSA-T (Lengua de Señas Argentina - Traducción)
LSA-T is the first continuous Argentinian Sign Language (LSA) dataset.
1 paper · 1 benchmark
LSDBench (Long-video Sampling Dilemma Benchmark)
A benchmark that focuses on the sampling dilemma in long-video tasks.
1 paper · 0 benchmarks
LSFB Datasets (French Belgian Sign Language Datasets)
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur.
1 paper · 0 benchmarks
LSICC (Large Scale Informal Chinese Corpus)
Large Scale Informal Chinese Corpus (LSICC) is a large-scale corpus of informal Chinese.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.