Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 51 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 2401–2448 of 3,130
The goal of InfoLossQA is to generate a series of QA pairs that reveal to lay readers what information a simplified text lacks compared to its original.
1 paper · 0 benchmarks
InstructOpenWiki is a substantial instruction tuning dataset for Open-world IE enriched with a comprehensive corpus, extensive annotations, and diverse instructions.
1 paper · 0 benchmarks
A dataset for image editing containing >450k samples of: 1.
1 paper · 0 benchmarks
For the purpose of training and evaluating our intent classification model for electric automation, we curated a dataset consisting of intent-based user instructions.
1 paper · 0 benchmarks
A dataset of sentence pairs annotated following the formalization.
1 paper · 0 benchmarks
IoT-23 (IoT-23: A labeled dataset with malicious and benign IoT network traffic)
IoT-23 is a dataset of network traffic from Internet of Things (IoT) devices.
1 paper · 0 benchmarks
The dataset includes source code vulnerabilities in some of the most commonly used IoT frameworks.
1 paper · 0 benchmarks
Text from Irish Wikipedia, an online encyclopedia.
1 paper · 0 benchmarks
JDDC 2.0 is a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform JD.com, containing about 246 thousand dialogue sessions, 3 million utterances, and 507 thousand images, along with…
1 paper · 0 benchmarks
JECC (Jericho Environment Commonsense Comprehension)
Jericho Environment Commonsense Comprehension (JECC) is a dataset for commonsense reasoning.
1 paper · 0 benchmarks
The JPersonaChat dataset is built by NTT CS LAB for Japanese dialog transformers models.
1 paper · 0 benchmarks
The Jejueo Single Speaker Speech (JSS) dataset consists of 10k high-quality audio files recorded by a native Jejueo speaker and a transcript file.
1 paper · 0 benchmarks
The information contained in JUSThink Dialogue and Actions Corpus dataset includes dialogue transcripts, event logs, and test responses of children aged 9 through 12, as they participate in a robot-mediated human-human collaborative…
1 paper · 0 benchmarks
📊 Dataset Details - Name: JamendoMaxCaps - URL: https://huggingface.co/datasets/amaai-lab/JamendoMaxCaps - Content: 362,238 songs with captions generated by Qwen2-Audio Metadata Fields - genre - speed - variable tags 🎯 Rationale 1.
1 paper · 0 benchmarks
JustLogic is a natural language deductive reasoning dataset.
1 paper · 0 benchmarks
Korean Multi-label Hate Speech Dataset We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns.
1 paper · 0 benchmarks
K-QA(fa) (persian translation of K-QA dataset)
persian translation of K-QA dataset
1 paper · 0 benchmarks
KCIF (Knowledge Conditioned Instruction Following (KCIF))
KCIF is a benchmark for evaluating the instruction-following capabilities of Large Language Models (LLM).
1 paper · 0 benchmarks
KD-EmoR (Korean Drama Scene Transcript Dataset for Emotion Recognition in Conversations)
KD-EmoR is socio-behavioral emotion dataset for emotion recognition in realistic conversation scenarios.
1 paper · 1 benchmark
KGRED (Knowledge-graph-enhanced relation extraction datasets--)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Dataset contains CS/Math articles abstracts (in Russian) obtained from two online sources.
1 paper · 0 benchmarks
The Kite database is a multi-modal dataset for the control of unmanned aerial vehicles (UAVs).
1 paper · 0 benchmarks
We introduce KnowledJe, an English-language knowledge graph of antisemitic history and language from the 20th century to the present.
1 paper · 0 benchmarks
Kor-Lang8 is a Korean grammatical error correction (GEC) dataset extracted from the NAIST Lang-8 Learner Corpora by the language label.
1 paper · 0 benchmarks
Kor-Learner is a Korean grammatical error correction (GEC) dataset made from the NIKL learner corpus containing essays written by Korean learners and their grammatical error correction annotations by their tutors in an morpheme-level XML…
1 paper · 0 benchmarks
Kor-Learner is a Korean grammatical error correction (GEC) dataset collected grammatically from two sources, and the correct sentences were read using Google Text-to-Speech(TTS) system.
1 paper · 0 benchmarks
APEACH is the first crowd-generated Korean evaluation dataset for hate speech detection.
1 paper · 0 benchmarks
1.9K Korean Online Hate Speech Comments for Multilabel Classification (Annotated by Three Independent Labelers per Data)
1 paper · 0 benchmarks
L3Cube-MahaCorpus is a Marathi monolingual data set scraped from different internet sources.
1 paper · 0 benchmarks
Raw negotiation transcripts generated for the paper "Evaluating Language Model Agency through Negotiations".
1 paper · 0 benchmarks
For our experiments, we collected a dataset of procedural knowledge of the LangChain Python library, unseen by many extant LLMs.
1 paper · 0 benchmarks
The datasets of "Towards Lightweight Cross-domain Sequential Recommendation via External Attention-enhanced Graph Convolution Network" (DASFAA 2023)
1 paper · 0 benchmarks
The dataset was collected from two courses offered on the University of Jordan's E-learning Portal during the second semester of 2020, namely "Computer Skills for Humanities Students" (CSHS) and "Computer Skills for Medical Students"…
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
LIGHT-Quests is an extension of LIGHT, a large-scale crowd-sourced fantasy text-game, to generate a dataset of quests.
1 paper · 0 benchmarks
This dataset comprises high-quality, targeted spear-phishing emails created using a proprietary system that harnesses the power of LLMs and knowledge graphs.
1 paper · 0 benchmarks
LLM Health Benchmarks Dataset The Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties.
1 paper · 0 benchmarks
To evaluate our proposed strategy of asynchronous communication for LLMs, we run games of Mafia with human players, incorporating an LLM-based agent as an additional player, within an asynchronous chat environment.
1 paper · 0 benchmarks
Three tasks were addressed in the LLMs4OL paradigm.
1 paper · 0 benchmarks
LLaVA-Rad MIMIC-CXR features more accurate section extractions from MIMIC-CXR free-text radiology reports.
1 paper · 0 benchmarks
Are Large Pre-Trained Language Models Leaking Your Personal Information?
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
LPSC (Planetary Science Data Set)
This data set contains annotated text versions of 1635 two-page abstracts published at the Lunar and Planetary Science Conference from 1998 to 2020 of relevance to four Mars missions.
1 paper · 2 benchmarks
LSA-T (Lengua de Señas Argentina - Traducción)
LSA-T is the first continuous Argentinian Sign Language (LSA) dataset.
1 paper · 1 benchmark
LSDBench (Long-video Sampling Dilemma Benchmark)
A benchmark that focuses on the sampling dilemma in long-video tasks.
1 paper · 0 benchmarks
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur.
1 paper · 0 benchmarks
LSICC (Large Scale Informal Chinese Corpus)
Large Scale Informal Chinese Corpus (LSICC) is a large-scale corpus of informal Chinese.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.