Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 64 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 3025–3072 of 3,130
The publicmeetings corpus contains meetings, made of pairs of automatic transcriptions from audio recordings and meeting reports written by a professional.
1 paper · 0 benchmarks
This dataset is made of 6366 threads collected from the r/AmITheAsshole community on Reddit.
1 paper · 0 benchmarks
Reader eye tracking and engagement scores for two short stories, aggregated by sentence.
1 paper · 0 benchmarks
sPBC (Super Parallel Bible Corpus)
This is a super-parallel Bible corpus containing 1401 language labels (languagescript pairs), meaning that for each verse, we have the translation in other languages.
1 paper · 0 benchmarks
satp-zsm-stage2 (Replication data for Crossing the Linguistic Causeway: Ethnonational Differences on Soundscape Attributes in Bahasa Melayu)
This is the replication data for the paper: "Crossing the Linguistic Causeway: Ethnonational Differences on Soundscape Attributes in Bahasa Melayu".
1 paper · 0 benchmarks
scb-mt-en-th-2020 is an English-Thai machine translation dataset with over 1 million segment pairs, curated from various sources, namely news, Wikipedia articles, SMS messages, task-based dialogs, web-crawled data and government documents.
1 paper · 0 benchmarks
These datasets, ComCo and SimCo, designed for evaluating multi-object representation in Vision-Language Models (VLMs).
1 paper · 0 benchmarks
The simply-CLEVR dataset aims to provide a benchmark dataset that can be used for transparent quantitative evaluation of explanation methods (aka heatmaps/XAI methods).
1 paper · 0 benchmarks
This dataset is a patched version of The Taste & Affect Music Database by D.
1 paper · 0 benchmarks
The datasets of "Time Interval-enhanced Graph Neural Network for Shared-account Cross-domain Sequential Recommendation" (TNNLs 2022)
1 paper · 0 benchmarks
Microscopy is a cornerstone of biomedical research, enabling detailed study of biological structures at multiple scales.
1 paper · 0 benchmarks
LLM-Based Vulnerability Classification in Police Narratives This repository contains datasets used in our research on applying large language models (LLMs) to identify indicators of vulnerability in police incident narratives.
1 paper · 0 benchmarks
VQA NLE synthetic dataset, made with LLaVA-1.5 using features from GQA dataset.
1 paper · 0 benchmarks
The dataset consists of 53,189 wikiHow articles across various categories of everyday tasks, 155,265 methods, and 772,294 steps with corresponding images.
1 paper · 1 benchmark
Here I provided the datasets I used for this analysis.
1 paper · 0 benchmarks
This is a stance detection dataset in the Zulu language.
1 paper · 0 benchmarks
ALFI (Annotations for Label-Free Images)
ALFI (Annotations for Label-Free Images) is a dataset of images and annotations for label-free microscopy imaging.
0 papers · 0 benchmarks
This dataset is described in the ALTA 2022 Shared Task and associated CodaLab competition.
0 papers · 0 benchmarks
ALTA 2023 Shared Task (Discriminate between human-authored and synthetic text generated by Large Language Models (LLMs))
This dataset is described in the ALTA 2023 Shared Task and associated CodaLab competition.
0 papers · 0 benchmarks
ARF (Artificial Relationships in Fiction)
Artificial Relationships in Fiction Dataset Description Artificial Relationships in Fiction (ARF) is a synthetically annotated dataset for Relation Extraction (RE) in fiction, created from a curated selection of literary texts sourced from…
0 papers · 0 benchmarks
The ASL-Phono introduces a novel linguistics-based representation, which describes the signs in the ASLLVD dataset in terms of a set of attributes of the American Sign Language phonology.
0 papers · 0 benchmarks
This open-source dataset consists of 5.04 hours of transcribed English conversational speech beyond telephony, where 13 conversations were contained.
0 papers · 0 benchmarks
ASSIN (Avaliação de Similaridade Semântica e INferência textual) is a dataset with semantic similarity score and entailment annotations.
0 papers · 0 benchmarks
Affective Text (Test Corpus of SemEval 2007) by Carlo Strapparava & Rada Mihalcea.
0 papers · 0 benchmarks
A new dataset for sentiment analysis, scraped from Allociné.fr user reviews.
0 papers · 0 benchmarks
AntM2C (Ant-Group Multi-Scenario Multi-Modal CTR dataset)
We release a large-scale Multi-Scenario Multi-Modal CTR dataset named AntM2C, built from real industrial data from Alipay.
0 papers · 0 benchmarks
ArcBench is a logically challenging dataset of 158 English question–answer pairs, derived from the RoR-Bench benchmark.
0 papers · 0 benchmarks
The Arena-Hard benchmark is a high-quality benchmarking tool for Language Learning Models (LLMs) developed by LMSYS Org¹.
0 papers · 0 benchmarks
This dataset is comprised of the dynamic analysis reports generated by CAPEv2, from both malware and goodware.
0 papers · 0 benchmarks
Datasets for Bangla Natural Language Processing tasks.
0 papers · 0 benchmarks
A Filipino multi-modal language dataset for text+visual tasks.
0 papers · 0 benchmarks
The dataset consists of 3265 text samples corresponding to the concatenation of lines spoken by fictional characters.
0 papers · 0 benchmarks
Medical report generation (MRG), which aims to automatically generate a textual description of a specific medical image (e.g., a chest X-ray), has recently received increasing research interest.
0 papers · 0 benchmarks
We introduce CCI4.0, a large-scale bilingual pre-training dataset engineered for superior data quality and diverse human-like reasoning trajectory.
0 papers · 0 benchmarks
CIDII Dataset (Correct Information and Disinformation about Islamic Issues)
The CIDII dataset is a binary classification, consisting of two classes of correct information and disinformation related to Islamic issues.
0 papers · 0 benchmarks
COCO-Facet is a benchmark for attribute-focused text-to-image retrieval, comprising 9,112 queries with 100 candidate images for each.
0 papers · 0 benchmarks
With the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic.
0 papers · 0 benchmarks
The CTV-Dataset (CTV stands for Cyclist Top-View) is a trajectories dataset for cyclist behaviour in mixed-traffic environments (aka.
0 papers · 0 benchmarks
We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource.
0 papers · 0 benchmarks
This is a dataset of paraphrases created by ChatGPT.
0 papers · 0 benchmarks
Dataset Description Our dataset contains questions from a well-known software testing book Introduction to Software Testing 2nd Edition by Ammann and Offutt.
0 papers · 0 benchmarks
CodeSCAN (ScreenCast ANalysis for Video Programming Tutorials)
CodeSCAN is the first large-scale and diverse dataset of coding screenshots with pixel-perfect annotations.
0 papers · 0 benchmarks
This dataset includes CSV files that contain IDs and sentiment scores of the tweets related to the COVID-19 pandemic.
0 papers · 0 benchmarks
The Couples Therapy corpus contains audio, video recordings and manual transcriptions of conversations between 134 real-life couples attending marital therapy.
0 papers · 0 benchmarks
The data originate from the journalistic domain in the Czech language.
0 papers · 0 benchmarks
Dialog System Technology Challenges 8 (DSTC) Track 2 builds on the success of DSTC 7 Track 1 (NOESIS: Noetic End-to-End Response Selection Challenge).
0 papers · 0 benchmarks
The DUC 2005 data set is a dataset for summarization which consists of 50 document collections of 25 documents each; each document collection includes a human-written query.
0 papers · 0 benchmarks
DigiLeTs (Digit- and Letter Trajectories)
A dataset with 23 870 digital trajectories (i.e.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.