Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 19 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 865–912 of 3,130

Talk The Walk is a large-scale dialogue dataset grounded in action and perception.
11 papers · 0 benchmarks
TaxiNLI is a dataset collected based on the principles and categorizations of the aforementioned taxonomy.
11 papers · 0 benchmarks
TimeDial presents a crowdsourced English challenge set, for temporal commonsense reasoning, formulated as a multiple choice cloze task with around 1.5k carefully curated dialogs.
11 papers · 0 benchmarks
VLEP (Video-and-Language Event Prediction)
VLEP contains 28,726 future event prediction examples (along with their rationales) from 10,234 diverse TV Show and YouTube Lifestyle Vlog video clips.
11 papers · 1 benchmark
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
11 papers · 9 benchmarks
methods2test is a supervised dataset consisting of Test Cases and their corresponding Focal Methods from a set of Java software repositories.
11 papers · 0 benchmarks
nvBench is a large-scale NL2VIS (natural languagge to visualisations) benchmark, containing 25,750 (NL, VIS) pairs from 750 tables over 105 domains, synthesized from (NL, SQL) benchmarks to support cross-domain NLPVIS (Natural Language…
11 papers · 0 benchmarks
ABCD (Action-Based Conversations Dataset)
10 papers · 1 benchmark
AnnoMI: A Dataset of Expert-Annotated Counselling Dialogues Dataset Introduction Research on natural language processing approaches to analysing counselling dialogues has seen substantial development in recent years, but access to this…
10 papers · 0 benchmarks
BiRD (Bigram Relatedness Dataset)
Bigram Relatedness Dataset (BiRD) is a large, fine-grained, bigram relatedness dataset, using a comparative annotation technique called Best Worst Scaling.
10 papers · 0 benchmarks
CHALET (Cornell House Agent Learning Environment)
CHALET is a 3D house simulator with support for navigation and manipulation.
10 papers · 0 benchmarks
CJRC (Chinese judicial reading comprehension)
The Chinese judicial reading comprehension (CJRC) dataset contains approximately 10K documents and almost 50K questions with answers.
10 papers · 0 benchmarks
CLEVR-Dialog is a large diagnostic dataset for studying multi-round reasoning in visual dialog.
10 papers · 0 benchmarks
CMeEE (Chinese Medical Named Entity Recognition Dataset)
Chinese Medical Named Entity Recognition, a dataset first released in CHIP20204, is used for CMeEE task.
10 papers · 1 benchmark
Chart2Text (Chart Summarization Dataset)
Chart2Text is a dataset that was crawled from 23,382 freely accessible pages from statista.com in early March of 2020, yielding a total of 8,305 charts, and associated summaries.
10 papers · 0 benchmarks
ChatHaruhi (ChatHaruhi: Reviving Anime Character in Reality via Large Language Model)
ChatHaruhi is a dataset covering 32 Chinese / English TV / anime characters with over 54k simulated dialogues.
10 papers · 0 benchmarks
Consist of 23,533 statements extracted from all U.S.
10 papers · 0 benchmarks
CriticBench is a comprehensive benchmark designed to assess the abilities of Large Language Models (LLMs) to critique and rectify their reasoning across various tasks.
10 papers · 0 benchmarks
Description Detection Dataset (D³, /dikju:b/) is an attempt at creating a next-generation object detection dataset.
10 papers · 1 benchmark
DiscoFuse was created by applying a rule-based splitting method on two corpora - sports articles crawled from the Web, and Wikipedia.
10 papers · 0 benchmarks
The Discovery datasets consists of adjacent sentence pairs (s1,s2) with a discourse marker (y) that occurred at the beginning of s2.
10 papers · 1 benchmark
This is the dataset for the 2020 Duolingo shared task on Simultaneous Translation And Paraphrase for Language Education (STAPLE).
10 papers · 0 benchmarks
EmoWOZ is the first large-scale open-source dataset for emotion recognition in task-oriented dialogues.
10 papers · 2 benchmarks
FM-IQA (Freestyle Multilingual Image Question Answering)
FM-IQA is a question-answering dataset containing over 150,000 images and 310,000 freestyle Chinese question-answer pairs and their English translations.
10 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
10 papers · 0 benchmarks
Fig-QA consists of 10256 examples of human-written creative metaphors that are paired as a Winograd schema.
10 papers · 0 benchmarks
The QMUL underGround Re-IDentification (GRID) dataset contains 250 pedestrian image pairs.
10 papers · 5 benchmarks
Griddly is an environment for grid-world based research.
10 papers · 0 benchmarks
LC-QuAD (Largescale Complex Question Answering Dataset)
LC-QuAD is a Large Question Answering dataset with 30,000 pairs of questions and its corresponding SPARQL query.
10 papers · 1 benchmark
LectureBank Dataset is a manually collected dataset of lecture slides.
10 papers · 0 benchmarks
MACS (Multi-Annotator Captioned Soundscapes)
This is a dataset containing audio captions and corresponding audio tags for a number of 3930 audio files of the TAU Urban Acoustic Scenes 2019 development dataset (airport, public square, and park).
10 papers · 0 benchmarks
MCoNaLa (Multilingual CoNaLa)
MCoNaLa is a multilingual dataset to benchmark code generation from natural language commands extending beyond English.
10 papers · 0 benchmarks
MP-DocVQA (Multipage Document Visual Question Answering)
The dataset is aimed to perform Visual Question Answering on multipage industry scanned documents.
10 papers · 0 benchmarks
MUGEN is a large-scale video-audio-text dataset MUGEN, collected using the open-sourced platform game CoinRun.
10 papers · 0 benchmarks
MUStARD++ is a multimodal sarcasm detection dataset (MUStARD) pre-annotated with 9 emotions.
10 papers · 1 benchmark
A large-scale dataset that consists of 21,184 claims, where each claim is assigned a truthfulness label and ruling statement, with 58,523 pieces of evidence in the form of text and images.
10 papers · 0 benchmarks
Multi-XScience is a large-scale dataset for multi-document summarization of scientific articles.
10 papers · 0 benchmarks
NLI4CT dataset consists of 2,400 annotated statements with accompanying labels, CTRs, and evidence.
10 papers · 0 benchmarks
OCW (Only Connect Wall Dataset and creative problem solving tasks)
The OCW dataset is for evaluating creative problem solving tasks by curating the problems and human performance results from the popular British quiz show Only Connect.
10 papers · 1 benchmark
OpenMEVA is a benchmark for evaluating open-ended story generation metrics.
10 papers · 0 benchmarks
PointQA is a set of datasets for Visual Question Datasets (VQA) that require a pointer to an object in the image to be answered correctly.
10 papers · 0 benchmarks
REFinD (REFinD: Relation Extraction Financial Dataset)
REFinD is a large-scale annotated dataset of relations, with ∼29K instances and 22 relations amongst 8 types of entity pairs, generated entirely over financial documents.
10 papers · 0 benchmarks
Rainbow is multi-task benchmark for common-sense reasoning that uses different existing QA datasets: aNLI, Cosmos QA, HellaSWAG.
10 papers · 0 benchmarks
ReQA (Retrieval Question-Answering)
Retrieval Question-Answering (ReQA) benchmark tests a model’s ability to retrieve relevant answers efficiently from a large set of documents.
10 papers · 0 benchmarks
ReaSCAN (ReaSCAN: Compositional Reasoning in Language Grounding)
ReaSCAN is a synthetic navigation task that requires models to reason about surroundings over syntactically difficult languages.
10 papers · 0 benchmarks
Rico is a public UI corpus with 72K Android UI screens mined from 9.7K Android apps (Deka et al., 2017).
10 papers · 0 benchmarks
SST-3 (Stanford Sentiment Treebank: 3-way)
SST-5 is the Stanford Sentiment Treebank 5-way classification dataset (positive, somewhat positive, neutral, somewhat negative, negative).
10 papers · 1 benchmark
SkillSpan (Hard and Soft Skill Extraction from English Job Postings)
SkillSpan is a dataset for Skill Extraction (SE).
10 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.