Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 16 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 721–768 of 3,130

ToolQA is a question answering benchmark for Large Language Models (LLMs) which is designed to faithfully evaluate LLMs' ability to use external tools for question answering.
16 papers · 0 benchmarks
TripClick is a large-scale dataset of click logs in the health domain, obtained from user interactions of the Trip Database health web search engine.
16 papers · 0 benchmarks
The TweepFake dataset consists of 25,572 social media messages posted either by bots or humans on Twitter.
16 papers · 1 benchmark
The ViGGO corpus is a set of 6,900 meaning representation to natural language utterance pairs in the video game domain.
16 papers · 1 benchmark
Wukong is a large-scale Chinese cross-modal dataset for benchmarking different multi-modal pre-training methods to facilitate the Vision-Language Pre-training (VLP).
16 papers · 0 benchmarks
X-FACT is a large publicly available multilingual dataset for factual verification of naturally existing real-world claims.
16 papers · 0 benchmarks
e-SNLI-VE is a large VL (vision-language) dataset with NLEs (natural language explanations) with over 430k instances for which the explanations rely on the image content.
16 papers · 2 benchmarks
ArSarcasm-v2 is an extension of the original ArSarcasm dataset published along with the paper From Arabic Sentiment Analysis to Sarcasm Detection: The ArSarcasm Dataset.
15 papers · 0 benchmarks
CMU DoG (CMU Document Grounded Conversations Dataset)
This is a document grounded dataset for text conversations.
15 papers · 0 benchmarks
CPED (Chinese Personalized and Emotional Dialogue)
We construct a dataset named CPED from 40 Chinese TV shows.
15 papers · 3 benchmarks
CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts generated using the proposed semi-automatic text generation strategy.
15 papers · 0 benchmarks
A large-scale video dataset, featuring clips from movies with detailed captions.
15 papers · 1 benchmark
The beginnings of a question answering dataset specifically designed for COVID-19, built by hand from knowledge gathered from Kaggle's COVID-19 Open Research Dataset Challenge.
15 papers · 0 benchmarks
CrossNER is a cross-domain NER (Named Entity Recognition) dataset, a fully-labeled collection of NER data spanning over five diverse domains (Politics, Natural Science, Music, Literature, and Artificial Intelligence) with specialized…
15 papers · 1 benchmark
The DUC2004 dataset is a dataset for document summarization.
15 papers · 4 benchmarks
FaVIQ (Fact Verification from Information-seeking Questions)
FaVIQ (Fact Verification from Information-seeking Questions) is a challenging and realistic fact verification dataset that reflects confusions raised by real users.
15 papers · 0 benchmarks
Fakeddit is a novel multimodal dataset for fake news detection consisting of over 1 million samples from multiple categories of fake news.
15 papers · 0 benchmarks
Humicroedit is a humorous headline dataset.
15 papers · 0 benchmarks
JEEBench is a considerably more challenging benchmark dataset for evaluating the problem solving abilities of LLMs.
15 papers · 0 benchmarks
MMHS150k (Multimodal Hate Speech)
Existing hate speech datasets contain only textual data.
15 papers · 0 benchmarks
MedVidQA (Medical Video Question Answering)
The MedVidQA dataset contains the collection of 3, 010 manually created health-related questions and timestamps as visual answers to those questions from trusted video sources, such as accredited medical schools with an established…
15 papers · 0 benchmarks
MuCGEC (Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction)
MuCGEC is a multi-reference multi-source evaluation dataset for Chinese Grammatical Error Correction (CGEC), consisting of 7,063 sentences collected from three different Chinese-as-a-Second-Language (CSL) learner sources.
15 papers · 1 benchmark
Large-scale American Sign Language (ASL) - English dataset collected from online video sites (e.g., YouTube).
15 papers · 0 benchmarks
Opusparcus is a paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish.
15 papers · 0 benchmarks
Pile of Law is a ∼256GB (and growing) dataset of legal and administrative data which can be used for assessing norms on data sanitization across legal and administrative settings.
15 papers · 0 benchmarks
Project CodeNet is a large-scale dataset with approximately 14 million code samples, each of which is an intended solution to one of 4000 coding problems.
15 papers · 0 benchmarks
PubTables-1M (PubMed Tables One Million)
The goal of PubTables-1M is to create a large, detailed, high-quality dataset for training and evaluating a wide variety of models for the tasks of table detection, table structure recognition, and functional analysis.
15 papers · 0 benchmarks
A large-scale dataset for retrieval and event localisation in video.
15 papers · 1 benchmark
A Benchmark for Robust Multi-Hop Spatial Reasoning in Texts
15 papers · 1 benchmark
xCodeEval is one of the largest executable multilingual multitask benchmarks consisting of 17 programming languages with execution-level parallelism.
15 papers · 0 benchmarks
AVSD (Audio-Visual Scene-Aware Dialog)
The Audio Visual Scene-Aware Dialog (AVSD) dataset, or DSTC7 Track 3, is a audio-visual dataset for dialogue understanding.
14 papers · 1 benchmark
AmazonQA consists of 923k questions, 3.6M answers and 14M reviews across 156k products.
14 papers · 0 benchmarks
ArSarcasm is a new Arabic sarcasm detection dataset.
14 papers · 0 benchmarks
BabyLM is a dataset for small scale language modeling, human language acquisition, low-resource NLP, and cognitive modeling.
14 papers · 0 benchmarks
The Collaborative Drawing game (CoDraw) dataset contains ~10K dialogs consisting of ~138K messages exchanged between human players in the CoDraw game.
14 papers · 0 benchmarks
A SemEval shared task in which participants must extract definitions from free text using a term-definition pair corpus that reflects the complex reality of definitions in natural language.
14 papers · 0 benchmarks
The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages.
14 papers · 0 benchmarks
DuLeMon (Baidu Long-term Memory Conversation)
DuLeMon is a large-scale Chinese Long-term Memory Conversation dataset, which simulates long-term memory conversations and focuses on the ability to actively construct and utilize the user's and the bot's persona in a long-term interaction.
14 papers · 0 benchmarks
GlobalOpinionQA consists of questions and answers from cross-national surveys designed to capture diverse opinions on global issues across different countries.
14 papers · 0 benchmarks
Inter-X is a large-scale dataset containing ~11K interaction sequences, more than 8.1M frames and 34K fine-grained human textual descriptions.
14 papers · 1 benchmark
KUAKE-QIC (Query Intent Classification Dataset)
KUAKE Query Intent Classification, a dataset for intent classification, is used for the KUAKE-QIC task.
14 papers · 1 benchmark
M3KE (Massive Multi-Level Multi-Subject Knowledge Evaluation Benchmark)
M3KE is a Massive Multi-Level Multi-Subject Knowledge Evaluation benchmark, which is developed to measure knowledge acquired by Chinese large language models by testing their multitask accuracy in zero- and few-shot settings.
14 papers · 0 benchmarks
MAVE (MAVE: : A Product Dataset for Multi-source Attribute Value Extraction)
The dataset contains 3 million attribute-value annotations across 1257 unique categories created from 2.2 million cleaned Amazon product profiles.
14 papers · 2 benchmarks
MIntRec is a novel dataset for multimodal intent recognition.
14 papers · 1 benchmark
MedDG is a large-scale high-quality Medical Dialogue dataset related to 12 types of common Gastrointestinal diseases.
14 papers · 0 benchmarks
With the same format as WikiHop, the MedHop dataset is based on research paper abstracts from PubMed, and the queries are about interactions between pairs of drugs.
14 papers · 0 benchmarks
OVAD benchmark (Open-Vocabulary Attribute Detection)
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner.
14 papers · 3 benchmarks
OVBench is a benchmark tailored for real-time video understanding: - Memory, Perception, and Prediction of Temporal Contexts: Questions are framed to reference the present state of entities, requiring models to memorize/perceive/predict…
14 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.