Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 15 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 673–720 of 3,130
The first parallel corpus composed from United Nations documents published by the original data creator.
18 papers · 0 benchmarks
VAST (VAried Stance Topics)
VAST consists of a large range of topics covering broad themes, such as politics (e.g., ‘a Palestinian state’), education (e.g., ‘charter schools’), and public health (e.g., ‘childhood vaccination’).
18 papers · 1 benchmark
Violin (VIdeO-and-Language INference)
Video-and-Language Inference is the task of joint multimodal understanding of video and text.
18 papers · 0 benchmarks
Weibo21 is a benchmark of fake news dataset for multi-domain fake news detection (MFND) with domain label annotated, which consists of 4,488 fake news and 4,640 real news from 9 different domains.
18 papers · 0 benchmarks
X-CSQA is a multilingual dataset for Commonsense reasoning research, based on CSQA.
18 papers · 0 benchmarks
xSID (Cross-lingual Slot and Intent Detection)
xSID, a new evaluation benchmark for cross-lingual (X) Slot and Intent Detection in 13 languages from 6 language families, including a very low-resource dialect, covering Arabic (ar), Chinese (zh), Danish (da), Dutch (nl), English (en),…
18 papers · 0 benchmarks
AVeriTeC (AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the Web)
AVeriTeC (Automated Verification of Textual Claims) is a dataset of 4568 real-world claims covering fact-checks by 50 different organizations.
17 papers · 1 benchmark
Extracted from the Tashkeela Corpus, the dataset consists of 55K lines containing about 2.3M words.
17 papers · 1 benchmark
BUG is a large-scale gender bias dataset of 108K diverse real-world English sentences, sampled semiautomatically from large corpora using lexical syntactic pattern matching
17 papers · 0 benchmarks
CLEVR-Ref+ is a synthetic diagnostic dataset for referring expression comprehension.
17 papers · 1 benchmark
COCO-Noisy (Microsoft Common Objects in Context with 20% of Noisy Correspondence and 1K test data)
This dataset is based on MS COCO that have 20% of data randomly shuffled to simulate noisy correspondence.
17 papers · 1 benchmark
ConvQuestions is the first realistic benchmark for conversational question answering over knowledge graphs.
17 papers · 0 benchmarks
DialoGLUE is a natural language understanding benchmark for task-oriented dialogue designed to encourage dialogue research in representation-based transfer, domain adaptation, and sample-efficient task learning.
17 papers · 2 benchmarks
FreebaseQA is a data set for open-domain QA over the Freebase knowledge graph.
17 papers · 0 benchmarks
By perturbing the widely used GSM8K dataset, an adversarial dataset for grade-school math called GSM-Plus is created.
17 papers · 1 benchmark
JESC (Japanese-English Subtitle Corpus)
Japanese-English Subtitle Corpus is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue.
17 papers · 0 benchmarks
Kleister NDA is a dataset for Key Information Extraction (KIE).
17 papers · 1 benchmark
LABR (Large-Scale Arabic Book Reviews)
LABR is a large sentiment analysis dataset to-date for the Arabic language.
17 papers · 1 benchmark
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
17 papers · 1 benchmark
MMDialog is a large-scale multi-turn dialogue dataset containing multi-modal open-domain conversations derived from real human-human chat content in social media.
17 papers · 1 benchmark
NoReC (Norwegian Review Corpus)
The Norwegian Review Corpus (NoReC) was created for the purpose of training and evaluating models for document-level sentiment analysis.
17 papers · 0 benchmarks
QAMR (Question-Answer Meaning Representation Dataset)
Question-Answer Meaning Representation (QAMR) represents a predicate-argument structure of a sentence with a set of question-answer pairs, so that annotations can be easily provided by non-experts.
17 papers · 0 benchmarks
Recent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online.
17 papers · 0 benchmarks
Quasimodo is commonsense knowledge base that focuses on salient properties of objects.
17 papers · 0 benchmarks
How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence.
17 papers · 2 benchmarks
VQA-E is a dataset for Visual Question Answering with Explanation, where the models are required to generate and explanation with the predicted answer.
17 papers · 0 benchmarks
A temporal counterfactual dataset composing of 1000 short and natural video-caption pairs.
17 papers · 1 benchmark
iSarcasm is a dataset of tweets, each labelled as either sarcastic or nonsarcastic.
17 papers · 1 benchmark
CMU-MOSI (Multimodal Corpus of Sentiment Intensity)
The Multimodal Corpus of Sentiment Intensity (CMU-MOSI) dataset is a collection of 2199 opinion video clips.
16 papers · 2 benchmarks
ChemProt consists of 1,820 PubMed abstracts with chemical-protein interactions annotated by domain experts and was used in the BioCreative VI text mining chemical-protein interactions shared task.
16 papers · 1 benchmark
FaithDial is a new benchmark for hallucination-free dialogues, by editing hallucinated responses in the Wizard of Wikipedia (WoW) benchmark.
16 papers · 0 benchmarks
GRIT (General Robust Image Task Benchmark)
The General Robust Image Task (GRIT) Benchmark is an evaluation-only benchmark for evaluating the performance and robustness of vision systems across multiple image prediction tasks, concepts, and data sources.
16 papers · 5 benchmarks
GeneCIS benchmark is designed for measuring models’ ability to adapt to a range of similarity conditions, which is zero-shot evaluation only.
16 papers · 1 benchmark
HumanEval-X is a benchmark for evaluating the multilingual ability of code generative models.
16 papers · 0 benchmarks
ICFG-PEDES (Identity-Centric and Fine-Grained Person Description Dataset)
One large-scale database for Text-to-Image Person Re-identification, i.e., Text-based Person Retrieval.
16 papers · 3 benchmarks
ILDC (Indian Legal Documents Corpus)
The ILDC dataset (Indian Legal Documents Corpus) is a large corpus of 35k Indian Supreme Court cases annotated with original court decisions.
16 papers · 0 benchmarks
JuICe is a corpus of 1.5 million examples with a curated test set of 3.7K instances based on online programming assignments.
16 papers · 0 benchmarks
MATRES (Multi-Axis Temporal RElations for Start-points)
This is the Multi-Axis Temporal RElations for Start-points (i.e., MATRES) dataset
16 papers · 2 benchmarks
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S.
16 papers · 1 benchmark
NewsCLIPpings is a dataset for detecting mismatched images and captions.
16 papers · 0 benchmarks
OntoNotes Release 4.0 contains the content of earlier releases -- OntoNotes Release 1.0 LDC2007T21, OntoNotes Release 2.0 LDC2008T04 and OntoNotes Release 3.0 LDC2009T24 -- and adds newswire, broadcast news, broadcast conversation and web…
16 papers · 1 benchmark
SMHD (Self-reported Mental Health Diagnoses)
A novel large dataset of social media posts from users with one or multiple mental health conditions along with matched control users.
16 papers · 0 benchmarks
SONAR, a new multilingual and multimodal fixed-size sentence embedding space, with a full suite of speech and text encoders and decoders.
16 papers · 0 benchmarks
SPoC (Pseudocode-to-Code)
Pseudocode-to-Code (SPoC) is a program synthesis dataset, containing 18,356 programs with human-authored pseudocode and test cases.
16 papers · 2 benchmarks
SemArt is a multi-modal dataset for semantic art understanding.
16 papers · 0 benchmarks
TV show Caption is a large-scale multimodal captioning dataset, containing 261,490 caption descriptions paired with 108,965 short video moments.
16 papers · 1 benchmark
To address the need for a standard open domain table benchmark dataset, the author propose a novel weak supervision approach to automatically create the TableBank, which is orders of magnitude larger than existing human labeled datasets…
16 papers · 0 benchmarks
TextComplexityDE is a dataset consisting of 1000 sentences in German language taken from 23 Wikipedia articles in 3 different article-genres to be used for developing text-complexity predictor models and automatic text simplification in…
16 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.