Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 33 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1537–1584 of 3,130
HatemojiCheck is a test suite for detecting emoji-based hate of 3,930 test cases covering seven functionalities of emoji-based hate and six identities.
3 papers · 0 benchmarks
HeadlineCause is a dataset for detecting implicit causal relations between pairs of news headlines.
3 papers · 0 benchmarks
Healthline is a nutrition related dataset for multi-document summarization, using scientific studies.
3 papers · 0 benchmarks
Hephaestus (Hephaestus: A large scale multitask dataset towards InSAR understanding)
Hephaestus is the first large-scale InSAR dataset.
3 papers · 0 benchmarks
HuRDL (Human-Robot Dialogue Learning Corpus)
The Human-Robot Dialogue Learning (HuRDL) Corpus is a dataset about asking questions in situated task-based interactions.
3 papers · 0 benchmarks
IAM Dataset (A Comprehensive and Large-Scale Dataset for Integrated Argument Mining Tasks)
We introduce a large and comprehensive dataset to facilitate the study of several essential AM tasks in the debating system.
3 papers · 2 benchmarks
ICSI Meeting Corpus in JSON format.
3 papers · 1 benchmark
Please refer: https://github.com/google/imageinwords/blob/main/datasets/IIW-400/README.md
3 papers · 0 benchmarks
IPAC (Icelandic Parallel Abstracts Corpus)
IPAC (Icelandic Parallel Abstracts Corpus ) is a new Icelandic-English parallel corpus, composed of abstracts from student theses and dissertations.
3 papers · 0 benchmarks
The IWSLT 2019 dataset contains source, Machine Translated, reference and Post-Edited text, which can be used to quantify and evaluate Post-editing effort after automatic MT.
3 papers · 0 benchmarks
IfAct (Identifying Human Actions Visible in Online Vlogs)
We consider the task of identifying human actions visible in online videos.
3 papers · 0 benchmarks
IllusionVQA is a Visual Question Answering (VQA) dataset with two sub-tasks.
3 papers · 2 benchmarks
The IndicNLP corpus is a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families.
3 papers · 0 benchmarks
Itihasa is a large-scale corpus for Sanskrit to English translation containing 93,000 pairs of Sanskrit shlokas and their English translations.
3 papers · 1 benchmark
JDsearch is a personalized product search dataset comprised of real user queries and diverse user-product interaction types (clicking, adding to cart, following, and purchasing) collected from JD.com, a popular Chinese online shopping…
3 papers · 0 benchmarks
JUSTICE (JUSTICE: A Dataset for Supreme Court’s Judgment Prediction)
The dataset contains 3304 cases from the Supreme Court of the United States from 1955 to 2021.
3 papers · 0 benchmarks
The Jamendo Corpus is a voice detection dataset consisting of 93 songs with Creative Commons license from the Jamendo free music sharing website.
3 papers · 0 benchmarks
KIND (Kessler Italian Named-entities Dataset)
KIND is an Italian dataset for Named-Entity Recognition.
3 papers · 0 benchmarks
KMIR (Knowledge Memorization, Identification, and Reasoning)
KMIR (Knowledge Memorization, Identification, and Reasoning) is a benchmark that covers 3 types of knowledge, including general knowledge, domain-specific knowledge, and commonsense, and provides 184,348 well-designed questions.
3 papers · 0 benchmarks
Collected by cleaning data from knowledge-intensive websites like Wikipedia and science and technology reports, and processing it using reverse engineering techniques.
3 papers · 0 benchmarks
Konzil dataset was created by specialists of the University of Greifswald.
3 papers · 0 benchmarks
The Kvasir-VQA dataset is an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations.
3 papers · 0 benchmarks
LARC (Language-annotated Abstraction and Reasoning)
LARC is a dataset built from ARC (Abstraction and Reasoning Corpus).
3 papers · 0 benchmarks
A Large Dataset for Remote Sensing Image Change Captioning.
3 papers · 0 benchmarks
The dataset was proposed in LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.
3 papers · 0 benchmarks
LLeQA (Long-form Legal Question Answering)
LLeQA is a French native dataset for studying information retrieval and long-form question answering in the legal domain.
3 papers · 0 benchmarks
LayoutBench is a diagnostic benchmark that examines 4 spatial control skills (number, position, size, shape), where each skill consists of 2 OOD layout splits, i.e., in total 8 tasks = 4 skills x 2 splits.
3 papers · 1 benchmark
LeNER-Br is a dataset for named entity recognition (NER) in Brazilian Legal Text.
3 papers · 2 benchmarks
The dataset contains the annotations of characters' visual appearances, in the form of tracks of face bounding boxes, and the associations with characters' textual mentions, when available.
3 papers · 1 benchmark
MCVQA (Multilingual and Code-mixed Visual Question Answering)
The MCVQA dataset consists of 248, 349 training questions and 121, 512 validation questions for real images in Hindi and Code-mixed.
3 papers · 0 benchmarks
MDIA is a large-scale multilingual benchmark for dialogue generation.
3 papers · 0 benchmarks
MDID (Multimodal Document Intent Dataset)
The Multimodal Document Intent Dataset (MDID) is a dataset for computing author intent from multimodal data from Instagram.
3 papers · 0 benchmarks
MFAQ is a multilingual FAQ dataset publicly available.
3 papers · 0 benchmarks
MIMIC-IV ICD-10 contains 122,279 discharge summaries—free-text medical documents—annotated with ICD-10 diagnosis and procedure codes.
3 papers · 1 benchmark
Question Answering (QA) is a widely-used framework for developing and evaluating an intelligent machine.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
MMCode is a multi-modal code generation dataset designed to evaluate the problem-solving skills of code language models in visually rich contexts (i.e.
3 papers · 0 benchmarks
The main goal of the data collection is to acquire highly natural conversations that cover a wide variety of styles and scenarios.
3 papers · 2 benchmarks
MOMA-LRG (Multi-Object Multi-Actor activity parsing with Language-Refined Graphs)
A dataset dedicated to multi-object, multi-actor activity parsing.
3 papers · 1 benchmark
MRS (Multilingual Reply Suggestion)
MRS, a multilingual reply suggestion dataset with ten languages.
3 papers · 0 benchmarks
MULTI-Benchmark is a cutting-edge benchmark for evaluating Multimodal Large Language Models (MLLMs).
3 papers · 0 benchmarks
MUTE (Multimodal Bengali Hateful Memes Dataset)
MUTE This is the first open-source Bengali Hateful Meme dataset, consisting of around 4200 memes annotated with two labels: hate and not hate.
3 papers · 0 benchmarks
We sample 2025 frames of images from the original KITTI for Mono3DRefer, containing 41,140 expressions in total and a vocabulary of 5,271 words.
3 papers · 0 benchmarks
MuLD (Multitask Long Document Benchmark)
MuLD (Multitask Long Document Benchmark) is a set of 6 NLP tasks where the inputs consist of at least 10,000 words.
3 papers · 5 benchmarks
MultiQ is a multi-hop QA dataset for Russian, suitable for general open-domain question answering, information retrieval, and reading comprehension tasks.
3 papers · 1 benchmark
Multilingual TOP is a dataset for multilingual semantic parsing with human-written sentences as opposed to machine translated ones.
3 papers · 0 benchmarks
Named Entity (NER) annotations of the Hebrew Treebank (Haaretz newspaper) corpus, including: morpheme and token level NER labels, nested mentions, and more.
3 papers · 3 benchmarks
NarraSum is a large-scale narrative summarization dataset.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.