Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 39 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1825–1872 of 3,130
We introduce KPI-EDGAR, a novel dataset for Joint Named Entity Recognition and Relation Extraction building on financial reports uploaded to the Electronic Data Gathering, Analysis, and Retrieval (EDGAR) system, where the main objective is…
2 papers · 1 benchmark
Kaleidoscope (Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation)
The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage.
2 papers · 0 benchmarks
KazNERD is a dataset for Kazakh named entity recognition.
2 papers · 0 benchmarks
General Corpora for the Maltese Language.
2 papers · 0 benchmarks
Nouns extracted automatically from Bible translations across 1580 languages.
2 papers · 0 benchmarks
LEPISZCZE is an open-source comprehensive benchmark for Polish NLP and a continuous-submission leaderboard, concentrating public Polish datasets (existing and new) in specific tasks.
2 papers · 0 benchmarks
LUMA (Learning from Uncertain and Multimodal Data)
LUMA is a multimodal dataset that consists of audio, image, and text modalities.
2 papers · 0 benchmarks
The Large-Scale CLIR Dataset is a retrieval dataset built for Cross-Language Information Retrieval (CLIR).
2 papers · 0 benchmarks
Description Dataset from the Law Stack Exchange, as used in "Parameter-Efficient Legal Domain Adaptation" (Li et al., 2022).
2 papers · 0 benchmarks
The Leipzig Corpora Collection presents corpora in different languages using the same format and comparable sources.
2 papers · 0 benchmarks
An RDF knowledge graph that provides comprehensive, current information about almost 400,000 machine learning publications.
2 papers · 0 benchmarks
The Live Comment Dataset is a large-scale dataset with 2,361 videos and 895,929 live comments that were written while the videos were streamed.
2 papers · 0 benchmarks
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks.
2 papers · 0 benchmarks
M2QA (Multi-domain Multilingual Question Answering)
M2QA (Multi-domain Multilingual Question Answering) is an extractive question answering benchmark for evaluating joint language and domain transfer.
2 papers · 0 benchmarks
MCSCSet is a large-scale specialist-annotated dataset, designed for the task of Medical-domain Chinese Spelling Correction that contains about 200k samples.
2 papers · 0 benchmarks
MDD (Movie Dialog dataset)
Movie Dialog dataset (MDD) is designed to measure how well models can perform at goal and non-goal orientated dialog centered around the topic of movies (question answering, recommendation and discussion).
2 papers · 0 benchmarks
MECD (Multi-Event Causal Discovery)
Provide: 1,105 lifestyle videos that span diverse scenarios.
2 papers · 1 benchmark
Persian-English parallel corpus with more than one million sentence pairs collected from masterpieces of literature.
2 papers · 0 benchmarks
MMSD2.0 (Towards a Reliable Multi-modal Sarcasm Detection System)
Multi-modal sarcasm detection has attracted much recent attention.
2 papers · 0 benchmarks
MMVax-Stance includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
MSDA (Multi-source domain adaptation dataset for text recognition)
5 domains: synthetic domain, document domain, street view domain, handwritten domain, and car license domain over five million images
2 papers · 2 benchmarks
The MUSE dataset contains bilingual dictionaries for 110 pairs of languages.
2 papers · 2 benchmarks
MUStARD (Multimodal Sarcasm Detection Dataset)
We release the MUStARD dataset which is a multimodal video corpus for research in automated sarcasm discovery.
2 papers · 0 benchmarks
The E-MASAC Dataset is a collection of code-mixed conversations sourced from an Indian TV series, focusing on Hindi-English interactions.
2 papers · 1 benchmark
The process by which sections in a document are demarcated and labeled is known as section identification.
2 papers · 2 benchmarks
MentSum (Mental Health Summarization Dataset)
Mental health remains a significant challenge of public health worldwide.
2 papers · 1 benchmark
MAUD is an expert-annotated merger agreement reading comprehension dataset based on the American Bar Association's 2021 Public Target Deal Points study, where lawyers and law students answered 92 questions about 152 merger agreements.
2 papers · 0 benchmarks
The Metaphorical Connections dataset is a poetry dataset that contains annotations between metaphorical prompts and short poems.
2 papers · 0 benchmarks
MiniWob++ is a suite of web-browser based tasks introduced in Liu et al.
2 papers · 0 benchmarks
MobIE is a German-language dataset which is human-annotated with 20 coarse- and fine-grained entity types and entity linking information for geographically linkable entities.
2 papers · 0 benchmarks
A large scale OCSR dataset, proposed in paper “MolParser: End-to-end Visual Recognition of Molecule Structures in the Wild“ MolParser-7M contains nearly 8 million paired image-SMILES data.
2 papers · 0 benchmarks
Morph Call is a suite of 46 probing tasks for four Indo-European languages that fall under different morphology: Russian, French, English, and German.
2 papers · 0 benchmarks
A version of the CMU Movie Summary Corpus (http://www.cs.cmu.edu/~ark/personas/), which was originally scraped from plot summaries from Wikipedia, with some cleaning and sentences turned into events & sorted into "genres" (via LDA).
2 papers · 0 benchmarks
MuCo-VQA consist of large-scale (3.7M) multilingual and code-mixed VQA datasets in multiple languages: Hindi (hi), Bengali (bn), Spanish (es), German (de), French (fr) and code-mixed language pairs: en-hi, en-bn, en-fr, en-de and en-es.
2 papers · 0 benchmarks
MultiOpEd is a corpus of multi-perspective news editorials.
2 papers · 0 benchmarks
MultiReQA is a cross-domain evaluation for retrieval question answering models.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
The dataset contains training and evaluation data for 12 languages: - Vietnamese - Romanian - Latvian - Czech - Polish - Slovak - Irish - Hungarian - French - Turkish - Spanish - Croatian For each language, one training, one development…
2 papers · 12 benchmarks
NCI (New Corpus for Ireland)
Contains a wide range of texts in Irish, including fiction, news reports, informative texts and official documents.
2 papers · 0 benchmarks
The dataset consists of titles and abstracts from NLP-related papers.
2 papers · 0 benchmarks
NPSC (Norwegian Parliamentary Speech Corpus)
The Norwegian Parliamentary Speech Corpus (NPSC) is a speech corpus made by the Norwegian Language Bank at the National Library of Norway in 2019-2021.
2 papers · 0 benchmarks
NQuAD (Nuclear Question Answering Dataset)
NQuAD is a Nuclear Question Answering Dataset, which contains 700+ nuclear Question Answer pairs developed and verified by expert nuclear researchers.
2 papers · 0 benchmarks
NYT-H is a dataset for distantly-supervised relation extraction, in which DS-labelled training data is used and several annotators to label test data are hired.
2 papers · 0 benchmarks
NewsPH-NLI is a sentence entailment benchmark dataset in the low-resource Filipino language.
2 papers · 0 benchmarks
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
2 papers · 0 benchmarks
NorDial is the first step to creating a corpus of dialectal variation of written Norwegian.
2 papers · 0 benchmarks
Scene-focused, multi-modal, episodic data of the images and symbolic world-states seen by an agent completing a pogo-stick assembly task within a video game world.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.