Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 52 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2449–2496 of 3,130

A corpus of 21,570 newspaper headlines written in European Spanish annotated with emergent anglicisms.
1 paper · 0 benchmarks
LeT-Mi (Levantine Twitter dataset for Misogynistic language)
Levantine Twitter dataset for Misogynistic language (LeT-Mi) is an Arabic Levantine Twitter dataset for misogynistic language to be the first benchmark dataset for Arabic misogyny.
1 paper · 0 benchmarks
Dataset Summary New dataset introduced in Parameter-Efficient Legal Domain Adaptation (Li et al., 2022) from the Legal Advice Reddit community (known as "/r/legaldvice"), sourcing the Reddit posts from the Pushshift Reddit dataset.
1 paper · 0 benchmarks
Recognizing events and their coreferential men- tions in a document is essential for understand- ing semantic meanings of text.
1 paper · 0 benchmarks
The Lenta Short Sentences dataset is a text dataset for language modelling for the Russian language.
1 paper · 0 benchmarks
LibriS2S is a Speech to Speech Translation (S2ST) dataset build further upon existing resources.
1 paper · 0 benchmarks
The Linguistic Benchmark (JSON), consisting of 30 questions was developed to be easy for human adults to answer but challenging for LLMs.
1 paper · 0 benchmarks
This is a dataset of 3 English books which do not contain the letter "e" in them.
1 paper · 1 benchmark
A listwise multi-response dataset for human preferences alignment.
1 paper · 0 benchmarks
An open-source online generative dictionary that takes a word and context containing the word as input and automatically generates a definition as output.
1 paper · 0 benchmarks
The Liu et al.
1 paper · 0 benchmarks
M-Phasis (A Feature-Based Corpus of Hate Online)
A corpus of 9k German and French user comments collected from migration-related news articles.
1 paper · 0 benchmarks
M3LS (Multi-Lingual Multi-Modal Summarization Dataset)
Significant developments in techniques such as encoder-decoder models have enabled us to represent information comprising multiple modalities.
1 paper · 0 benchmarks
MAKED (MultiModal MultiLingual Summarization and Keyword Extraction Dataset)
Keyword extraction is an integral task for many downstream problems like clustering, recommendation, search and classification.
1 paper · 0 benchmarks
MAQA (Multi-Answer Question Answering dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
MAST (Multi-Attributed Structured Text-to-face Dataset)
A new data consolidation called Multi-Attributed and Structured Text-to-face (MAST) dataset.
1 paper · 0 benchmarks
The MATHWELL Human Annotation Dataset contains 5,084 synthetic word problems and answers generated by MATHWELL, a reference-free educational grade school math word problem generator released in MATHWELL: Generating Educational Math Word…
1 paper · 0 benchmarks
MAVE - Attribute: Black Tea Variety (MAVE - Attribute: Black Tea Variety: A Product Dataset for Multi-source Attribute Value Extraction)
The dataset contains 3 million attribute-value annotations across 1257 unique categories created from 2.2 million cleaned Amazon product profiles.
1 paper · 0 benchmarks
MELON (Melodic Design)
1.
1 paper · 0 benchmarks
MENYO-20k is the first multi-domain parallel corpus with a special focus on clean orthography for Yorùbá--English with standardized train-test splits for benchmarking.
1 paper · 0 benchmarks
MF (Mathematical Formulas)
Mathematical dataset containing formulas based on the AMPS Khan dataset and the ARQMath dataset V1.3.
1 paper · 0 benchmarks
MF3QA (Medical Free Form Farsi Question Answering dataset)
real-world doctor-patient question- answering dataset cleaned manually and automatically
1 paper · 0 benchmarks
MF3QA_uncleaned (Medical Free Form Farsi Question Answering dataset (uncleaned))
real-world doctor-patient question- answering dataset
1 paper · 0 benchmarks
The dataset from the study "A Fully Generative Motivational Interviewing Counsellor Chatbot for Moving Smokers Towards the Decision to Quit".
1 paper · 0 benchmarks
To assess a model’s ability to create microcontroller-driven electronic devices, we developed a benchmark, MICRO25, that includes 25 tasks intended for the common ARDUINO microcontroller ecosystem..
1 paper · 0 benchmarks
This dataset contains 4606 articles from 1996 to 2024 that were presented in MIE (Medical Informatics Europe Conference) conferences.
1 paper · 0 benchmarks
MILU (Multi-task Indic Language Understanding Benchmark)
Overview MILU (Multi-task Indic Language Understanding Benchmark) is a comprehensive evaluation dataset designed to assess the performance of Large Language Models (LLMs) across 11 Indic languages.
1 paper · 0 benchmarks
MIMIC-IV ICD-9 contains 209,326 discharge summaries—free-text medical documents—annotated with ICD-9 diagnosis and procedure codes.
1 paper · 1 benchmark
Retrospectively collected medical data has the opportunity to improve patient care through knowledge discovery and algorithm development.
1 paper · 0 benchmarks
MIMIC-IV-Note (MIMIC-IV-Note: Deidentified free-text clinical notes)
The advent of large, open access text databases has driven advances in state-of-the-art model performance in natural language processing (NLP).
1 paper · 0 benchmarks
MINT (a Multi-modal Image and Narrative Text Dubbing Dataset)
Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape.
1 paper · 0 benchmarks
MIPD (Manipulation and Intention In a Novel Corpus of Polish Disinformation)
A novel corpus of 15,356 Polish web articles, including articles identified as containing disinformation.
1 paper · 0 benchmarks
An example dataset of 110,000 question/query pairs across four WikiData domains.
1 paper · 0 benchmarks
MLQuestions is a domain-adaptation dataset for the machine learning domain containing 50K unaligned passages and 35K unaligned questions, and 3K aligned passage and question pairs.
1 paper · 0 benchmarks
MM-Eval (Modern Mongolian Evaluation)
Large language models (LLMs) excel in high-resource languages but face notable challenges in low-resource languages like Mongolian.
1 paper · 0 benchmarks
MM-Locate-News is a dataset for location estimation of news.
1 paper · 0 benchmarks
MMCOMPOSITION is a high-quality benchmark specifically designed to comprehensively evaluate the compositionality of pre-trained Vision-Language Models (VLMs) across three main dimensions—VL compositional perception, reasoning, and…
1 paper · 0 benchmarks
MMInstruct-GPT4V (MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity)
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: 1.
1 paper · 0 benchmarks
Mix of Minimal Optimal Sets (MMOS) of dataset has two advantages for two aspects, higher performance and lower construction costs on math reasoning.
1 paper · 0 benchmarks
MMSQL (Multi-Turn Multi-Type Text-to-SQL test suit)
A dataset for training and testing tin various problem types and multi-turn Q&A scenarios, including a training set, test set, and test scripts.
1 paper · 1 benchmark
The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations.
1 paper · 0 benchmarks
This is a challenging dataset from real courtrooms to predict the legal judgment in a reasonably encyclopedic manner by leveraging the genuine input of the case -- plaintiff's claims and court debate data, from which the case's facts are…
1 paper · 0 benchmarks
MSNER (Multilingual Spoken Named Entity Recognition)
This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish.
1 paper · 0 benchmarks
MSVD-Indonesian is derived from the MSVD dataset, which is obtained with the help of a machine translation service.
1 paper · 3 benchmarks
MT (Mathematical Texts)
Mathematical dataset containing mathematical texts, i.e., texts containing LaTeX formulas, based on the AMPS Khan dataset and the ARQMath dataset V1.3.
1 paper · 0 benchmarks
MT-Bench in Thai.
1 paper · 0 benchmarks
The MT40K dataset for predicting malware threat intelligence is a collection of 40,000 triples generated from 27,354 unique entities and 34 relations.
1 paper · 0 benchmarks
MT560 (MT560 - A Many-to-English Machine Translation Dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.