Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 53 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 2497–2544 of 3,130
MTC is a financial-domain dataset of the multi-label topic classification task.
1 paper · 0 benchmarks
MTTN is a large scale derived and synthesized dataset built with on real prompts and indexed with popular image-text datasets like MS-COCO, Flickr, etc.
1 paper · 0 benchmarks
MUSIED is a large-scale Chinese event detection dataset based on user reviews, text conversations, and phone conversations in a leading e-commerce platform for food service, designed for event detection tasks.
1 paper · 0 benchmarks
MVME (Multi-View Medical Evaluation Benchmark)
The benchmark assesses the real-time interactive consultation capabilities of LLMs across three critical dimensions.
1 paper · 0 benchmarks
Multiword expressions (MWEs) represent lexemes that should be treated as single lexical units due to their idiosyncratic nature.
1 paper · 0 benchmarks
This dataset is used to train and evaluate models for the detection of machine-paraphrased text.
1 paper · 0 benchmarks
Dataset introduction There are four dimension in MBTI.
1 paper · 0 benchmarks
A collection of over 700 games of Mafia, in which players are randomly assigned either deceptive or non-deceptive roles and then interact via forum postings.
1 paper · 0 benchmarks
MalayalamMixSentiment is a Sentiment Analysis Dataset for Code-Mixed Malayalam-English.
1 paper · 0 benchmarks
MapEval contains 700 question-answer pairs.
1 paper · 0 benchmarks
MapEval-Textual contains 300 question-answer pairs.
1 paper · 1 benchmark
MapEval-Textual contains 300 context-question-answer triplets.
1 paper · 1 benchmark
MapEval-Visual contains 400 image-question-answer triplets.
1 paper · 1 benchmark
MapReader in GeoHumanities workshop (SIGSPATIAL 2022): Gold standards and outputs Refer to: https://github.com/Living-with-machines/MapReader/wiki/GeoHumanities-workshop-in-SIGSPATIAL-2022
1 paper · 0 benchmarks
This dataset accompanies the ICWSM 2022 paper "Mapping Topics in 100,000 Real-Life Moral Dilemmas".
1 paper · 0 benchmarks
Describe the Marmara Turkish Coreference Corpus, which is an annotation of the whole METU-Sabanci Turkish Treebank with mentions and coreference chains.
1 paper · 0 benchmarks
pymatgencodeqa benchmark: qabenchmark/generatedqa/generationresultscode.json, which consists of 34,621 QA pairs.
1 paper · 0 benchmarks
MathEquiv (mathematical statement equivalence)
MathEquiv dataset is accompanied to EquivPruner .
1 paper · 0 benchmarks
Mathematical dataset based on 71 famous mathematical identities.
1 paper · 0 benchmarks
MatriVasha the largest dataset of handwritten Bangla compound characters for research on handwritten Bangla compound character recognition.
1 paper · 0 benchmarks
McQueen dataset contains 15k visual conversations and over 80k queries where each one is associated with a fully-specified rewrite version.
1 paper · 0 benchmarks
MeSHup (A Corpus for Full Text Biomedical Document Indexing)
Contains 1,342,667 full text articles in English, together with the associated MeSH labels and metadata, authors, and publication venues that are collected from the MEDLINE database.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
A comprehensive Turkish dataset for question-answering tasks in medical domain
1 paper · 1 benchmark
The MedVidCL dataset contains a collection of 6, 617 videos annotated into ‘medical instructional’, ‘medical non-instructional' and ‘non-medical’ classes.
1 paper · 0 benchmarks
MediBeng (Synthetic Code-Switched Bengali-English Speech Conversations for Healthcare Applications)
MediBeng Dataset The MediBeng dataset contains synthetic code-switched dialogues in Bengali and English for training models in speech recognition (ASR), text-to-speech (TTS), and machine translation in clinical settings.
1 paper · 1 benchmark
MediConfusion is a challenging medical Visual Question Answering (VQA) benchmark dataset, that probes the failure modes of medical Multimodal Large Language Models (MLLMs) from a vision perspective.
1 paper · 0 benchmarks
Mediapi-RGB is a bilingual corpus of French Sign Language (LSF) and written French in the form of subtitled videos, accompanied by complementary data (various representations, segmentation, vocabulary, etc.).
1 paper · 1 benchmark
Medical Case Report Corpus is a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library.
1 paper · 0 benchmarks
A medical Wiki paralell corpus for medical text simplification.
1 paper · 0 benchmarks
Mega-COV is a billion-scale dataset from Twitter for studying COVID-19.
1 paper · 0 benchmarks
MerRec (MerRec Recommendation Dataset)
A large scale, C2C marketplace e-commerce dataset.
1 paper · 0 benchmarks
We introduce a new task of rephrasing for amore natural virtual assistant.
1 paper · 0 benchmarks
This data adds textual meta-infomation data to two existing corpora for cross language information retrieval: BoostCLIR, and the Large Scale CLIR Dataset (wiki-clir).
1 paper · 0 benchmarks
MetaEval is a collection of 101 NLP tasks.
1 paper · 0 benchmarks
Metric-Type of Numerical Tables is a dataset extracted from scientific papers (ACL anthology website) consisting of header tables, captions, and metric-types.
1 paper · 0 benchmarks
MiMIC (Multi-Modal Indian Earnings Calls Dataset)
Predicting stock market prices following corporate earnings calls remains a significant challenge for investors and researchers alike, requiring innovative approaches that can process diverse information sources.
1 paper · 0 benchmarks
MiST (Modals In Scientific Text) is a dataset containing 3737 modal instances in five scientific domains annotated for their semantic, pragmatic, or rhetorical function.
1 paper · 0 benchmarks
We present a comprehensive dataset comprising a vast collection of raw mineral samples for the purpose of mineral recognition.
1 paper · 0 benchmarks
The MiniHAREM, a reiteration of the 2005 evaluation, used the same methodology and platform.
1 paper · 0 benchmarks
Mint (Multilingual Intimacy analysis)
Mint is a new Multilingual intimacy analysis dataset covering 13,384 tweets in 10 languages including English, French, Spanish, Italian, Portuguese, Korean, Dutch, Chinese, Hindi, and Arabic.
1 paper · 0 benchmarks
MoToMQA (Multi-Order Theory of Mind Question & Answer)
The MoToMQA (Multi-Order Theory of Mind Question & Answer) benchmark is a test suite introduced to examine the extent to which large language models (LLMs) have developed higher-order theory of mind (ToM); the human ability to reason about…
1 paper · 0 benchmarks
Modern Hebrew Sentiment Dataset is a sentiment analysis benchmark for Hebrew, based on 12K social media comments, and provide two instances of these data: in token-based and morpheme-based settings.
1 paper · 0 benchmarks
This dataset of medical misinformation was collected and is published by Kempelen Institute of Intelligent Technologies (KInIT).
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
This is a dataset of robot motions based on physics simulations.
1 paper · 0 benchmarks
Movie Reviews (Movie Review Polarity Dataset Enriched with "Annotator Rationales")
This dataset is based on the movie review polarity dataset (v2.0) collected and maintained by Bo Pang and Lillian Lee.
1 paper · 0 benchmarks
Given an ongoing dialogue between a user and a dialogue assistant, for the user query, the model is required to predict both coreference links between the query and the dialogue context, and the self-contained rewritten user query that is…
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.