Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 54 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2545–2592 of 3,130

MuLMS (Multi-Layer Materials Science)
The Multi-Layer Materials Science corpus (MuLMS) consists of 50 documents (licensed CC BY) from the materials science domain, spanning across the following 7 subareas: "Electrolysis", "Graphene", "Polymer Electrolyte Fuel Cell (PEMFC)",…
1 paper · 0 benchmarks
This is the large version of the MuMiN dataset.
1 paper · 1 benchmark
This is the medium version of the MuMiN dataset.
1 paper · 1 benchmark
This is the small version of the MuMiN dataset.
1 paper · 1 benchmark
Dataset Description The dataset used in this study comprises bug reports extracted from the Visual Studio Code GitHub repository, specifically focusing on those labeled with the english-please tag.
1 paper · 1 benchmark
Multi-CrossRE is a broadest multi-lingual dataset for Relation Extraction (RE) including 26 languages in addition to English, and covering six text domains.
1 paper · 0 benchmarks
The Multi-Eup is a new multilingual benchmark dataset, comprising 22K multilingual documents collected from the European Parliament, spanning 24 languages.
1 paper · 0 benchmarks
The Multi-domain Image Characteristic Dataset consists of thousands of images sourced from the internet.
1 paper · 0 benchmarks
MultiCite is a dataset of 12,653 citation contexts from over 1,200 computational linguistics papers used for Citation context analysis (CCA).
1 paper · 0 benchmarks
MultiRefKGC (multi-reference KGC)
MultiRefKGC is a dataset created from conversations from Reddit designed for Knowledge-Grounded Dialogue Generation tasks.
1 paper · 0 benchmarks
MultiSum is a dataset for multimodal summarization (MSMO).
1 paper · 0 benchmarks
MultiWOZ-coref, (or MultiWOZ 2.3) is an extension of the MultiWOZ dataset that adds co-reference annotations in addition to corrections of dialogue acts and dialogue states.
1 paper · 0 benchmarks
This dataset contains dialogue lines from the games Knights of the Old Republic 1 & 2 and Neverwinter Nights 1.
1 paper · 0 benchmarks
The first annotated corpus for multilingual analysis of potentially unfair clauses in online Terms of Service.
1 paper · 0 benchmarks
Multimedia Goal-oriented Generative Script Learning Dataset This link contains a dataset consisting of multimedia steps for two categories: gardening and crafts.
1 paper · 0 benchmarks
MuseChat Dataset (MuseChat: A Conversational Music Recommendation System for Videos (CVPR 2024 Highlight Paper))
Music recommendation for videos attracts growing interest in multi-modal research.
1 paper · 0 benchmarks
M²ConceptBase is a concept-centric multimodal knowledge base designed to bridge the gap between visual and linguistic semantics.
1 paper · 0 benchmarks
N15News is a large-scale multimodal news dataset comprising 200K imagetext pairs and 15 categories, which exceeding the previous news dataset in both the number of categories and samples.
1 paper · 1 benchmark
This dataset contains names that are exclusively associated with a single gender and that have no ambiguous meanings, therefore being exact with respect to both gender and meaning.
1 paper · 0 benchmarks
This dataset extends NAMEXACT by including words that can be used as names, but may not exclusively be used as names in every context.
1 paper · 0 benchmarks
NC-SentNoB (Noise Classification on SentNoB)
This is a multilabel dataset used for Noise Identification purpose in the paper "A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts" accepted in 2024 The 9th Workshop on Noisy and User-generated…
1 paper · 0 benchmarks
NCSE v2.0 (NCSE v2.0: A Dataset of OCR-Processed 19th Century English Newspapers)
The NCSE v2.0 is a digitized collection of six 19th-century English periodicals The ground truth contains 358 cropped images of text blocks from 31 pages of 19th century newspaper data
1 paper · 0 benchmarks
NCTE Transcripts consists of 1,660 45-60 minute long 4th and 5th grade elementary mathematics observations collected by the National Center for Teacher Effectiveness (NCTE) between 2010-2013.
1 paper · 0 benchmarks
NELA-GT-2021 is the fourth installment of the NELA-GT datasets, NELA-GT-2021.
1 paper · 0 benchmarks
NEREL-BIO is an annotation scheme and corpus of PubMed abstracts in Russian and English.
1 paper · 0 benchmarks
NJH (Not Just Hate)
NJH is a dataset of over 40,000 tweets about immigration from the US and UK, annotated with six labels for different aspects of incivility and intolerance.
1 paper · 0 benchmarks
NL2GQL Dataset (NL2GQL developped for R3-NL2GQL)
A bilingual (English and Chinese natural language queries) dataset which has NL queries annotated with their corresponding GQL queries (i.e.
1 paper · 0 benchmarks
The first Portuguese dataset compiled for Native Language Identification (NLI), the task of identifying an author's first language based on their second language writing.
1 paper · 0 benchmarks
NLI4Wills Corpus can be used to train transformers and sentence-transformer models for the validity evaluation of the legal will statements.
1 paper · 0 benchmarks
NMF (Named Mathematical Formulas)
Mathematical dataset based on 71 famous mathematical identities.
1 paper · 0 benchmarks
This corpus contains data files that were generated as part of the NOVIC paper (see above).
1 paper · 0 benchmarks
NText is an eight million words dataset extracted and preprocessed from nuclear research papers and thesis.
1 paper · 0 benchmarks
NaSGEC is a new dataset to facilitate research on Chinese grammatical error correction (CGEC) for native speaker texts from multiple domains.
1 paper · 0 benchmarks
Diacritized texts in Modern Hebrew, collected from eleven different sources.
1 paper · 0 benchmarks
A collection of diacritized Hebrew text in a variety of registers and from different sources.
1 paper · 0 benchmarks
Includes co-referent name string pairs along with their similarities.
1 paper · 0 benchmarks
Narvik Road Dataset (DIT4BEARs Smart Road Dataset)
DIT4BEARs Internship Project (at UiT-The Arctic University of Norway) Dataset The dataset contains data of 5 months including weather conditions, friction coefficient, distance traveled, wind speed, surface temperature, air temperature,…
1 paper · 0 benchmarks
We scraped the Gutenberg Project and a subset of English Wikipedia to obtain the list of sentences that contain any.
1 paper · 0 benchmarks
The dataset comprises 1641 questions and answers generated as three separate parts.
1 paper · 0 benchmarks
Collected by cleaning data from daily Xinwen Lianbo transcripts over the past three months and processing it using reverse engineering techniques.
1 paper · 0 benchmarks
NewsMTSC is a dataset for target-dependent sentiment classification (TSC) on news articles reporting on policy issues.
1 paper · 0 benchmarks
A benchmark for legal question answering.
1 paper · 0 benchmarks
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models.
1 paper · 0 benchmarks
Dataset composed of two main parts 1.
1 paper · 0 benchmarks
Automatic language identification is a challenging problem.
1 paper · 1 benchmark
OAGT (Paper Topic Dataset)
OAGL is a paper topic dataset consisting of 6942930 records which comprise various scientific publication attributes like abstracts, titles, keywords, publication years, venues, etc.
1 paper · 0 benchmarks
The OFEQ-10k dataset contains 12,548 detailed questions with corresponding math headlines from MathOverflow.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.