Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 29 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1345–1392 of 3,130

FewSOL (A Dataset for Few-Shot Object Learning in Robotic Environments)
The Few-Shot Object Learning (FewSOL) dataset can be used for object recognition with a few images per object.
4 papers · 0 benchmarks
FiNER-139 is comprised of 1.1M sentences annotated with eXtensive Business Reporting Language (XBRL) tags extracted from annual and quarterly reports of publicly-traded companies in the US.
4 papers · 0 benchmarks
Pretrain: 200k Instruction: 100k
4 papers · 0 benchmarks
The first NER dataset in the field of traffic, which is to extract the characteristics and attributes of the vehicle on the road.
4 papers · 2 benchmarks
Finer (Finnish News Corpus for Named Entity Recognition)
Finnish News Corpus for Named Entity Recognition (Finer) is a corpus that consists of 953 articles (193,742 word tokens) with six named entity classes (organization, location, person, product, event,and date).
4 papers · 0 benchmarks
Finnish Paraphrase Corpus is a fully manually annotated paraphrase corpus for Finnish containing 53,572 paraphrase pairs harvested from alternative subtitles and news headings.
4 papers · 0 benchmarks
GDSC (Genomics of Drug Sensitivity in Cancer)
We have characterized 1000 human cancer cell lines and screened them with 100s of compounds.
4 papers · 1 benchmark
GGPONC (German Guideline Program in Oncology NLP Corpus)
German Guideline Program in Oncology NLP Corpus (GGPONC) is a German language corpus based on clinical practice guidelines for oncology.
4 papers · 0 benchmarks
Gazeta is a dataset for automatic summarization of Russian news.
4 papers · 1 benchmark
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
Goal is a novel dataset of football (or 'soccer') highlights videos with transcribed live commentaries in English.
4 papers · 0 benchmarks
HKR (Handwritten Kazakh and Russian (HKR) Database for Text Recognition)
The database is written in Cyrillic and shares the same 33 characters.
4 papers · 1 benchmark
Models character profiles and gives dialogue agents the ability to learn characters' language styles through their HLAs.
4 papers · 0 benchmarks
Images with paired ground-truth caption hierarchies
4 papers · 0 benchmarks
The IS-A dataset is a dataset of relations extracted from a medical ontology.
4 papers · 0 benchmarks
ITALIC: An ITALian Intent Classification Dataset ITALIC is an intent classification dataset for the Italian language, which is the first of its kind.
4 papers · 0 benchmarks
ImgEdit is a large-scale, high-quality image-editing dataset comprising 1.2 million carefully curated edit pairs, which contain both novel and complex single-turn edits, as well as challenging multi-turn tasks.
4 papers · 0 benchmarks
IndoNLI is the first human-elicited NLI dataset for Indonesian consisting of nearly 18K sentence pairs annotated by crowd workers and experts.
4 papers · 0 benchmarks
InferWiki is a Knowledge Graph Completion (KGC) dataset that improves upon existing benchmarks in inferential ability, assumptions, and patterns.
4 papers · 0 benchmarks
This is a large-scale dataset of tweets associated to thousands of news articles published on Italian disinformation websites in the context of 2019 European elections.
4 papers · 0 benchmarks
JaQuAD (Japanese Question Answering Dataset) is a question answering dataset in Japanese that consists of 39,696 extractive question-answer pairs on Japanese Wikipedia articles.
4 papers · 1 benchmark
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior.
4 papers · 1 benchmark
JobStack is a new corpus for de-identification of personal data in job vacancies on Stackoverflow.
4 papers · 0 benchmarks
KRAUTS (Korpus of newspapeR Articles with Underlinded Temporal expressionS)
KRAUTS (Korpus of newspapeR Articles with Underlinded Temporal expressionS) is a German temporally annotated news corpus accompanied with TimeML annotation guidelines for German.
4 papers · 1 benchmark
KazakhTTS is an open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide.
4 papers · 0 benchmarks
This is the dataset for knowledge editing.
4 papers · 0 benchmarks
LAM(line-level) (The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text Recognition)
Handwritten Text Recognition (HTR) is an open problem at the intersection of Computer Vision and Natural Language Processing.
4 papers · 1 benchmark
The LIAR dataset has been widely followed by fake news detection researchers since its release, and along with a great deal of research, the community has provided a variety of feedback on the dataset to improve it.
4 papers · 1 benchmark
LLM-Seg40K dataset contains 14K images in total.
4 papers · 0 benchmarks
Laptop-ACOS is a brand new Laptop dataset collected from the Amazon platform in the years 2017 and 2018 (covering ten types of laptops under six brands such as ASUS, Acer, Samsung, Lenovo, MBP, MSI, and so on).
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
LayoutBench-COCO is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.
4 papers · 1 benchmark
This dataset is a benchmark for complex reasoning abilities in large language models, drawing on United Kingdom Linguistics Olympiad problems which cover a wide range of languages.
4 papers · 1 benchmark
M³-VOS (M³-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation)
💡 Description A new benchmark, Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation (M³-VOS), to verify the ability of models to understand object phases, which consists of 479 high-resolution videos spanning over 10…
4 papers · 1 benchmark
The M5Product dataset is a large-scale multi-modal pre-training dataset with coarse and fine-grained annotations for E-products.
4 papers · 0 benchmarks
MACSum a human-annotated summarization dataset for controlling mixed attributes.
4 papers · 0 benchmarks
MMToM-QA (Multimodal Theory of Mind Question Answering)
MMToM-QA is the first multimodal benchmark to evaluate machine Theory of Mind (ToM), the ability to understand people's minds.
4 papers · 0 benchmarks
MN-DS (Multilabeled News Dataset)
Multilabeled News Dataset (MN-DS) is a dataset for news classification.
4 papers · 0 benchmarks
MOPRD, a multidisciplinary open peer review dataset consists of paper metadata, multiple version manuscripts, review comments, meta-reviews, author's rebuttal letters, and editorial decisions from 6578 papers.
4 papers · 0 benchmarks
MSRVTT-CTN (MSRVTT Causal-Temporal Narrative)
MSRVTT-CTN Dataset This dataset contains CTN annotations for the MSRVTT-CTN benchmark dataset in JSON format.
4 papers · 1 benchmark
MSVD-CTN (MSVD Causal-Temporal Narrative)
MSVD-CTN Dataset This dataset contains CTN annotations for the MSVD-CTN benchmark dataset in JSON format.
4 papers · 1 benchmark
MS^2 (Multi-Document Summarization of Medical Studies)
MS^2 (Multi-Document Summarization of Medical Studies) is a dataset of over 470k documents and 20k summaries derived from the scientific literature.
4 papers · 1 benchmark
math 500
4 papers · 1 benchmark
Medical Abstracts (Medical Abstracts Text Classification Dataset)
The Medical Abstracts dataset contains 14,438 medical abstracts describing 5 different classes of patient conditions, with all of the dataset being annotated.
4 papers · 1 benchmark
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection This is MetaHate: a meta-collection of 36 hate speech datasets from social media comments.
4 papers · 0 benchmarks
To construct the MICROSOFT RESEARCH MULTIMODAL ALIGNED RECIPE CORPUS the authors first extract a large number of text and video recipes from the web.
4 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.