Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 35 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1633–1680 of 3,130

TaL Corpus (The Tongue and Lips Corpus)
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
A novel dataset of document-grounded task-based dialogues, where an Information Giver (IG) provides instructions (by consulting a document) to an Information Follower (IF), so that the latter can successfully complete the task.
3 papers · 0 benchmarks
Huggingface Datasets is a great library, but it lacks standardization, and datasets require preprocessing work to be used interchangeably.
3 papers · 0 benchmarks
TextBox 2.0 is a comprehensive and unified library for text generation, focusing on the use of pre-trained language models (PLMs).
3 papers · 0 benchmarks
It is a freely available resource for research on handling negation and uncertainty in biomedical texts .
3 papers · 5 benchmarks
The Little Prince (The Little Prince Corpus)
This corpus is an annotation of the novel The Little Prince by Antoine de Saint-Exupéry, published in 1943.
3 papers · 1 benchmark
A vast amount of information in the biomedical domain is available as natural language free text.
3 papers · 0 benchmarks
Title2Event is a large-scale sentence-level dataset for benchmarking Open Event Extraction without restricting event types.
3 papers · 0 benchmarks
The ToolLens dataset consists of 18,770 concise yet intentionally multifaceted queries, each associated with 1 to 3 verified tools out of a total of 464, designed to better mimic real-world user interactions.
3 papers · 1 benchmark
TriBERT dataset consists of 12,049 training, 2,527 validation and 2,560 test Human-Machine collaborative texts.
3 papers · 1 benchmark
Dataset of restaurant reviews from TripAdvisor that includes images and texts uploaded in reviews by users.
3 papers · 0 benchmarks
Nowadays, individuals tend to engage in dialogues with Large Language Models, seeking answers to their questions.
3 papers · 0 benchmarks
Detecting out-of-context media, such as "mis-captioned" images on Twitter, is a relevant problem, especially in domains of high public significance.
3 papers · 0 benchmarks
TyDiP (A Dataset for Politeness Classification in Nine Typologically Diverse Languages)
A Dataset for Politeness Classification in Nine Typologically Diverse Languages (TyDiP) is a dataset containing three-way politeness annotations for 500 examples in each language, totaling 4.5K examples.
3 papers · 0 benchmarks
UIIS10K (General Underwater Image Instance Segmentation dataset 10K)
We propose a large-scale underwater instance segmentation dataset, UIIS10K, which includes 10,048 images with pixel-level annotations for 10 categories.
3 papers · 0 benchmarks
The dataset comprises 4,500 question-answer pairs collected from trusted medical sources, with at least one answer and at most four unique paraphrased answers per question
3 papers · 0 benchmarks
The UIT-ViWikiQA is a dataset for evaluating sentence extraction-based machine reading comprehension in the Vietnamese language.
3 papers · 0 benchmarks
UPFD-GOS (User Preference-aware Fake News Detection)
The Gossipcop variant of the UPFD dataset for benchmarking.
3 papers · 1 benchmark
UzWordnet (The Uzbek Wordnet)
UzWordnet is a lexical-semantic database, or a “word-net”, for the (Northern) Uzbek language (native: O’zbek till) compatible with Princeton Wordnet.
3 papers · 0 benchmarks
ViMQ is a Vietnamese dataset of medical questions from patients with sentence-level and entity-level annotations for the Intent Classification and Named Entity Recognition tasks.
3 papers · 0 benchmarks
Video Localized Narratives is a new form of multimodal video annotations connecting vision and language.
3 papers · 0 benchmarks
Hugging Face Datasets (New!) | Website | Github Repository | arXiv e-Print The Visual Writing Prompts (VWP) dataset contains almost 2K selected sequences of movie shots, each including 5-10 images.
3 papers · 0 benchmarks
WDC-Dialogue is a dataset built from the Chinese social media to train EVA.
3 papers · 0 benchmarks
Test-driven benchmark to challenge LLMs to write JavaScript React application GitHub Script
3 papers · 1 benchmark
The WebVid-CoVR dataset is a collection of video-text-video triplets that can be used for the task of composed video retrieval (CoVR).
3 papers · 1 benchmark
Dataset Description The dataset described in the provided text is focused on social media polls collected from Weibo, a popular Chinese microblogging platform.
3 papers · 3 benchmarks
WiLI-2018 is a benchmark dataset for monolingual written natural language identification.
3 papers · 0 benchmarks
A dataset of single-sentence edits crawled from Wikipedia.
3 papers · 0 benchmarks
The WikiScenes dataset consists of paired images and language descriptions capturing world landmarks and cultural sites, with associated 3D models and camera poses.
3 papers · 0 benchmarks
WikiText-TL-39 is a benchmark language modeling dataset in Filipino that has 39 million tokens in the training set.
3 papers · 0 benchmarks
WikiWiki is a dataset for understanding entities and their place in a taxonomy of knowledge—their types.
3 papers · 0 benchmarks
Wikipedia Title is a dataset for learning character-level compositionality from the character visual characteristics.
3 papers · 0 benchmarks
XLEnt consists of parallel entity mentions in 120 languages aligned with English.
3 papers · 0 benchmarks
We present XHate-999, a multi-domain and multilingual evaluation data set for abusive language detection.
3 papers · 0 benchmarks
YTD-18M is a large-scale corpus of 18M video-based dialogues, constructed from web videos: crucial to the data collection pipeline is a pretrained language model that converts error-prone automatic transcripts to a cleaner dialogue format…
3 papers · 0 benchmarks
Youtbean is a dataset created from closed captions of YouTube product review videos.
3 papers · 0 benchmarks
This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages.
3 papers · 0 benchmarks
esXNLI is a bilingual NLI dataset.
3 papers · 0 benchmarks
A large, crowd-sourced dataset for the Native Language Identification (NLI) task.
3 papers · 1 benchmark
xMIND (A Multilingual Dataset for Cross-lingual News Recommendation)
xMIND is an open, large-scale multilingual news dataset for multi- and cross-lingual news recommendation.
3 papers · 0 benchmarks
30MQA (30M Factoid Question-Answer Corpus)
An enormous question answer pair corpus produced by applying a novel neural network architecture on the knowledge base Freebase to transduce facts into natural language questions.
2 papers · 0 benchmarks
5k_presetation_slides (5000 presentation slide pairs)
We crawled 5000 paper, slide pairs from conference proceeding websites.
2 papers · 0 benchmarks
AG’s Corpus (AG's corpus of news articlesNews)
Antonio Gulli’s corpus of news articles is a collection of more than 1 million news articles.
2 papers · 0 benchmarks
AIDA/testc is a new challenging test set for entity linking systems containing 131 Reuters news articles published between December 5th and 7th, 2020.
2 papers · 1 benchmark
ALM-Bench (All Languages Matter Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
APE (Automatic Post-Editing)
APE is useful to evaluate Machine Translation automatic post-editing (APE), which is the task of improving the output of a blackbox MT system by automatically fixing its mistakes.
2 papers · 0 benchmarks
ATUE is an antibody study benchmark with four real-world supervised tasks covering therapeutic antibody engineering, B cell analysis, and antibody discovery.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.