Home › Datasets › modality › Texts
Texts datasets
archive 2025-07-28
3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 38 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Texts datasets 1777–1824 of 3,130
FINDSum (Financial Report Document Summarization)
FINDSum is a large-scale dataset for long text and multi-table summarization.
2 papers · 0 benchmarks
FSVQA (Full-Sentence Visual Question Answering)
Full-Sentence Visual Question Answering (FSVQA) dataset, consisting of nearly 1 million pairs of questions and full-sentence answers for images, built by applying a number of rule-based natural language processing techniques to original…
2 papers · 0 benchmarks
FairPrism is a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harms relating to gender and sexuality.
2 papers · 0 benchmarks
FarsBase-KBP contains 22015 sentences, in which the entities and relation types are linked to the FarsBase ontology.
2 papers · 0 benchmarks
Fashion-MNT is large-scale bilingual product description dataset called Fashion-MMT, which contains over 114k noisy and 40k manually cleaned description translations with multiple product images.
2 papers · 0 benchmarks
📄 Read 💾 Code 🔗 Webpage 💻 Demo 🤗 Huggingface Dataset 💬 Discussions Overview Users interact with QA systems and leave feedback.
2 papers · 0 benchmarks
FinBench is a benchmark for evaluating the performance of machine learning models with both tabular data inputs and profile text inputs.
2 papers · 0 benchmarks
This dataset enriches the benchmark Room-to-Room (R2R) dataset by dividing the instructions into sub-instructions and pairing each of those with their corresponding viewpoints in the path.
2 papers · 0 benchmarks
FinnSentiment introduces a 27,000 sentence dataset (in Finnish) annotated independently with sentiment polarity by three native annotators.
2 papers · 0 benchmarks
Given a sentence in the source language, generate a translation in the target language.
2 papers · 0 benchmarks
FormulaNet FormulaNet is a new large-scale Mathematical Formula Detection dataset.
2 papers · 0 benchmarks
This dataset is dialog dataset collected in a Wizard-of-Oz fashion.
2 papers · 0 benchmarks
FuLG is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully extracted from Common Crawl.
2 papers · 0 benchmarks
G-VUE (General-purpose Visual Understanding Evaluation)
General-purpose Visual Understanding Evaluation (G-VUE) is a comprehensive benchmark covering the full spectrum of visual cognitive abilities with four functional domains -- Perceive, Ground, Reason, and Act.
2 papers · 0 benchmarks
GATITOS (Google's Additional Translations Into Tail-languages: Often Short)
The GATITOS (Google's Additional Translations Into Tail-languages: Often Short) dataset is a high-quality, multi-way parallel dataset of tokens and short phrases, intended for training and improving machine translation models.
2 papers · 0 benchmarks
GEOBench-VLM, a comprehensive benchmark specifically designed to evaluate VLMs on geospatial tasks, including scene understanding, object counting, localization, fine-grained categorization, and temporal analysis.
2 papers · 0 benchmarks
GIRT-Data (GitHub Issue Report Template Dataset)
GIRT-Data is the first and largest dataset of issue report templates (IRTs) in both YAML and Markdown format.
2 papers · 0 benchmarks
GPI corpus (Government Privacy Instructions Corpus)
The GPI Corpus is a collection of 1,043 privacy laws, regulations, and guidelines ("GPIs") covering 182 jurisdictions around the world.
2 papers · 0 benchmarks
GQA-OOD is a new dataset and benchmark for the evaluation of VQA models in OOD (out of distribution) settings.
2 papers · 0 benchmarks
Geoclidean-Elements dataset is derived from definitions in the first book of Euclid’s Elements, which focuses on plane geometry.
2 papers · 0 benchmarks
The GermEval dataset is a valuable resource for natural language processing (NLP) tasks, specifically named entity recognition (NER), conducted in the German language.
2 papers · 0 benchmarks
GermanDPR is a dataset for passage retrieval in German.
2 papers · 0 benchmarks
The Gun Violence Corpus (GVC) consists of 241 unique incidents for which we have structured data on a) location, b) time c) the name, gender and age of the victims and d) the status of the victims after the incident: killed or injured.
2 papers · 0 benchmarks
HERDPhobia is an annotated hate speech detection dataset on Fulani herders in Nigeria -- in three languages: English, Nigerian-Pidgin, and Hausa.
2 papers · 0 benchmarks
HLGD (Headline Grouping Dataset)
The Headline Grouping dataset is a binary classification dataset on pairs of news headline.
2 papers · 0 benchmarks
HalluEditBench is a comprehensive benchmark for evaluating knowledge editing methods' effectiveness in correcting real-world hallucinations.
2 papers · 0 benchmarks
Hansel is a human-annotated Chinese entity linking (EL) dataset, focusing on tail entities and emerging entities: - The test set contains Few-shot (FS) and zero-shot (ZS) slices, has 10K examples and uses Wikidata as the corresponding…
2 papers · 0 benchmarks
This is a Twitter dataset of 100,386 users along with up to 200 tweets from their timelines with a random-walk-based crawler on the retweet graph, with a subsample of 4,972 which is manually annotated as hateful or not through…
2 papers · 0 benchmarks
HeriGraph (Multimodal Machine Learning Datasets on Graphs of Heritage Values and Attributes)
The dataset contains constructed multi-modal features (visual and textual), pseudo-labels (on heritage values and attributes), and graph structures (with temporal, social, and spatial links) constructed using User-Generated Content data…
2 papers · 0 benchmarks
The Horne 2017 Fake News Data contains two independed news datasets: 1.
2 papers · 0 benchmarks
Timely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data, a process that can highly benefit from expert-assisted NLP systems trained on validated and annotated data in the…
2 papers · 0 benchmarks
Human Simulacra is a virtual character dataset that contains 129k texts across 11 virtual characters, with each character having unique attributes, biographies, and stories.
2 papers · 0 benchmarks
IG-3.5B-17k is an internal Facebook AI Research dataset for training image classification models.
2 papers · 0 benchmarks
The IRFL dataset consists of idioms, similes, and metaphors with matching figurative and literal images, as well as two novel tasks of multimodal figurative understanding and preference.
2 papers · 2 benchmarks
The ISOT Fake News dataset is a compilation of several thousands fake news and truthful articles, obtained from different legitimate news sites and sites flagged as unreliable by Politifact.com.
2 papers · 0 benchmarks
Iconary dataset is for testing multimodal communication with drawings and text.
2 papers · 0 benchmarks
InLegalNER is a corpus of 46545 annotated legal named entities mapped to 14 legal entity types.
2 papers · 1 benchmark
IndiaPoliceEvents is a corpus of 21,391 sentences from 1,257 English-language Times of India articles about events in the state of Gujarat during March 2002.
2 papers · 0 benchmarks
InfiniteBench (∞Bench: Extending Long Context Evaluation Beyond 100K Tokens)
Introduction Welcome to InfiniteBench, a cutting-edge benchmark tailored for evaluating the capabilities of language models to process, understand, and reason over super long contexts (100k+ tokens).
2 papers · 0 benchmarks
InspiRe (Inspiring and non-inspiring posts from Reddit)
We analyze social media posts to tease out what makes a post inspiring and what topics are inspiring.
2 papers · 0 benchmarks
Instantiation is a dataset for the task of instantiation detection
2 papers · 0 benchmarks
JAMUL (JApanese MUlti-Length Headline Corpus)
A large-scale evaluation dataset for headlines of three different lengths composed by professional editors.
2 papers · 0 benchmarks
JEMMA is an Extensible Java Dataset for ML4Code Applications, which is a large-scale dataset targeted at ML4 code.
2 papers · 0 benchmarks
The Jejueo Interview Transcripts (JIT) dataset is a parallel corpus containing 170k+ Jejueo-Korean sentences.
2 papers · 0 benchmarks
JNC (Japanese News Corpus)
The JNC data provides common supervision data for headline generation.
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
This dataset contains information about Japanese word similarity including rare words.
2 papers · 0 benchmarks
K-SportsSum is a sports game summarization dataset with two characteristics: (1) K-SportsSum collects a large amount of data from massive games.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.