Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 49 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2305–2352 of 3,130

Fallout New Vegas Dialog is a multilingual sentiment annotated dialog dataset from Fallout New Vegas.
1 paper · 0 benchmarks
The "Famous Keyword Twitter Replies Dataset" is a comprehensive collection of Twitter data that focuses on popular keywords and their associated replies.
1 paper · 0 benchmarks
FashionRec (Fashion Recommendation Dataset)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
FedNLP (FOMC Docs and Speeches)
We collect the various forms of Federal Reserve communications.
1 paper · 0 benchmarks
FewDR is a dataset for Few-shot dense retrieval (DR).
1 paper · 0 benchmarks
FewGLUE_64_labeled (A new version of FewGLUE with 64 training examples)
Introduction The FewGLUE64labeled dataset is a new version of FewGLUE dataset.
1 paper · 0 benchmarks
Filipino CrowS-Pairs and Filipino WinoQueer assess sexist and homophobic biases in language models handling Filipino.
1 paper · 0 benchmarks
FinDKG: The Global Financial Dynamic Knowledge Graph Dataset FinDKG is an open-source dataset focused on creating a temporally-resolved Financial Dynamic Knowledge Graph.
1 paper · 0 benchmarks
Financial Language Understanding Evaluation is an open-source comprehensive suite of benchmarks for the financial domain.
1 paper · 0 benchmarks
First HAREM (Primeiro HAREM)
HAREM, an initiative by Linguateca, boasts a Golden Collection—a meticulously curated repository of annotated Portuguese texts.
1 paper · 0 benchmarks
This dataset contains: (1) Slforge Generated Simulink Models : Synthetic Simulink Models (2) Source of Real World Simulink Models The .txt file is a combined text file that contains all the real world Simulink models based on SLGPT's…
1 paper · 0 benchmarks
FFR Dataset is an ongoing project to collect, clean and store corpora of Fon and French sentences for machine translation from Fon-French.
1 paper · 0 benchmarks
The Food Recall Incidents dataset consists of 7,546 short texts (from 5 to 360 characters each), which are the titles of food recall announcements (therefore referred to as title), crawled from 24 public food safety authority websites by…
1 paper · 0 benchmarks
Food.com Recipes and Interactions consists of 270K recipes and 1.4M user-recipe interactions (reviews) scraped from Food.com, covering a period of 18 years (January 2000 to December 2018).
1 paper · 0 benchmarks
This dataset contains news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY.
1 paper · 0 benchmarks
Open-source dataset
1 paper · 0 benchmarks
FreCDo (French cross-domain)
FreCDo is a corpus for French dialect identification comprising 413,522 French text samples collected from public news websites in Belgium, Canada, France and Switzerland.
1 paper · 0 benchmarks
GASP is a dataset composed by a list of cited abstracts associated with the corresponding source abstract.
1 paper · 0 benchmarks
GD-NLI (Generated Debiased NLI Datasets)
This is a set of debiased Natural Language Inference (NLI) datasets produced by the paper Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets.
1 paper · 0 benchmarks
Dataset Card for Dataset Name This dataset is a filtered version of BookCorpus containing only gender-neutral words.
1 paper · 0 benchmarks
GENTER (GEnder Name TEmplates with pRonouns)
This dataset consists of template sentences associating first names ([NAME]) with third-person singular pronouns ([PRONOUN]), e.g., [NAME] asked , not sounding as if [PRONOUN] cared about the answer .
1 paper · 0 benchmarks
GENTYPES (Gender Stereotypes)
This dataset contains short sentences linking a first name, represented by the template mask [NAME], to stereotypical associations.
1 paper · 0 benchmarks
The released GIF Reply dataset contains 1,562,701 real text-GIF conversation turns on Twitter.
1 paper · 1 benchmark
GLAMI-1M (A Multilingual Image-Text Fashion Dataset)
We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark.
1 paper · 1 benchmark
GLARE (Guided LexRank for Advanced Retrieval in Legal Analysis)
The Guided Lexrank algorithm is applied to dataset specialappeal.csv to summarize the texts of legal documents.
1 paper · 0 benchmarks
GLARE is an Arabic Apps Reviews dataset collected from Saudi Google PlayStore.
1 paper · 0 benchmarks
GPR-bench (General‑Purpose Reproducibility Benchmark)
GPR‑bench is an open‑source, multilingual benchmark for regression testing and reproducibility tracking in generative‑AI systems.
1 paper · 0 benchmarks
GPTKB is a large general-domain knowledge base (KB) constructed entirely from a large language model (LLM).
1 paper · 0 benchmarks
GQNLI-FR is a manually translated French version of the GQNLI challenge dataset, originally written in English.
1 paper · 0 benchmarks
Gambling Address Dataset is a collection of 10,423 gambling addresses that have transactions with gambling contracts.
1 paper · 0 benchmarks
Gambling Contract Dataset is a collection of 260 gambling smart contracts from decentralized gambling websites, such as Dicether, Degens.
1 paper · 0 benchmarks
GameQA (GameQA-140K)
GameQA is a large-scale, diverse, and challenging multimodal reasoning dataset designed to enhance the general reasoning capabilities of Vision Language Models (VLMs).
1 paper · 0 benchmarks
GameWikiSum is a domain-specific (video game) dataset for multi-document summarization, which is one hundred times larger than commonly used datasets, and in another domain than news.
1 paper · 0 benchmarks
GeBiD (Geometric shapes Bimodal Dataset)
We provide a custom synthetic bimodal dataset, called GeBiD, designed specifically for the comparison of the joint- and cross-generative capabilities of Multimodal Variational Autoencoders.
1 paper · 0 benchmarks
This dataset encompasses 265 speeches (over 200,000 tokens) from the German Bundestag, primarily from the 19th legislative term (2017-2021), given by 195 distinct speakers representing 6 political parties.
1 paper · 2 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
GenAIPABench is a specialized dataset designed to evaluate Generative AI-based Privacy Assistants (GenAIPAs).
1 paper · 0 benchmarks
GenPlot (GenPlot: 500k pre-generated plots)
This dataset contains the pre-generated dataset referenced in the GenPlot Paper.
1 paper · 0 benchmarks
Genomics Adversarial Attack Sample dataset
1 paper · 0 benchmarks
GenoTEX (An LLM Agent Benchmark for Automated Gene Expression Data Analysis)
GenoTEX (Genomics Data Automatic Exploration Benchmark) is a benchmark dataset for the automated analysis of gene expression data to identify disease-associated genes while considering the influence of other biological factors.
1 paper · 0 benchmarks
GeoJEPAD (GeoJEPA Dataset)
GeoJEPAD is a multimodal dataset combining OpenStreetMap (OSM) data (attributes and geometries) with high-resolution aerial imagery from diverse urban areas.
1 paper · 0 benchmarks
GeoQuestions1089 is a crowdsourced geospatial question-answering dataset that targets the Knowledge Graph YAGO2geo.
1 paper · 1 benchmark
GerMS-AT (GERMS-AT: A Sexism/Misogyny Dataset of Forum Comments from an Austrian Online Newspaper)
This dataset contains 7984 user comments from an Austrian online newspaper.
1 paper · 2 benchmarks
The Gigaword Entailment dataset is a dataset for entailment prediction between an article and its headline.
1 paper · 0 benchmarks
Collection of news websites in low-resource languages.
1 paper · 0 benchmarks
A Brazilian Portuguese TTS dataset featuring a female voice recorded with high quality in a controlled environment, with neutral emotion and more than 20 hours of recordings.
1 paper · 0 benchmarks
A database containing high sampling rate recordings of a single speaker reading sentences in Brazilian Portuguese with neutral voice, along with the corresponding text corpus.
1 paper · 0 benchmarks
Google Local review (Google Local Data)
Description This Dataset contains review information on Google map (ratings, text, images, etc.), business metadata (address, geographical info, descriptions, category information, price, open hours, and MISC info), and links (relative…
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.