Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 59 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2785–2832 of 3,130

This dataset consists of virtual scenes rendered in MuJoCo with multiple views each presented in multiple modalities: image, and synthetic or natural language descriptions.
1 paper · 0 benchmarks
SoccerNet-Echoes (SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset.
1 paper · 0 benchmarks
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity 🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain Dataset Statistics | Category | Samples | Description |…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
SoliDiffy Differencing Contract Pairs and Edit Scripts Dataset The project creates and maintains two main datasets to assist with research and evaluation of Solidity smart contract differencing: Mutated Contracts Dataset: The mutated…
1 paper · 0 benchmarks
Ensemble Tagger Training and Testing Set This data includes two files: The training set used to create the SCANL Ensemble tagger [1] and the "unseen" testing set that includes words from systems that are not available in the training set.
1 paper · 0 benchmarks
Scene Graph Generation (SGG) converts visual scenes into structured graph representations, providing deeper scene understanding for complex vision tasks.
1 paper · 0 benchmarks
Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehension (MRC) or natural language inference (NLI).
1 paper · 0 benchmarks
Spanish Corpus XIX (19th Century Spanish Corpus)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Spectral Detection and Analysis Based Paper(SDAAP) dataset is the first open-source textual knowledge dataset for spectral analysis and detection and contains annotated literature data as well as corresponding knowledge instruction data,…
1 paper · 0 benchmarks
Dataset Summary Speech Brown is a comprehensive, synthetic, and diverse paired speech-text dataset in 15 categories, covering a wide range of topics from fiction to religion.
1 paper · 0 benchmarks
Spiced is a paraphrase dataset of scientific findings annotated for degree of information change.
1 paper · 0 benchmarks
This dataset contains over 47,000 LEGO structures of over 28,000 unique 3D objects accompanied by detailed captions.
1 paper · 0 benchmarks
The dataset identifies the shortcomings of existing benchmarks in evaluating the problem of compositional generalization, which underscores the need for the development of datasets tailored to assess compositional generalization in open…
1 paper · 1 benchmark
> The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents > > Xing Han Lu, Siva Reddy, Harm de Vries > > EACL 2023 | | | | | | | :--: | :--: | :--: | :--: | :--: | | Code | Huggingface | Request on…
1 paper · 1 benchmark
When arriving at each state, each observation token gets a coin toss to see whether it will appear in the output observation string.
1 paper · 0 benchmarks
In this Coding and Coding Scheme spreadsheet: Student answers to reflection questions from pre-context workshops; coding scheme for student reflections; and coding of student reflection
1 paper · 0 benchmarks
The StudyAbroadGPT-Dataset is a collection of conversational data focused on university application requirements for various programs, including MBA, MS in Computer Science, Data Science, and Bachelor of Medicine.
1 paper · 0 benchmarks
Digital Edition: Sturm Edition Source: Schrade, Torsten: „Startseite“, in: DER STURM.
1 paper · 0 benchmarks
SubSumE Dataset This repository contains the SubSumE dataset for subjective document summarization.
1 paper · 0 benchmarks
Subjective Perception of Active Noise Reduction (SPANR) (Replication Data for: Anti-noise window: subjective perception of active noise reduction and effect of informational masking)
This repository contains replication data to the paper titled: "Anti-noise window: subjective perception of active noise reduction and effect of informational masking"
1 paper · 0 benchmarks
A large dataset of around 40000 Reddit posts was collected from r/suicidewatch and other non-suicidal subreddits.
1 paper · 0 benchmarks
The dataset contains 140 paragraphs from climate change reports with associated aspect-based (i.e.
1 paper · 0 benchmarks
SumeCzech-NER contains named entity annotations of SumeCzech 1.0, a Czech news-based summarization dataset.
1 paper · 0 benchmarks
SummZoo, a benchmark consists of 8 diverse summarization tasks with multiple sets of few-shot samples for each task, covering both monologue and dialogue domains.
1 paper · 0 benchmarks
Super-CLEVR-3D is a visual question answering (VQA) dataset where the questions are about the explicit 3D configuration of the objects from images (i.e.
1 paper · 0 benchmarks
Dataset Generation - Base Model: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2 - Seed Instructions: Selected from the databricks/databricks-dolly-15k dataset - Generation Approach: Iterative evolution of instructions using a conversational…
1 paper · 0 benchmarks
Overview The LaMini Dataset is an instruction dataset generated using h2ogpt-gm-oasst1-en-2048-falcon-40b-v2.
1 paper · 0 benchmarks
Dataset Generation - Base Model: h2oai/h2ogpt-gm-oasst1-en-2048-falcon-40b-v2 - Seed Instructions: Derived from the FLAN-v2 Collection.
1 paper · 0 benchmarks
There are 9,321 survey papers with high quality included in the SurvayBank in the domain of computer science.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Switchboard Dialog Act Corpus
1 paper · 1 benchmark
Overview: This collection contains three synthetic datasets produced by gpt-4o-mini for sentiment analysis and PDT (Product Desirability Toolkit) testing.
1 paper · 0 benchmarks
Synthetic Reasoning - Natural Language (Vietnamese Synthetic Reasoning - Natural Language)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We used the following procedure.
1 paper · 0 benchmarks
TACO-BAAI (Topics in Algorithmic Code generation dataset)
TACO (Topics in Algorithmic Code generation dataset) is a dataset focused on algorithmic code generation, designed to provide a more challenging training dataset and evaluation benchmark for the code generation model field.
1 paper · 1 benchmark
TADAC (Text Annotated Distortion, Appearance and Content Dataset)
We have developed a systematic method for constructing large text annotated image databases designed for exploiting vision-language modeling for image quality assessment and present the Text Annotated Distortion, Appearance and Content…
1 paper · 0 benchmarks
TARA is a dataset for tool-augmented reward modeling, which includes comprehensive comparison data of human preferences and detailed tool invocation processes.
1 paper · 0 benchmarks
TASTEset Recipe Dataset and Food Entities Recognition is a dataset for Named Entity Recognition (NER) which consists of 700 recipes with more than 13,000 entities to extract.
1 paper · 0 benchmarks
TBCOV is a large-scale Twitter dataset comprising more than two billion multilingual tweets related to the COVID-19 pandemic collected worldwide over a continuous period of more than one year.
1 paper · 0 benchmarks
The TED VCR Video Retrieval Dataset is a multimodal collection derived from publicly available TED Talks.
1 paper · 0 benchmarks
TF1-EN-3M (klusai/ds-tf1-en-3m)
TF1-EN-3M: Three Million Synthetic Moral Fables for Open Language Models TF1-EN-3M is a large-scale synthetic dataset of 3,000,000 English-language moral fables, generated by instruction-tuned language models with no more than 8 billion…
1 paper · 0 benchmarks
The dataset contains more than 100k code patch pairs extracted from open source projects on GitHub.
1 paper · 1 benchmark
THRED (Two-Hop Relation Extraction Dataset)
This is two-hop relation extraction dataset derived from WikiHop dataset [1].
1 paper · 0 benchmarks
TQBA++ (Tiny QA Benchmark++)
Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.
1 paper · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
1 paper · 1 benchmark
TREx-2p is a dataset to probe whether a pretrained LM possesses “indirect” 2-hop knowledge.
1 paper · 0 benchmarks
TUMTraffic-VideoQA is a novel dataset designed to understand spatiotemporal video in complex roadside traffic scenarios.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.