Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 44 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2065–2112 of 3,130

Contains a large number of online videos and subtitles.
1 paper · 0 benchmarks
The Apiza Corpus is a WoZ-like (Wizard of Oz) set of dialogues between 30 programmers and a simulated virtual assistant.
1 paper · 0 benchmarks
The AppealCase dataset is the first large-scale resource specifically designed to support LegalAI research in appellate judgment scenarios.
1 paper · 0 benchmarks
Sentiment analysis is pivotal in Natural Language Processing for understanding opinions and emotions in text.
1 paper · 0 benchmarks
ArVoice (ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis)
We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration,…
1 paper · 0 benchmarks
AraCovid19-SSD is a manually annotated Arabic COVID-19 sarcasm and sentiment detection dataset containing 5,162 tweets.
1 paper · 0 benchmarks
Digital Edition: Essays from Hannah Arendt We have created a NER dataset from the digital edition "Sechs Essays" by Hannah Arendt.
1 paper · 0 benchmarks
ArgSciChat is an argumentative dialogue dataset.
1 paper · 2 benchmarks
This dataset contains 2,360 paraphrases in Armenian that can be used for paraphrase detection.
1 paper · 0 benchmarks
The task of Visual Question Answering (VQA) has been studied extensively on general-domain real-world images.
1 paper · 1 benchmark
AskParents is a dataset for advice classification extracted from Reddit.
1 paper · 0 benchmarks
Collects all the courses from XuetangX5, one of the largest MOOCs in China, and this results in 1951 courses.
1 paper · 0 benchmarks
This is a benchmark for neural paraphrase detection, to differentiate between original and machine-generated content.
1 paper · 0 benchmarks
This dataset is used to evaluate a predictive consent model for users’ information shared in social media.
1 paper · 0 benchmarks
Dataset Overview: 998 images and 4,208 annotations focusing on interaction with in-vehicle infotainment (IVI) systems.
1 paper · 0 benchmarks
For more details see https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset
1 paper · 0 benchmarks
AuxAD is a a distantly supervised dataset for acronym disambiguation.
1 paper · 0 benchmarks
AuxAI is a distantly supervised dataset for acronym identification.
1 paper · 0 benchmarks
AviationQA is introduced in the paper titled- There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering The paper is accepted in the main conference of ICON 2022.
1 paper · 1 benchmark
BAH (Behavioural Ambivalence/Hesitancy)
Recognizing complex emotions linked to ambivalence and hesitancy (A/H) can play a critical role in the personalization and effectiveness of digital behaviour change interventions.
1 paper · 0 benchmarks
BASIR (BASIR_Budget_Assisted_Sectoral_Impact_Ranking)
Government fiscal policies, particularly annual union budgets, exert significant influence on financial markets.
1 paper · 0 benchmarks
BBAI Dataset (Black-box Agent Integration)
This dataset is for evaluating the task of Black-box Multi-agent Integration which focuses on combining the capabilities of multiple black-box conversational agents at scale.
1 paper · 1 benchmark
BCWS (Bilingual Contextual Word Similarity)
Dataset for evaluating English-Chinese Bilingual Contextual Word Similarity.
1 paper · 0 benchmarks
BEAMetrics (Benchmark to Evaluate Automatic Metrics) is resource to make research into new metrics for evaluation of generated language easier to evaluate.
1 paper · 0 benchmarks
BEAR-probe (Benchmark for Evaluating Associative Reasoning)
The BEAR dataset and its larger version, BEARbig, are benchmarks for evaluating common factual knowledge contained in language models.
1 paper · 0 benchmarks
BLANCA (Benchmarks for LANguage models on Coding Artifacts) is a collection of benchmarks that assess code understanding based on tasks such as predicting the best answer to a question in a forum post, finding related forum posts, or…
1 paper · 0 benchmarks
BLM-17m is a labeled dataset for topic detection that contains 17 million tweets.
1 paper · 0 benchmarks
BLN600 (BLN600: A Parallel Corpus of Machine/Human Transcribed Nineteenth Century Newspaper Texts)
A publicly available corpus of nineteenth-century newspaper text focused on crime in London, derived from the Gale British Library Newspapers corpus parts 1 and 2.
1 paper · 0 benchmarks
BLP (Blackout Poetry Dataset)
A blackout poetry dataset constructed from publicly available short stories and large poems.
1 paper · 1 benchmark
BN-AuthProf (Bangla Author Profiling Dataset)
Although research on author profiling has quite progressed in abundant resources languages, it is still infancy for limited resources languages such as Bengali.
1 paper · 1 benchmark
This is a dataset for Bengali Captioning from Images.
1 paper · 0 benchmarks
BPersona-chat is an evaluation dataset based on the English multiturn chat corpus Persona-chat and the Japanese multiturn chat corpus JPersona-chat.
1 paper · 0 benchmarks
BS-Objaverse 660k Dataset is a set of GPT4-Vision-powered multi-modal captions data.
1 paper · 0 benchmarks
The BWB corpus consists of Chinese novels translated by experts into English, and the annotated test set is designed to probe the ability of machine translation systems to model various discourse phenomena.
1 paper · 0 benchmarks
A Bambara dialectal dataset dedicated for Sentiment Analysis, available freely for Natural Language Processing research purposes
1 paper · 0 benchmarks
BanglaBook (Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews)
This repository contains the code, data, and models of the paper titled "BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews" published in the Findings of the Association for Computational Linguistics: ACL…
1 paper · 1 benchmark
BanglaEmotion (BanglaEmotion: A Benchmark Dataset for Bangla Textual Emotion Analysis)
BanglaEmotion is a manually annotated Bangla Emotion corpus, which incorporates the diversity of fine-grained emotion expressions in social-media text.
1 paper · 0 benchmarks
The dataset identifies the shortcomings of existing benchmarks in evaluating the problem of compositional generalization, which underscores the need for the development of datasets tailored to assess compositional generalization in open…
1 paper · 1 benchmark
Beemo (Benchmark of expert-edited machine-generated outputs)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Belfort (The Belfort dataset: Handwritten Text Recognition from Crowdsourced Annotations)
The Belfort dataset This dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946.
1 paper · 1 benchmark
BenBench is designed to benchmark the potential for data leakage in benchmark datasets, which can lead to biased and inequitable comparisons.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
The dataset contains 36000 Bangla data based on Ekman's six basic emotions.
1 paper · 1 benchmark
Our dataset, BSMDD, was collected from various open social media platforms and translated and annotated by native Bengali speakers with expertise in both language and mental health.
1 paper · 0 benchmarks
BestRev (Understanding Peer Review of Software Engineering Papers)
Survey instrument, analysis code, and anonymized responses for the paper on review practices in SE.
1 paper · 0 benchmarks
Bianet is a parallel news corpus in Turkish, Kurdish and English It contains 3,214 Turkish articles with their sentence-aligned Kurdish or English translations from the Bianet online newspaper.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.