Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 63 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 2977–3024 of 3,130

The Winograd schema challenge composes tasks with syntactic ambiguity, which can be resolved with logic and reasoning.
1 paper · 1 benchmark
WordNet-feelings, is an affective dataset that identifies 3664 word senses as feelings, and associates each of these with one of the 9 categories of feeling.
1 paper · 0 benchmarks
The goal of this dataset is to understand how people experience sexism and sexual harassment in the workplace by discovering themes in 2,362 experiences posted on the Everyday Sexism Project's website Source:…
1 paper · 0 benchmarks
We present the World Wide Dishes dataset which seeks to assess disparities in representations of food through a decentralised data collection effort to gather perspectives directly from people with a wide variety of backgrounds from around…
1 paper · 0 benchmarks
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
We provide a new data set XWikiRef for the task of Cross-lingual Multi-document Summarization.
1 paper · 0 benchmarks
Xamarin Q&A consists of two datasets of questions and answers for studying the development of cross-platform mobile applications using the Xamarin framework.
1 paper · 0 benchmarks
XiaChuFang Recipe Corpus contains recipes are from 下厨房 (XiaChuFang), a popular Chinese recipe sharing website.
1 paper · 0 benchmarks
YesBut Dataset (https://yesbut-dataset.github.io) Understanding satire and humor is a challenging task for even current Vision-Language models.
1 paper · 0 benchmarks
YouwikiHow is a dataset for Weakly-Supervised temporal Article Grounding (WSAG).
1 paper · 0 benchmarks
The first and the one open dataset for Russian finger- spelling, contained 1,593 annotated phrases and over 37 thousand HD+ videos.
1 paper · 1 benchmark
abc_cc (ABC CC)
Dataset Summary The dataset used to train and evaluate TunesFormer is collected from two sources: The Session and ABCnotation.com.
1 paper · 0 benchmarks
approved_drug_target (Approved Drug SMILES and Protein Sequence Dataset)
This dataset provides a curated collection of approved drug Simplified Molecular Input Line Entry System (SMILES) strings and their associated protein sequences.
1 paper · 0 benchmarks
This is a second public release of the arXMLiv dataset generated by the KWARC research group.
1 paper · 0 benchmarks
arXiv Categories (arXiv Categories Multi-label Text Classification Dataset)
This is a dataset of scientific documents derived from arXiv.
1 paper · 0 benchmarks
bSDD (buildingSMART Data Dictionary)
The buildingSMART Data Dictionary (bSDD) is an online service that hosts classifications and their properties, allowed values, units and translations.
1 paper · 0 benchmarks
bajer_danish_misogyny (Bajer Online Misogyny)
This is a high-quality dataset of annotated posts sampled from social media posts and annotated for misogyny.
1 paper · 1 benchmark
bigscience/P3 (bigscience/P3, split='ai2_arc_ARC_Challenge_pick_the_most_correct_option')
This datasets consists of challenging reasoning questions in multiple choice format.
1 paper · 0 benchmarks
blbooks (The British Library Books)
This dataset consists of books digitised by the British Library in partnership with Microsoft.
1 paper · 0 benchmarks
catbAbI LM-mode (concatenated-bAbI)
We aim to improve the bAbI benchmark as a means of developing intelligent dialogue agents.
1 paper · 1 benchmark
catbAbI QA-mode (concatenated-bAbI)
We aim to improve the bAbI benchmark as a means of developing intelligent dialogue agents.
1 paper · 1 benchmark
Download free fonts in DaFont style from our extensive collection.
1 paper · 0 benchmarks
Release decompile-ghidra-100k, a subset of 100k training samples (25k per optimization level).
1 paper · 0 benchmarks
diaforge-utc-r-0725 (DiaFORGE UTC: Unified Tool-Calling Conversations Dataset)
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks
Diderot’s Encyclopédie is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era.
1 paper · 0 benchmarks
ec-darkpattern is a dataset for dark pattern detection and prepared its baseline detection performance with state-of-the-art machine learning methods.
1 paper · 0 benchmarks
fake (Real / Fake Job Posting Prediction)
[Real or Fake] : Fake Job Description Prediction This dataset contains 18K job descriptions out of which about 800 are fake.
1 paper · 1 benchmark
A collection of natural language prompt-completion pairs pertaining to multiple-choice Q&A on benchmark tasks based on US census products.
1 paper · 0 benchmarks
gENder-IT is an English-Italian challenge set focusing on the resolution of natural gender phenomena by providing word-level gender tags on the English source side and multiple gender alternative translations, where needed, on the Italian…
1 paper · 0 benchmarks
Online web communities often face bans for violating platform policies, encouraging their migration to alternative platforms.
1 paper · 0 benchmarks
This is the Big-Bench version of our language-based movie recommendation dataset https://github.com/google/BIG-bench/tree/main/bigbench/benchmarktasks/movierecommendation GPT-2 has a 48.8% accuracy, chance is 25%.
1 paper · 1 benchmark
This dataset links all the entries describing named entities of Petit Larousse illustré, a French dictionary published in 1905, to wikidata identifiers.
1 paper · 0 benchmarks
medisim is a collection of new large-scale medical term similarity datasets based on SNOMED-CT.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks
🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.
1 paper · 0 benchmarks
We introduce misinfo-general, a benchmark dataset for evaluating misinformation models’ ability to perform out-of-distribution generalisation.
1 paper · 0 benchmarks
A modification on the ShEMO dataset with help of an Automatic Speech Recognition (ASR) system.
1 paper · 0 benchmarks
needadvice is a dataset for advice classification extracted from Reddit.
1 paper · 0 benchmarks
[Dataset on HF] [Project Page] [Subjective LeaderBoard] [Objective LeaderBoard] CriticBench is a novel benchmark designed to comprehensively and reliably evaluate the critique abilities of Large Language Models (LLMs).
1 paper · 0 benchmarks
A collection of datasets and benchmarks for large-scale Performance Modeling with LLMs.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
In this dataset we teleoperated UR5 arm to collect manipulation data for picking up a screwdriver in a cluttered tabletop environment.
1 paper · 0 benchmarks
polstance (Political Stance in Danish)
Political stance in Danish.
1 paper · 0 benchmarks
This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP, also called verbal probabilities), e.g.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.