Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 28 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 1297–1344 of 3,130

ARC-DA (ARC Direct Answer Questions)
ARC Direct Answer Questions (ARC-DA) dataset consists of 2,985 grade-school level, direct-answer ("open response", "free form") science questions derived from the ARC multiple-choice question set released as part of the AI2 Reasoning…
4 papers · 0 benchmarks
ATIS (vi) (Vietnamese Intent Detection and Slot Filling)
This is a dataset for intent detection and slot filling for the Vietnamese language.
4 papers · 2 benchmarks
Advising Corpus is a dataset based on an entirely new collection of dialogues in which university students are being advised which classes to take.
4 papers · 1 benchmark
AndroidHowTo contains 32,436 data points from 9,893 unique How-To instructions and split into training (8K), validation (1K) and test (900).
4 papers · 0 benchmarks
ArCOV19-Rumors is an Arabic COVID-19 Twitter dataset for misinformation detection composed of tweets containing claims from 27th January till the end of April 2020.
4 papers · 0 benchmarks
AutoChart is a dataset for chart-to-text generation, a task that consists on generating analytical descriptions of visual plots.
4 papers · 0 benchmarks
Bentham (Bentham project)
Bentham manuscripts refers to a large set of documents that were written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832).
4 papers · 1 benchmark
CAsT-snippets is a high-quality dataset for conversational information seeking containing snippet-level annotations for all queries in the TREC CAsT 2020 and 2022 datasets.
4 papers · 0 benchmarks
CC-DBP is a dataset for knowledge base population research using Common Crawl and DBpedia.
4 papers · 0 benchmarks
CCPE-M (Coached Conversational Preference Elicitation dataset for Movies)
A dataset consisting of 502 English dialogs with 12,000 annotated utterances between a user and an assistant discussing movie preferences in natural language.
4 papers · 0 benchmarks
CCv2 (Casual Conversations v2)
Casual Conversations v2 (CCv2) is composed of over 5,567 participants (26,467 videos) and intended mainly to be used for assessing the performance of already trained models in computer vision and audio applications for the purposes…
4 papers · 0 benchmarks
CHOCOLATE (Captions Have Often ChOsen Lies About The Evidence)
CHOCOLATE is a benchmark for detecting and correcting factual inconsistency in generated chart captions.
4 papers · 4 benchmarks
CI-MNIST (Correlated and Imbalanced MNIST)
CI-MNIST (Correlated and Imbalanced MNIST) is a variant of MNIST dataset with introduced different types of correlations between attributes, dataset features, and an artificial eligibility criterion.
4 papers · 0 benchmarks
CODA-19 is a human-annotated dataset that denotes the Background, Purpose, Method, Finding/Contribution, and Other for 10,966 English abstracts in the COVID-19 Open Research Dataset.
4 papers · 0 benchmarks
COMPARE is a taxonomy and a dataset of comparison discussions in peer reviews of research papers in the domain of experimental deep learning.
4 papers · 0 benchmarks
CS1QA is a dataset for code-based question answering in the programming education domain.
4 papers · 0 benchmarks
CSPubSum is a dataset for summarisation of computer science publications, created by exploiting a large resource of author provided summaries and show straightforward ways of extending it further.
4 papers · 0 benchmarks
CUGE is a Chinese Language Understanding and Generation Evaluation benchmark with the following features: (1) Hierarchical benchmark framework, where datasets are principally selected and organized with a language capability-task-dataset…
4 papers · 0 benchmarks
Cambridge Law Corpus (The Cambridge Law Corpus: A Dataset for Legal AI Research)
We introduce the Cambridge Law Corpus (CLC), a corpus for legal AI research.
4 papers · 0 benchmarks
The Chilean Waiting List corpus comprises de-identified referrals from the waiting list in Chilean public hospitals.
4 papers · 1 benchmark
CiteSum is a large-scale scientific extreme summarization benchmark.
4 papers · 1 benchmark
CoDesc is a large dataset of 4.2m Java source code and parallel data of their description from code search, and code summarization studies.
4 papers · 2 benchmarks
CoNaLa-Ext (CoNaLa Extended With Question Text)
The CoNaLa Extended With Question Text is an extension to the original CoNaLa Dataset (Papers With Code Link) proposed in the NLP4Prog workshop paper "Reading StackOverflow Encourages Cheating: Adding Question Text Improves Extractive Code…
4 papers · 1 benchmark
CoVERT (A Corpus of Fact-checked Biomedical COVID-19 Tweets)
CoVERT is a fact-checked corpus of tweets with a focus on the domain of biomedicine and COVID-19-related (mis)information.
4 papers · 0 benchmarks
CommitChronicle is a dataset for commit message generation (and/or completion).
4 papers · 0 benchmarks
CompMix is a crowdsourced QA benchmark which naturally demands the integration of a mixture of input sources.
4 papers · 0 benchmarks
Concadia is a publicly available Wikipedia-based corpus, which consists of 96,918 images with corresponding English-language descriptions, captions, and surrounding context.
4 papers · 0 benchmarks
ConvRef is a conversational QA benchmark with reformulations.
4 papers · 0 benchmarks
D3 (DBLP Discovery Dataset)
DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues.
4 papers · 0 benchmarks
DAGW (Danish Gigaword)
It’s hard to develop good tools for processing Danish with computers when no large and wide-coverage dataset of Danish text is readily available.
4 papers · 0 benchmarks
DDRel is a dataset for interpersonal relation classification in dyadic dialogues.
4 papers · 1 benchmark
DISRPT2019 (DISRPT2019 shared task on Discourse Unit Segmentation and Connective Detection)
The DISRPT 2019 workshop introduces the first iteration of a cross-formalism shared task on discourse unit segmentation.
4 papers · 0 benchmarks
MIT licenseDPCSpell-Bangla-SEC-Corpus is a large-scale parallel corpus for Bangla spelling error correction.
4 papers · 1 benchmark
DSTC7 Task 2 (Dialog System Technology Challenges Task 2)
DSTC Task 2 is a dataset and task for end-to-end conversation modeling.
4 papers · 0 benchmarks
DaReCzech (Dataset for text relevance ranking in Czech)
DareCzech DaReCzech is a dataset for text relevance ranking in Czech.
4 papers · 1 benchmark
A large-scale anime image database with 4.2m+ images annotated with 130m+ text tags describing image contents in detail; it can be useful for machine learning purposes such as image recognition and generation.
4 papers · 0 benchmarks
Databricks Dolly 15k (databricks-dolly-15k)
Databricks Dolly 15k is a dataset containing 15,000 high-quality human-generated prompt / response pairs specifically designed for instruction tuning large language models.
4 papers · 0 benchmarks
DiS-ReX is a multilingual dataset for distantly supervised (DS) relation extraction (RE).
4 papers · 0 benchmarks
DiSCQ (Discharge Summary Clinical Questions)
DiSCQ is a newly curated question dataset composed of 2,000+ questions paired with the snippets of text (triggers) that prompted each question.
4 papers · 0 benchmarks
DiaASQ (Conversational Aspect-based Sentiment Quadruple Extraction)
DiaASQ is a fine-grained Aspect-based Sentiment Analysis (ABSA) benchmark under the conversation scenario.
4 papers · 2 benchmarks
Diamante is a novel and efficient framework consisting of a data collection strategy and a learning method to boost the performance of pre-trained dialogue models.
4 papers · 0 benchmarks
The images in DukeMTMC-attribute dataset comes from Duke University.
4 papers · 1 benchmark
EMU (Edited Media Understanding)
48k question-answer pairs written in rich natural language.
4 papers · 0 benchmarks
EVJVQA (English-Japanese-Vietnamese Visual Question Answering)
EVJVQA, the first multilingual Visual Question Answering dataset with three languages: English, Vietnamese, and Japanese, is released in this task.
4 papers · 0 benchmarks
FACTIFY (a dataset on multi-modal fact verification)
FACTIFY is a dataset on multi-modal fact verification.
4 papers · 0 benchmarks
FLAG3D is a large-scale 3D fitness activity dataset with language instruction containing 180K sequences of 60 categories.
4 papers · 0 benchmarks
FRMT (Few-shot Region-aware Machine Translation)
FRMT is a dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation.
4 papers · 4 benchmarks
FacetSum is a faceted summarization dataset for scientific documents.
4 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.