Home › Datasets › modality › Texts

Texts datasets

archive 2025-07-28

3,130 datasets carry the modality tag "Texts", ordered by the archive's paper count. Page 6 of 66: 48 shown of 3,130. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Texts datasets 241–288 of 3,130

The CoNLL-2012 shared task involved predicting coreference in English, Chinese, and Arabic, using the final version, v5.0, of the OntoNotes corpus.
89 papers · 0 benchmarks
PathVQA consists of 32,799 open-ended questions from 4,998 pathology images where each question is manually checked to ensure correctness.
89 papers · 0 benchmarks
Logical reasoning is an important ability to examine, analyze, and critically evaluate arguments as they occur in ordinary language as the definition from Law School Admission Council.
88 papers · 4 benchmarks
ASPEC (Asian Scientific Paper Excerpt Corpus)
ASPEC, Asian Scientific Paper Excerpt Corpus, is constructed by the Japan Science and Technology Agency (JST) in collaboration with the National Institute of Information and Communications Technology (NICT).
87 papers · 0 benchmarks
This dataset is for evaluating the performance of intent classification systems in the presence of "out-of-scope" queries, i.e., queries that do not fall into any of the system-supported intent classes.
87 papers · 5 benchmarks
KP20k is a large-scale scholarly articles dataset with 528K articles for training, 20K articles for validation and 20K articles for testing.
87 papers · 3 benchmarks
FLoRes-101 is an evaluation benchmark for low-resource and multilingual machine translation.
86 papers · 57 benchmarks
The MovieQA dataset is a dataset for movie question answering.
86 papers · 1 benchmark
A large-scale and machine-generated dataset of 274,186 toxic and benign statements about 13 minority groups.
85 papers · 0 benchmarks
CodeContests is a competitive programming dataset for machine-learning.
84 papers · 1 benchmark
The How2 dataset contains 13,500 videos, or 300 hours of speech, and is split into 185,187 training, 2022 development (dev), and 2361 test utterances.
84 papers · 2 benchmarks
MiniF2F is a dataset of formal Olympiad-level mathematics problems statements intended to provide a unified cross-system benchmark for neural theorem proving.
84 papers · 2 benchmarks
MusicCaps is a dataset composed of 5.5k music-text pairs, with rich text descriptions provided by human experts.
84 papers · 1 benchmark
TweetEval introduces an evaluation framework consisting of seven heterogeneous Twitter-specific classification tasks.
84 papers · 1 benchmark
WikiBio (Wikipedia Biography Dataset)
This dataset gathers 728,321 biographies from English Wikipedia.
84 papers · 1 benchmark
GovReport is a dataset for long document summarization, with significantly longer documents and summaries.
83 papers · 2 benchmarks
NLVR (Natural Language Visual Reasoningnatural language for visual reasoning)
NLVR contains 92,244 pairs of human-written English sentences grounded in synthetic images.
83 papers · 3 benchmarks
Our task is to localize and provide a pixel-level mask of an object on all video frames given a language referring expression obtained either by looking at the first frame only or the full video.
82 papers · 1 benchmark
Douban (Douban Conversation Corpus)
We release Douban Conversation Corpus, comprising a training data set, a development set and a test set for retrieval based chatbot.
81 papers · 4 benchmarks
MetaQA (MoviE Text Audio QA)
The MetaQA dataset consists of a movie ontology derived from the WikiMovies Dataset and three sets of question-answer pairs written in natural language: 1-hop, 2-hop, and 3-hop queries.
81 papers · 1 benchmark
The ProofWriter dataset contains many small rulebases of facts and rules, expressed in English.
80 papers · 0 benchmarks
RealNews is a large corpus of news articles from Common Crawl.
80 papers · 0 benchmarks
The ReferIt dataset contains 130,525 expressions for referring to 96,654 objects in 19,894 images of natural scenes.
80 papers · 0 benchmarks
VQG (Visual Question Generation)
VQG is a collection of datasets for visual question generation.
80 papers · 1 benchmark
CoNLL-2014 will continue the CoNLL tradition of having a high profile shared task in natural language processing.
79 papers · 0 benchmarks
WIT (Wikipedia-based Image Text)
Wikipedia-based Image Text (WIT) Dataset is a large multimodal multilingual dataset.
79 papers · 1 benchmark
WikiTableQuestions is a question answering dataset over semi-structured tables.
79 papers · 2 benchmarks
RadGraph (RadGraph: Extracting Clinical Entities and Relations from Radiology Reports)
RadGraph is a dataset of entities and relations in radiology reports based on our novel information extraction schema, consisting of 600 reports with 30K radiologist annotations and 221K reports with 10.5M automatically generated…
78 papers · 0 benchmarks
CoNaLa (CMU CoNaLa, the Code/Natural Language Challenge)
The CMU CoNaLa, the Code/Natural Language Challenge dataset is a joint project from the Carnegie Mellon University NeuLab and Strudel labs.
77 papers · 1 benchmark
Few-NERD is a large-scale, fine-grained manually annotated named entity recognition dataset, which contains 8 coarse-grained types, 66 fine-grained types, 188,200 sentences, 491,711 entities, and 4,601,223 tokens.
77 papers · 3 benchmarks
ECB+ (extension to the EventCorefBank)
The ECB+ corpus is an extension to the EventCorefBank (ECB, Bejan and Harabagiu, 2010).
76 papers · 0 benchmarks
TAT-QA (Tabular And Textual dataset for Question Answering) is a large-scale QA dataset, aiming to stimulate progress of QA research over more complex and realistic tabular and textual data, especially those requiring numerical reasoning.
76 papers · 1 benchmark
The Semantic Scholar corpus (S2) is composed of titles from scientific papers published in machine learning conferences and journals from 1985 to 2017, split by year (33 timesteps).
75 papers · 0 benchmarks
TREC-COVID is a community evaluation designed to build a test collection that captures the information needs of biomedical researchers using the scientific literature during a pandemic.
73 papers · 1 benchmark
TrecQA (Text Retrieval Conference Question Answering)
Text Retrieval Conference Question Answering (TrecQA) is a dataset created from the TREC-8 (1999) to TREC-13 (2004) Question Answering tracks.
73 papers · 3 benchmarks
DS-1000 is a code generation benchmark with a thousand data science questions spanning seven Python libraries that (1) reflects diverse, realistic, and practical use cases, (2) has a reliable metric, (3) defends against memorization by…
72 papers · 0 benchmarks
MASSIVE is a parallel dataset of > 1M utterances across 51 languages with annotations for the Natural Language Understanding tasks of intent prediction and slot annotation.
72 papers · 3 benchmarks
The shared task of CoNLL-2002 concerns language-independent named entity recognition.
70 papers · 3 benchmarks
CMRC (Chinese Machine Reading Comprehension)
CMRC is a dataset is annotated by human experts with near 20,000 questions as well as a challenging set which is composed of the questions that need reasoning over multiple clues.
69 papers · 0 benchmarks
QMSum is a new human-annotated benchmark for query-based multi-domain meeting summarisation task, which consists of 1,808 query-summary pairs over 232 meetings in multiple domains.
69 papers · 1 benchmark
VSR (Visual Spatial Reasoning)
The Visual Spatial Reasoning (VSR) corpus is a collection of caption-image pairs with true/false labels.
69 papers · 1 benchmark
WikiLarge comprise 359 test sentences, 2000 development sentences and 300k training sentences.
69 papers · 0 benchmarks
BeaverTails is a dataset aimed at fostering research on safety alignment in large language models (LLMs).
68 papers · 0 benchmarks
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants.
68 papers · 0 benchmarks
DREAM is a multiple-choice Dialogue-based REAding comprehension exaMination dataset.
68 papers · 2 benchmarks
Recipe1M+ is a dataset which contains one million structured cooking recipes with 13M associated images.
68 papers · 3 benchmarks
HatEval (SemEval 2019 Task 5: Multilingual Detection of Hate Speech Against Immigrants and Women in Twitter)
Hate Speech is commonly defined as any communication that disparages a person or a group on the basis of some characteristic such as race, color, ethnicity, gender, sexual orientation, nationality, religion, or other characteristics.
67 papers · 1 benchmark
HSOL is a dataset for hate speech detection.
67 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.