Home › Datasets › language › Vietnamese

Vietnamese datasets

archive 2025-07-28

71 datasets carry the language tag "Vietnamese", ordered by the archive's paper count. Page 1 of 2: 48 shown of 71. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets

Vietnamese datasets 1–48 of 71

MultiNLI (Multi-Genre Natural Language Inference)
The Multi-Genre Natural Language Inference (MultiNLI) dataset has 433K sentence pairs.
1,830 papers · 4 benchmarks
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
XNLI (Cross-lingual Natural Language Inference)
The Cross-lingual Natural Language Inference (XNLI) corpus is the extension of the Multi-Genre NLI (MultiNLI) corpus to 15 languages.
349 papers · 7 benchmarks
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance.
190 papers · 1 benchmark
MLQA (MultiLingual Question Answering)
MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance.
167 papers · 1 benchmark
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
The LIMA dataset is a valuable resource used in natural language processing (NLP) research.
107 papers · 0 benchmarks
MusicCaps is a dataset composed of 5.5k music-text pairs, with rich text descriptions provided by human experts.
84 papers · 1 benchmark
WikiANN (PAN-X)
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
ICDAR 2015 was a scene text detection used for the ICDAR 2015 conference.
52 papers · 2 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
MKQA (Multilingual Knowledge Questions and Answers)
Multilingual Knowledge Questions and Answers (MKQA) is an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically diverse languages (260k question-answer pairs in total).
49 papers · 0 benchmarks
IGLUE (Image-Grounded Language Understanding Evaluation)
The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded…
31 papers · 0 benchmarks
XM 3600 (Crossmodal 3600)
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
A new dataset for the low-resource language as Vietnamese to evaluate MRC models.
15 papers · 0 benchmarks
ViHSD (Vietnamese Hate Speech Detection Dataset)
This dataset contains 33,400 annotated comments used for hate speech detection on social network sites.
12 papers · 0 benchmarks
Synbols is a dataset generator designed for probing the behavior of learning algorithms.
11 papers · 0 benchmarks
This is the dataset for the 2020 Duolingo shared task on Simultaneous Translation And Paraphrase for Language Education (STAPLE).
10 papers · 0 benchmarks
PhoMT is a high-quality and large-scale Vietnamese-English parallel dataset of 3.02M sentence pairs for machine translation.
7 papers · 1 benchmark
UIT-ViCTSD (UIT Vietnamese Constructive and Toxic Speech Detection)
UIT-ViCTSD (Vietnamese Constructive and Toxic Speech Detection) is a dataset for constructive and toxic speech detection in Vietnamese.
7 papers · 0 benchmarks
VIVOS (VIVOS Corpus)
VIVOS is a free Vietnamese speech corpus consisting of 15 hours of recording speech prepared for Automatic Speech Recognition task.
7 papers · 1 benchmark
VNHSGE (VietNamese High School Graduation Examination Dataset for Large Language Models)
The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article.
7 papers · 9 benchmarks
ViHOS (Hate Speech Spans Detection for Vietnamese)
The first human-annotated corpus containing 26k spans on 11k comments
7 papers · 0 benchmarks
UIT-ViIC contains manually written captions for images from Microsoft COCO dataset relating to sports played with ball.
6 papers · 0 benchmarks
ViNLI (Vietnamese Natural Language Inference Dataset)
A large-scale and high-quality corpus is necessary for studies on NLI for Vietnamese, which can be considered a low-resource language.
6 papers · 1 benchmark
UIT-ViNewsQA is a new corpus for the Vietnamese language to evaluate healthcare reading comprehension models.
5 papers · 0 benchmarks
ATIS (vi) (Vietnamese Intent Detection and Slot Filling)
This is a dataset for intent detection and slot filling for the Vietnamese language.
4 papers · 2 benchmarks
EVJVQA (English-Japanese-Vietnamese Visual Question Answering)
EVJVQA, the first multilingual Visual Question Answering dataset with three languages: English, Vietnamese, and Japanese, is released in this task.
4 papers · 0 benchmarks
MultiSpider is a large multilingual text-to-SQL dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese).
4 papers · 0 benchmarks
OpenViVQA (Open-domain Visual Question Answering in Vietnamese)
In recent years, visual question answering (VQA) has attracted attention from the research community because of its highly potential applications (such as virtual assistance on intelligent cars, assistant devices for blind people, or…
4 papers · 0 benchmarks
UIT-VSMEC (Vietnamese Social Media Emotion Corpus)
Emotion recognition is a higher approach or special case of sentiment analysis.
4 papers · 0 benchmarks
ViMMRC (Vietnamese Multiple-choice Machine Reading Comprehension Corpus)
A challenging machine comprehension corpus with multiple-choice questions, intended for research on the machine comprehension of Vietnamese text.
4 papers · 0 benchmarks
VietMed (VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain)
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
4 papers · 2 benchmarks
DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants.
3 papers · 0 benchmarks
PhoNERCOVID19 is a dataset for recognising COVID-19 related named entities in Vietnamese, consisting of 35K entities over 10K sentences.
3 papers · 1 benchmark
TyDiP (A Dataset for Politeness Classification in Nine Typologically Diverse Languages)
A Dataset for Politeness Classification in Nine Typologically Diverse Languages (TyDiP) is a dataset containing three-way politeness annotations for 500 examples in each language, totaling 4.5K examples.
3 papers · 0 benchmarks
The dataset comprises 4,500 question-answer pairs collected from trusted medical sources, with at least one answer and at most four unique paraphrased answers per question
3 papers · 0 benchmarks
The UIT-ViWikiQA is a dataset for evaluating sentence extraction-based machine reading comprehension in the Vietnamese language.
3 papers · 0 benchmarks
ViMQ is a Vietnamese dataset of medical questions from patients with sentence-level and entity-level annotations for the Intent Classification and Named Entity Recognition tasks.
3 papers · 0 benchmarks
ViSpamReviews (Vietnamese Spam Reviews Detection)
This dataset is used for spam review detection (opinion spam reviews) on Vietnamese E-commerce website
3 papers · 0 benchmarks
xMIND (A Multilingual Dataset for Cross-lingual News Recommendation)
xMIND is an open, large-scale multilingual news dataset for multi- and cross-lingual news recommendation.
3 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.