Home › Datasets › language › Vietnamese
Vietnamese datasets
archive 2025-07-28
71 datasets carry the language tag "Vietnamese", ordered by the archive's paper count. Page 2 of 2: 23 shown of 71. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Vietnamese datasets 49–71 of 71
MGSM8KInstruct, the multilingual math reasoning instruction dataset, encompassing ten distinct languages, thus addressing the issue of training data scarcity in multilingual math reasoning.
2 papers · 0 benchmarks
The dataset contains training and evaluation data for 12 languages: - Vietnamese - Romanian - Latvian - Czech - Polish - Slovak - Irish - Hungarian - French - Turkish - Spanish - Croatian For each language, one training, one development…
2 papers · 12 benchmarks
UIT-ViSFD (Vietnamese Aspect-Based Sentiment Analysis Dataset)
UIT-ViSFD is a Vietnamese Smartphone Feedback Dataset as a new benchmark corpus built based on strict annotation schemes for evaluating aspect-based sentiment analysis, consisting of 11,122 human-annotated comments for mobile e-commerce,…
2 papers · 0 benchmarks
VNDS (VNDS: A Vietnamese Dataset for Summarization)
A single-document Vietnamese summarization dataset
2 papers · 1 benchmark
ViText2SQL is a dataset for the Vietnamese Text-to-SQL semantic parsing task, consisting of about 10K question and SQL query pairs.
2 papers · 0 benchmarks
In AISIA-VN-Review-S and AISIA-VN-Review-F datasets, we first collect 450K customer reviewing comments from various e–commerce websites.
1 paper · 0 benchmarks
BKEE (BKEE: Pioneering Event Extraction in the Vietnamese Language)
A novel event extraction dataset for Vietnamese.
1 paper · 0 benchmarks
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks
MVALUE (Multilingual human VALUE dataset)
Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and…
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
UIT-ViCoQA (Conversational machine reading comprehension in the Vietnamese language)
UIT-ViCoQA is a new corpus for conversational machine reading comprehension in the Vietnamese language.
1 paper · 0 benchmarks
VNEMOS (Vietnamese Speech Emotion Dataset)
This research introduces the dataset that we created to test voice emotional recognition models with Vietnamese data.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ViSR (Vietnamese Synthetic Reasoning)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
ViTHSD (Vietnamese Targeted-Hate-Speech-Detection)
A Vietnamese dataset for hate speech detection by the specific target.
1 paper · 0 benchmarks
Spoken Named Entity Recognition (NER) aims to extracting named entities from speech and categorizing them into types like person, location, organization, etc.
1 paper · 0 benchmarks
In doctor-patient conversations, identifying medically relevant information is crucial, posing the need for conversation summarization.
1 paper · 0 benchmarks
We introduce a first Vietnamese Spelling Correction dataset containing manual labelling mistakes and corresponding correct words.
1 paper · 0 benchmarks
VlogQA (Vietnamese Spoken-Based Machine Reading Comprehension)
The VlogQA consists of 10,076 question-answer pairs based on 1,230 transcript documents sourced from YouTube - an extensive source of user-uploaded content, covering the topics of food and travel in the Vietnamese language.
1 paper · 0 benchmarks
WEATHub is a dataset containing 24 languages.
1 paper · 0 benchmarks
This dataset contains images and annotations for scene text detection and recognition.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.