Papers › VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models

VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models

20 May 2023arXiv:2305.12199archive 2025-07-28

Dao Xuan-Quy, Le Ngoc-Bich, Vo The-Duy, Phan Xuan-Dung, Ngo Bac-Bien, Nguyen Van-Tien, Nguyen Thi-My-Thanh, Nguyen Hong-Phuoc

The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article. The dataset, which covers nine subjects, was generated from the Vietnamese National High School Graduation Examination and comparable tests. 300 literary essays have been included, and there are over 19,000 multiple-choice questions on a range of topics. The dataset assesses LLMs in multitasking situations such as question answering, text generation, reading comprehension, visual question answering, and more by including both textual data and accompanying images. Using ChatGPT and BingChat, we evaluated LLMs on the VNHSGE dataset and contrasted their performance with that of Vietnamese students to see how well they performed. The results show that ChatGPT and BingChat both perform at a human level in a number of areas, including literature, English, history, geography, and civics education. They still have space to grow, though, especially in the areas of mathematics, physics, chemistry, and biology. The VNHSGE dataset seeks to provide an adequate benchmark for assessing the abilities of LLMs with its wide-ranging coverage and variety of activities. We intend to promote future developments in the creation of LLMs by making this dataset available to the scientific community, especially in resolving LLMs' limits in disciplines involving mathematics and the natural sciences.

PaperPDFCode

Code

xdao85/vnhsge officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Multiple-choiceQuestion AnsweringQuestion RewritingReading ComprehensionText GenerationVisual Question Answering

Datasets

Introduced by this paper, per the archive.

VNHSGE

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Question Answering VNHSGE Mathematics Bing Chat Accuracy 60 #1 of 2 Archive leaderboard report
Question Answering VNHSGE Mathematics ChatGPT Accuracy 58.8 #2 of 2 Archive leaderboard report
Question Answering VNHSGE-Biology Bing Chat Accuracy 69 #1 of 2 Archive leaderboard report
Question Answering VNHSGE-Biology ChatGPT Accuracy 58 #2 of 2 Archive leaderboard report
Question Answering VNHSGE-Chemistry Bing Chat Accuracy 52.5 #1 of 2 Archive leaderboard report
Question Answering VNHSGE-Chemistry ChatGPT Accuracy 48 #2 of 2 Archive leaderboard report
Question Answering VNHSGE-Civic Bing Chat Accuracy 85.5 #1 of 2 Archive leaderboard report
Question Answering VNHSGE-Civic ChatGPT Accuracy 70.5 #2 of 2 Archive leaderboard report
Question Answering VNHSGE-English Bing Chat Accuracy 92.4 #1 of 3 Archive leaderboard report
Question Answering VNHSGE-English ChatGPT Accuracy 79.2 #3 of 3 Archive leaderboard report
Question Answering VNHSGE-Geography Bing Chat Accuracy 85.5 #1 of 2 Archive leaderboard report
Question Answering VNHSGE-Geography ChatGPT Accuracy 61.5 #2 of 2 Archive leaderboard report
Question Answering VNHSGE-History Bing Chat Accuracy 88.5 #1 of 2 Archive leaderboard report
Question Answering VNHSGE-History ChatGPT Accuracy 56.5 #2 of 2 Archive leaderboard report
Question Answering VNHSGE-Literature ChatGPT Accuracy 68 #1 of 2 Archive leaderboard report
Question Answering VNHSGE-Literature Bing Chat Accuracy 56.8 #2 of 2 Archive leaderboard report
Question Answering VNHSGE-Physics Bing Chat Accuracy 66 #1 of 2 Archive leaderboard report
Question Answering VNHSGE-Physics ChatGPT Accuracy 61 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections