Papers › Hierarchical Pronunciation Assessment with Multi-Aspect Attention

Hierarchical Pronunciation Assessment with Multi-Aspect Attention

15 Nov 2022arXiv:2211.08102archive 2025-07-28

Heejin Do, Yunsu Kim, Gary Geunbae Lee

Automatic pronunciation assessment is a major component of a computer-assisted pronunciation training system. To provide in-depth feedback, scoring pronunciation at various levels of granularity such as phoneme, word, and utterance, with diverse aspects such as accuracy, fluency, and completeness, is essential. However, existing multi-aspect multi-granularity methods simultaneously predict all aspects at all granularity levels; therefore, they have difficulty in capturing the linguistic hierarchy of phoneme, word, and utterance. This limitation further leads to neglecting intimate cross-aspect relations at the same linguistic unit. In this paper, we propose a Hierarchical Pronunciation Assessment with Multi-aspect Attention (HiPAMA) model, which hierarchically represents the granularity levels to directly capture their linguistic structures and introduces multi-aspect attention that reflects associations across aspects at the same level to create more connotative representations. By obtaining relational information from both the granularity- and aspect-side, HiPAMA can take full advantage of multi-task learning. Remarkable improvements in the experimental results on the speachocean762 datasets demonstrate the robustness of HiPAMA, particularly in the difficult-to-assess aspects.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

doheejin/HiPAMA officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Multi-Task LearningPhone-level pronunciation scoringUtterance-level pronounciation scoringWord-level pronunciation scoring

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Phone-level pronunciation scoring speechocean762 HiPAMA-Librispeech Pearson correlation coefficient (PCC) 0.62 #6 of 8 Archive leaderboard report
Utterance-level pronounciation scoring speechocean762 HiPAMA-Librispeech Pearson correlation coefficient (PCC) 0.754 #3 of 5 Archive leaderboard report
Word-level pronunciation scoring speechocean762 HiPAMA-Librispeech Pearson correlation coefficient (PCC) 0.59 #4 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections