Papers › Empirical Interpretation of Speech Emotion Perception with Attention Based Model for...

Empirical Interpretation of Speech Emotion Perception with Attention Based Model for Speech Emotion Recognition

28 Oct 2020Interspeech 2020 10archive 2025-07-28

Md AsifJalal, Rosanna Milner, Thomas Hain Speech

Speech emotion recognition is essential for obtaining emotional intelligence which affects the understanding of context and meaning of speech. Harmonically structured vowel and consonant sounds add indexical and linguistic cues in spoken in- formation. Previous research argued whether vowel sound cues were more important in carrying the emotional context from a psychological and linguistic point of view. Other research also claimed that emotion information could exist in small over- lapping acoustic cues. However, these claims are not corroborated in computational speech emotion recognition systems. In this research, a convolution-based model and a long-short- term memory-based model, both using attention, are applied to investigate these theories of speech emotion on computational models. The role of acoustic context and word importance is demonstrated for the task of speech emotion recognition. The IEMOCAP corpus is evaluated by the proposed models, and 80.1% unweighted accuracy is achieved on pure acoustic data which is higher than current state-of-the-art models on this task. The phones and words are mapped to the attention vectors and it is seen that the vowel sounds are more important for defining emotion acoustic cues than the consonants, and the model can assign word importance based on acoustic context

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Emotion RecognitionEmotional IntelligenceSpeech Emotion Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Emotion Recognition IEMOCAP SYSCOMB: BLSTMATT with CSA (session5) F1 - #5 of 8 Archive leaderboard report
Speech Emotion Recognition IEMOCAP SYSCOMB: BLSTMATT with CSA (session5) UA 0.740 #5 of 8 Archive leaderboard report
Speech Emotion Recognition IEMOCAP SYSCOMB: BLSTMATT with CSA (session5) WA 0.805 #5 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections