Papers › BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition

BERSting at the Screams: A Benchmark for Distanced, Emotional and Shouted Speech Recognition

30 Apr 2025arXiv:2505.00059archive 2025-07-28

Paige Tuttösí, Mantaj Dhillon, Luna Sang, Shane Eastwood, Poorvi Bhatia, Quang Minh Dinh, Avni Kapoor, Yewon Jin, Angelica Lim

Some speech recognition tasks, such as automatic speech recognition (ASR), are approaching or have reached human performance in many reported metrics. Yet, they continue to struggle in complex, real-world, situations, such as with distanced speech. Previous challenges have released datasets to address the issue of distanced ASR, however, the focus remains primarily on distance, specifically relying on multi-microphone array systems. Here we present the B(asic) E(motion) R(andom phrase) S(hou)t(s) (BERSt) dataset. The dataset contains almost 4 hours of English speech from 98 actors with varying regional and non-native accents. The data was collected on smartphones in the actors homes and therefore includes at least 98 different acoustic environments. The data also includes 7 different emotion prompts and both shouted and spoken utterances. The smartphones were places in 19 different positions, including obstructions and being in a different room than the actor. This data is publicly available for use and can be used to evaluate a variety of speech recognition tasks, including: ASR, shout detection, and speech emotion recognition (SER). We provide initial benchmarks for ASR and SER tasks, and find that ASR degrades both with an increase in distance and shout level and shows varied performance depending on the intended emotion. Our results show that the BERSt dataset is challenging for both ASR and SER tasks and continued work is needed to improve the robustness of such systems for more accurate real-world use.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion RecognitionSpeech Recognitionspeech-recognition

Datasets

Introduced by this paper, per the archive.

BERSt

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Emotion Recognition BERSt DAWN-hidden-SVM Unweighted Accuracy (UA) 32.1 #1 of 3 Archive leaderboard report
Speech Emotion Recognition BERSt DAWN-hidden-SVM Weighted Accuracy (WA) 32.2 #1 of 3 Archive leaderboard report
Speech Emotion Recognition BERSt Wav2Small-VAD-SVM Unweighted Accuracy (UA) 23.3 #2 of 3 Archive leaderboard report
Speech Emotion Recognition BERSt Wav2Small-VAD-SVM Weighted Accuracy (WA) 22.3 #2 of 3 Archive leaderboard report
Speech Emotion Recognition BERSt Speechbrain Wav2Vec2 Unweighted Accuracy (UA) 20.7 #3 of 3 Archive leaderboard report
Speech Emotion Recognition BERSt Speechbrain Wav2Vec2 Weighted Accuracy (WA) 20.8 #3 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Focus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections