Papers › Hierarchical Pre-training for Sequence Labelling in Spoken Dialog
Hierarchical Pre-training for Sequence Labelling in Spoken Dialog
Emile Chapuis, Pierre Colombo, Matteo Manica, Matthieu Labeau, Chloe Clavel
Sequence labelling tasks like Dialog Act and Emotion/Sentiment identification are a key component of spoken dialog systems. In this work, we propose a new approach to learn generic representations adapted to spoken dialog, which we evaluate on a new benchmark we call Sequence labellIng evaLuatIon benChmark fOr spoken laNguagE benchmark (\texttt{SILICONE}). \texttt{SILICONE} is model-agnostic and contains 10 different datasets of various sizes. We obtain our representations with a hierarchical encoder based on transformer architectures, for which we extend two well-known pre-training objectives. Pre-training is performed on OpenSubtitles: a large corpus of spoken dialog containing over $2.3$ billion of tokens. We demonstrate how hierarchical encoders achieve competitive results with consistently fewer parameters compared to state-of-the-art models and we show their importance for both pre-training and fine-tuning.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Dialogue Act Classification | ICSI Meeting Recorder Dialog Act (MRDA) corpus | Pretrained Hierarchical Transformer | Accuracy | 92.4 | #1 of 8 | Archive leaderboard | report |
| Dialogue Act Classification | Switchboard corpus | Pretrained Hierarchical Transformer | Accuracy | 79.2 | #8 of 11 | Archive leaderboard | report |
| Emotion Recognition in Conversation | DailyDialog | Pretrained Hierarchical Transformer | Micro-F1 | 60.14 | #6 of 22 | Archive leaderboard | report |
| Emotion Recognition in Conversation | IEMOCAP | Pretrained Hierarchical Transformer | Accuracy | 66.05 | #42 of 59 | Archive leaderboard | report |
| Emotion Recognition in Conversation | IEMOCAP | Pretrained Hierarchical Transformer | Weighted-F1 | 65.37 | #42 of 59 | Archive leaderboard | report |
| Emotion Recognition in Conversation | MELD | Pretrained Hierarchical Transformer | Weighted-F1 | 61.90 | #51 of 68 | Archive leaderboard | report |
| Emotion Recognition in Conversation | SEMAINE | Pretrained Hierarchical Transformer | MAE (Arousal) | 0.16 | #2 of 3 | Archive leaderboard | report |
| Emotion Recognition in Conversation | SEMAINE | Pretrained Hierarchical Transformer | MAE (Expectancy) | 0.16 | #2 of 3 | Archive leaderboard | report |
| Emotion Recognition in Conversation | SEMAINE | Pretrained Hierarchical Transformer | MAE (Power) | 7.70 | #2 of 3 | Archive leaderboard | report |
| Emotion Recognition in Conversation | SEMAINE | Pretrained Hierarchical Transformer | MAE (Valence) | 0.16 | #2 of 3 | Archive leaderboard | report |
| Text Classification | SILICONE Benchmark | Pretrained Hierarchical Transformer | 1:1 Accuracy | 71.25 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections