Papers › ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations
ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations
Yuanhang Zhang, Shuang Yang, Shiguang Shan, Xilin Chen
We propose a novel strategy ES3 for self-supervised learning of robust audio-visual speech representations from unlabeled talking face videos. While many recent approaches for this task primarily rely on guiding the learning process using the audio modality alone to capture information shared between audio and video we reframe the problem as the acquisition of shared unique (modality-specific) and synergistic speech information to address the inherent asymmetry between the modalities. Based on this formulation we propose a novel "evolving" strategy that progressively builds joint audio-visual speech representations that are strong for both uni-modal (audio & visual) and bi-modal (audio-visual) speech. First we leverage the more easily learnable audio modality to initialize audio and visual representations by capturing audio-unique and shared speech information. Next we incorporate video-unique speech information and bootstrap the audio-visual representations on top of the previously acquired shared knowledge. Finally we maximize the total audio-visual speech information including synergistic information to obtain robust and comprehensive representations. We implement ES3 as a simple Siamese framework and experiments on both English benchmarks and a newly contributed large-scale Mandarin dataset show its effectiveness. In particular on LRS2-BBC our smallest model is on par with SoTA models with only 1/2 parameters and 1/8 unlabeled data (223h).
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Audio-Visual Speech Recognition | CAS-VSR-S101 | ES³ Base* | Word Error Rate (WER) | 11.0 | #1 of 1 | Archive leaderboard | report |
| Lipreading | CAS-VSR-S101 | ES³ Base* | Word Error Rate (WER) | 55.6 | #1 of 1 | Archive leaderboard | report |
| Lipreading | LRS2 | ES³ Large + extLM | Word Error Rate (WER) | 24.6 | #6 of 25 | Archive leaderboard | report |
| Lipreading | LRS2 | ES³ Large | Word Error Rate (WER) | 26.7 | #8 of 25 | Archive leaderboard | report |
| Lipreading | LRS2 | ES³ Base + extLM | Word Error Rate (WER) | 28.7 | #9 of 25 | Archive leaderboard | report |
| Lipreading | LRS2 | ES³ Base* + extLM | Word Error Rate (WER) | 29.3 | #12 of 25 | Archive leaderboard | report |
| Lipreading | LRS2 | ES³ Base | Word Error Rate (WER) | 30.7 | #13 of 25 | Archive leaderboard | report |
| Lipreading | LRS2 | ES³ Base* | Word Error Rate (WER) | 31.4 | #14 of 25 | Archive leaderboard | report |
| Lipreading | LRS3-TED | ES³ Large | Word Error Rate (WER) | 37.1 | #15 of 23 | Archive leaderboard | report |
| Lipreading | LRS3-TED | ES³ Base | Word Error Rate (WER) | 40.3 | #16 of 23 | Archive leaderboard | report |
| Speech Recognition | CAS-VSR-S101 | ES³ Base* | Word Error Rate (WER) | 11.6 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections