Papers › SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio...
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
Young Jin Ahn, Jungwoo Park, Sangha Park, Jonghyun Choi, Kee-Eung Kim
Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visually similar lip gestures that represent different phonemes. Prior approaches have sought to distinguish fine-grained visemes by aligning visual and auditory semantics, but often fell short of full synchronization. To address this, we present SyncVSR, an end-to-end learning framework that leverages quantized audio for frame-level crossmodal supervision. By integrating a projection layer that synchronizes visual representation with acoustic data, our encoder learns to generate discrete audio tokens from a video sequence in a non-autoregressive manner. SyncVSR shows versatility across tasks, languages, and modalities at the cost of a forward pass. Our empirical evaluations show that it not only achieves state-of-the-art results but also reduces data usage by up to ninefold.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Landmark-based Lipreading | LRS2 | SyncVSR | Word Error Rate (WER) | 74.6 | #1 of 1 | Archive leaderboard | report |
| Landmark-based Lipreading | LRW | SyncVSR (Word Boundary) | Top 1 Accuracy | 80.3 | #1 of 5 | Archive leaderboard | report |
| Landmark-based Lipreading | LRW | SyncVSR | Top 1 Accuracy | 75.1 | #2 of 5 | Archive leaderboard | report |
| Lipreading | CAS-VSR-W1k (LRW-1000) | SyncVSR (Word Boundary) | Top-1 Accuracy | 58.2 | #1 of 9 | Archive leaderboard | report |
| Lipreading | LRS2 | SyncVSR | Word Error Rate (WER) | 16.5 | #3 of 25 | Archive leaderboard | report |
| Lipreading | LRS2 | SyncVSR | Word Error Rate (WER) | 28.9 | #11 of 25 | Archive leaderboard | report |
| Lipreading | LRS3-TED | SyncVSR | Word Error Rate (WER) | 21.5 | #3 of 23 | Archive leaderboard | report |
| Lipreading | LRS3-TED | SyncVSR | Word Error Rate (WER) | 31.2 | #12 of 23 | Archive leaderboard | report |
| Lipreading | Lip Reading in the Wild | SyncVSR (Word Boundary) | Top-1 Accuracy | 95.0 | #1 of 22 | Archive leaderboard | report |
| Lipreading | Lip Reading in the Wild | SyncVSR | Top-1 Accuracy | 93.2 | #3 of 22 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections