Browse State-of-the-Art › Audio-Visual Synchronization
Audio-Visual Synchronization
11 papers with code · 0 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
11 shown of 11 papers with code (32 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Jan 2024 2 repositories listed Syntology ran 2 of 11 samples · 9 unverifiedOur objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse.
-
27 Oct 2022 2 repositories listedThis paper proposed an MTDVocaLiST model, which is trained by our proposed multimodal Transformer distillation (MTD) loss.
-
13 Oct 2022 2 repositories listedThis contrasts with the case of synchronising videos of talking heads, where audio-visual correspondence is dense in both time and space.
-
6 May 2025 1 repository listedThey also enable strong performance in Visual Speech Recognition (VSR) with a WER of 20.
-
19 Dec 2024 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedWe propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio.
-
14 Oct 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverified · 8 pointer-only (licence)Real-time video dubbing that preserves identity consistency while achieving accurate lip synchronization remains a critical challenge.
-
30 Apr 2024 1 repository listedWith the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes.
-
10 Apr 2024 1 repository listedRecent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks.
-
22 May 2023 1 repository listedTo benefit from both facial cue and reference speech, we propose the Target Speaker TalkNet (TS-TalkNet), which leverages a pre-enrolled speaker embedding to complement the audio-visual synchronization cue in detecting…
-
5 Apr 2022 1 repository listedFinally, we use the frozen visual features learned by our lip synchronisation model in the singing voice separation task to outperform a baseline audio-visual model which was trained end-to-end.
-
5 Oct 2021 1 repository listedModifying the pitch and timing of an audio signal are fundamental audio editing operations with applications in speech manipulation, audio-visual synchronization, and singing voice editing and synthesis.
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections