Browse State-of-the-Art › Active Speaker Detection
Active Speaker Detection
30 papers with code · 1 benchmark · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| LRS3-TED (1 row) | GestSync | GestSync: Determining who is speaking without a talking head | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 30 papers with code (63 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection14 Jul 2021 4 repositories listed Syntology ran 2 of 9 samples · 7 unverified · 1 pointer-only (licence)Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers.
-
27 Mar 2022 3 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 5 pointer-only (licence)Recent advances in the Active Speaker Detection (ASD) problem build upon a two-stage process: feature extraction and spatio-temporal context aggregation.
-
19 Jan 2023 2 repositories listedThese two contexts are complementary to each other and can help infer the active speaker.
-
15 Jul 2022 2 repositories listedActive speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows.
-
5 Jan 2019 2 repositories listedThe dataset contains temporally labeled face tracks in video, where each face instance is labeled as speaking or not, and whether the speech is audible.
-
28 May 2025 1 repository listedWe present UniTalk, a novel dataset specifically designed for the task of active speaker detection, emphasizing challenging scenarios to enhance model generalization.
-
6 May 2025 1 repository listedThey also enable strong performance in Visual Speech Recognition (VSR) with a WER of 20.
-
21 Jan 2025 1 repository listedActive Speaker Detection (ASD) aims to identify speaking individuals in complex visual scenes.
-
11 Dec 2024 1 repository listedState-of-the-art Active Speaker Detection (ASD) approaches mainly use audio and facial features as input.
-
6 Dec 2024 1 repository listedThe results show that BIAS is state-of-the-art in challenging conditions where body-based features are of utmost importance (Columbia, open-settings, and WASD), and yields competitive results in AVA-ActiveSpeaker, where…
-
20 Nov 2024 1 repository listedActive speaker detection (ASD) in multimodal environments is crucial for various applications, from video conferencing to human-robot interaction.
-
16 Jul 2024 1 repository listedHead movements are crucial for social human-human interaction.
-
20 Feb 2024 1 repository listedIn order to promote research on low-resource languages for audio-visual speech technologies, we present AnnoTheia, a semi-automatic annotation toolkit that detects when a person speaks on the scene and the corresponding…
-
21 Dec 2023 1 repository listedThe multichannel audio ``student'' network is trained to generate the same results.
-
8 Oct 2023 1 repository listedIn this paper we introduce a new synchronisation task, Gesture-Sync: determining if a person's gestures are correlated with their speech or not.
-
21 Sep 2023 1 repository listedThe goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames.
-
22 May 2023 1 repository listedTo benefit from both facial cue and reference speech, we propose the Target Speaker TalkNet (TS-TalkNet), which leverages a pre-enrolled speaker embedding to complement the audio-visual synchronization cue in detecting…
-
9 Mar 2023 1 repository listedThe results show that: 1) AVA trained models maintain a state-of-the-art performance in WASD Easy group, while underperforming in the Hard one, showing the 2) similarity between AVA and Easy data; and 3) training in…
-
8 Mar 2023 1 repository listed Syntology ran 5 of 6 samples · 1 unverifiedExperimental results on the AVA-ActiveSpeaker dataset show that our framework achieves competitive mAP performance (94.
-
1 Dec 2022 1 repository listedActive speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality.
-
23 Nov 2022 1 repository listedFinally, we devise a model for emotion recognition in conversations trained on the realigned MELD-FAIR videos, which outperforms state-of-the-art models for ERC based on vision alone.
-
24 Sep 2022 1 repository listedWe leverage speaker identity information from speech and faces, and formulate active speaker detection as a speech-face assignment task such that the active speaker's face and the underlying speech identify the same…
-
4 Mar 2022 1 repository listedActive speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding.
-
17 Aug 2021 1 repository listedFace tracks are extracted from the videos and active segments are annotated based on the timestamps of VoxConverse in a semi-automatic way.
-
7 Jun 2021 1 repository listedSuccessful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers…
-
1 Jun 2021 1 repository listedActive speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers.
-
11 Jan 2021 1 repository listedActive speaker detection requires a solid integration of multi-modal cues.
-
10 Aug 2020 1 repository listed Syntology ran 1 of 14 samples · 13 unverifiedOur objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning.
-
20 May 2020 1 repository listedCurrent methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker.
-
28 Feb 2020 1 repository listedFurthermore, neuroscience has successfully identified the superior colliculus region in the brain as the one responsible for this modality fusion, with a handful of biological models having been proposed to approach its…
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections