Papers › Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection

Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection

15 Jul 2022arXiv:2207.07783archive 2025-07-28

Kyle Min, Sourya Roy, Subarna Tripathi, Tanaya Guha, Somdeb Majumdar

Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we present SPELL, a novel spatial-temporal graph learning framework that can solve complex tasks such as ASD. To this end, each person in a video frame is first encoded in a unique node for that frame. Nodes corresponding to a single person across frames are connected to encode their temporal dynamics. Nodes within a frame are also connected to encode inter-person relationships. Thus, SPELL reduces ASD to a node classification task. Importantly, SPELL is able to reason over long temporal contexts for all nodes without relying on computationally expensive fully connected graph neural networks. Through extensive experiments on the AVA-ActiveSpeaker dataset, we demonstrate that learning graph-based representations can significantly improve the active speaker detection performance owing to its explicit spatial and temporal structure. SPELL outperforms all previous state-of-the-art approaches while requiring significantly lower memory and computational resources. Our code is publicly available at https://github.com/SRA2/SPELL

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

kylemin/SPELL officialmentioned in papermentioned on GitHubpytorch report
sra2/spell officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Active Speaker DetectionAudio-Visual Active Speaker DetectionGraph LearningNode Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio-Visual Active Speaker Detection AVA-ActiveSpeaker SPELL+ validation mean average precision 94.9% #4 of 20 Archive leaderboard report
Audio-Visual Active Speaker Detection AVA-ActiveSpeaker SPELL validation mean average precision 94.2% #6 of 20 Archive leaderboard report
Node Classification AVA ASDNet [ASDNet_ICCV2021] mAP 93.5 #1 of 4 Archive leaderboard report
Node Classification AVA TalkNet [tao2021someone] mAP 92.3 #2 of 4 Archive leaderboard report
Node Classification AVA UniCon [zhang2021unicon] mAP 92 #3 of 4 Archive leaderboard report
Node Classification AVA MAAS-TAN [MAAS2021] mAP 88.8 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections