Papers › Multi-Task Learning for Audio Visual Active Speaker Detection

Multi-Task Learning for Audio Visual Active Speaker Detection

1 Jun 2019The ActivityNet Large-Scale Activity Recognition Challenge Workshop, CVPR 2019 6archive 2025-07-28

Yuanhang Zhang, Jingyun Xiao, Shuang Yang, Shiguang Shan

This report describes the approach underlying our submission to the active speaker detection task (task B-2) of ActivityNet Challenge 2019. We introduce a new audio-visual model which builds upon a 3D-ResNet18 visual model pretrained for lipreading and a VGG-M acoustic model pretrained for audio-to-video synchronization. The model is trained with two losses in a multi-task learning fashion: a contrastive loss to enforce matching between audio and video features for active speakers, and a regular crossentropy loss to obtain speaker / non-speaker labels. This model obtains 84.0% mAP on the validation set of AVAActiveSpeaker. Experimental results showcase the pretrained embeddings' abilities to transfer across tasks and data formats, as well as the advantage of the proposed multi-task learning strategy.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Active Speaker DetectionAudio-Visual Active Speaker DetectionLipreadingMulti-Task LearningVideo Synchronization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio-Visual Active Speaker Detection AVA-ActiveSpeaker 3D-ResNet-GRU validation mean average precision 84.0% #20 of 20 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections