{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/maas-multi-modal-assignation-for-active","title":"MAAS: Multi-modal Assignation for Active Speaker Detection","arxiv_id":"2101.03682","date":"2021-01-11","proceeding":"ICCV 2021 10","authors":["Juan León-Alcázar","Fabian Caba Heilbron","Ali Thabet","Bernard Ghanem"],"abstract":"Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling their temporal progression. Despite its inherent muti-modal nature, current methods still focus on modeling and fusing short-term audiovisual features for individual speakers, often at frame level. In this paper we present a novel approach to active speaker detection that directly addresses the multi-modal nature of the problem, and provides a straightforward strategy where independent visual features from potential speakers in the scene are assigned to a previously detected speech event. Our experiments show that, an small graph data structure built from a single frame, allows to approximate an instantaneous audio-visual assignment problem. Moreover, the temporal extension of this initial graph achieves a new state-of-the-art on the AVA-ActiveSpeaker dataset with a mAP of 88.8\\%.","url_abs":"https://arxiv.org/abs/2101.03682v2","url_pdf":"https://arxiv.org/pdf/2101.03682v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"maas-multi-modal-assignation-for-active","repo_url":"https://github.com/fuankarion/maas","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"active-speaker-detection","task_name":"Active Speaker Detection"},{"task_slug":"audio-visual-active-speaker-detection","task_name":"Audio-Visual Active Speaker Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-visual-active-speaker-detection-on-ava","task":"Audio-Visual Active Speaker Detection","dataset":"AVA-ActiveSpeaker","model":"MAAS-TAN","rank_in_archive_order":16,"of":20,"metrics":{"validation mean average precision":"88.8%"},"uses_additional_data":false},{"leaderboard":"/sota/audio-visual-active-speaker-detection-on-ava","task":"Audio-Visual Active Speaker Detection","dataset":"AVA-ActiveSpeaker","model":"MAAS-LAN","rank_in_archive_order":19,"of":20,"metrics":{"validation mean average precision":"85.1%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2101.03682","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}