{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-active-speaker-detection","title":"End-to-End Active Speaker Detection","arxiv_id":"2203.14250","date":"2022-03-27","proceeding":null,"authors":["Juan Leon Alcazar","Moritz Cordes","Chen Zhao","Bernard Ghanem"],"abstract":"Recent advances in the Active Speaker Detection (ASD) problem build upon a two-stage process: feature extraction and spatio-temporal context aggregation. In this paper, we propose an end-to-end ASD workflow where feature learning and contextual predictions are jointly learned. Our end-to-end trainable network simultaneously learns multi-modal embeddings and aggregates spatio-temporal context. This results in more suitable feature representations and improved performance in the ASD task. We also introduce interleaved graph neural network (iGNN) blocks, which split the message passing according to the main sources of context in the ASD problem. Experiments show that the aggregated features from the iGNN blocks are more suitable for ASD, resulting in state-of-the art performance. Finally, we design a weakly-supervised strategy, which demonstrates that the ASD problem can also be approached by utilizing audiovisual data but relying exclusively on audio annotations. We achieve this by modelling the direct relationship between the audio signal and the possible sound sources (speakers), as well as introducing a contrastive loss. All the resources of this project will be made available at: https://github.com/fuankarion/end-to-end-asd.","url_abs":"https://arxiv.org/abs/2203.14250v2","url_pdf":"https://arxiv.org/pdf/2203.14250v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-active-speaker-detection","repo_url":"https://github.com/fuankarion/end-to-end-asd","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"end-to-end-active-speaker-detection","repo_url":"https://github.com/tiago-roxo/asdnb","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"end-to-end-active-speaker-detection","repo_url":"https://github.com/tiago-roxo/bias","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"active-speaker-detection","task_name":"Active Speaker Detection"},{"task_slug":"audio-visual-active-speaker-detection","task_name":"Audio-Visual Active Speaker Detection"},{"task_slug":"graph-neural-network","task_name":"Graph Neural Network"}],"methods":[{"method_slug":"graph-neural-network","method_name":"Graph Neural Network"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-visual-active-speaker-detection-on-ava","task":"Audio-Visual Active Speaker Detection","dataset":"AVA-ActiveSpeaker","model":"EASEE-50","rank_in_archive_order":7,"of":20,"metrics":{"validation mean average precision":"94.1%"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2203.14250","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2203.14250"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/fuankarion/end-to-end-asd","reach":null}],"summary":{"ran":1,"ran_draft_wrong":2,"unverified":2},"by_repo_kind":{"official":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"e50cfa477d25dece","entry":"LinearPathPreact","repo":"fuankarion/end-to-end-asd","repo_kind":"official","path":"models/graph_models.py","file_url":"https://github.com/fuankarion/end-to-end-asd/blob/HEAD/models/graph_models.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e50cfa477d25dece"}},{"code_sha256_prefix":"213ccc8fb87ee8d9","entry":"conv1x1x1","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"213ccc8fb87ee8d9"}},{"code_sha256_prefix":"93b21cfea9c9966c","entry":"generate_av_mask","repo":"fuankarion/end-to-end-asd","repo_kind":"official","path":"models/graph_models.py","file_url":"https://github.com/fuankarion/end-to-end-asd/blob/HEAD/models/graph_models.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"93b21cfea9c9966c"}},{"code_sha256_prefix":"42f183f82d339d43","entry":"GraphTwoStreamResNet3D","repo":"fuankarion/end-to-end-asd","repo_kind":"official","path":"models/graph_models.py","file_url":"https://github.com/fuankarion/end-to-end-asd/blob/HEAD/models/graph_models.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"42f183f82d339d43"}},{"code_sha256_prefix":"90adc8d3d7ed1195","entry":"GraphTwoStreamResNet3DTwoGraphs4LVLRes","repo":"fuankarion/end-to-end-asd","repo_kind":"official","path":"models/graph_models.py","file_url":"https://github.com/fuankarion/end-to-end-asd/blob/HEAD/models/graph_models.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"90adc8d3d7ed1195"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}