{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/audio-visual-event-recognition-through-the","title":"Audio-Visual Event Recognition through the lens of Adversary","arxiv_id":"2011.07430","date":"2020-11-15","proceeding":null,"authors":["Juncheng B Li","Kaixin Ma","Shuhui Qu","Po-Yao Huang","Florian Metze"],"abstract":"As audio/visual classification models are widely deployed for sensitive tasks like content filtering at scale, it is critical to understand their robustness along with improving the accuracy. This work aims to study several key questions related to multimodal learning through the lens of adversarial noises: 1) The trade-off between early/middle/late fusion affecting its robustness and accuracy 2) How do different frequency/time domain features contribute to the robustness? 3) How do different neural modules contribute to the adversarial noise? In our experiment, we construct adversarial examples to attack state-of-the-art neural models trained on Google AudioSet. We compare how much attack potency in terms of adversarial perturbation of size $\\epsilon$ using different $L_p$ norms we would need to \"deactivate\" the victim model. Using adversarial noise to ablate multimodal models, we are able to provide insights into what is the best potential fusion strategy to balance the model parameters/accuracy and robustness trade-off and distinguish the robust features versus the non-robust features that various neural networks model tend to learn.","url_abs":"https://arxiv.org/abs/2011.07430v1","url_pdf":"https://arxiv.org/pdf/2011.07430v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"audio-visual-event-recognition-through-the","repo_url":"https://github.com/lijuncheng16/AudioSetDoneRight","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2011.07430","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2011.07430"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lijuncheng16/AudioSetDoneRight","reach":null}],"summary":{"ran_draft_wrong":1,"ran_violates":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"a454df1ab6677cd8","entry":"get_linear_schedule_with_warmup","repo":"lijuncheng16/AudioSetDoneRight","repo_kind":"official","path":"train_multimodal_late_fusion.py","file_url":"https://github.com/lijuncheng16/AudioSetDoneRight/blob/HEAD/train_multimodal_late_fusion.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a454df1ab6677cd8"}},{"code_sha256_prefix":"1d124cd8be27a2fb","entry":"mybool","repo":"lijuncheng16/AudioSetDoneRight","repo_kind":"official","path":"train_multimodal_late_fusion.py","file_url":"https://github.com/lijuncheng16/AudioSetDoneRight/blob/HEAD/train_multimodal_late_fusion.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1d124cd8be27a2fb"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}