{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/audio-visual-class-incremental-learning","title":"Audio-Visual Class-Incremental Learning","arxiv_id":"2308.11073","date":"2023-08-21","proceeding":"ICCV 2023 1","authors":["Weiguo Pian","Shentong Mo","Yunhui Guo","Yapeng Tian"],"abstract":"In this paper, we introduce audio-visual class-incremental learning, a class-incremental learning scenario for audio-visual video recognition. We demonstrate that joint audio-visual modeling can improve class-incremental learning, but current methods fail to preserve semantic similarity between audio and visual features as incremental step grows. Furthermore, we observe that audio-visual correlations learned in previous tasks can be forgotten as incremental steps progress, leading to poor performance. To overcome these challenges, we propose AV-CIL, which incorporates Dual-Audio-Visual Similarity Constraint (D-AVSC) to maintain both instance-aware and class-aware semantic similarity between audio-visual modalities and Visual Attention Distillation (VAD) to retain previously learned audio-guided visual attentive ability. We create three audio-visual class-incremental datasets, AVE-Class-Incremental (AVE-CI), Kinetics-Sounds-Class-Incremental (K-S-CI), and VGGSound100-Class-Incremental (VS100-CI) based on the AVE, Kinetics-Sounds, and VGGSound datasets, respectively. Our experiments on AVE-CI, K-S-CI, and VS100-CI demonstrate that AV-CIL significantly outperforms existing class-incremental learning methods in audio-visual class-incremental learning. Code and data are available at: https://github.com/weiguoPian/AV-CIL_ICCV2023.","url_abs":"https://arxiv.org/abs/2308.11073v3","url_pdf":"https://arxiv.org/pdf/2308.11073v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"audio-visual-class-incremental-learning","repo_url":"https://github.com/weiguopian/av-cil_iccv2023","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"class-incremental-learning","task_name":"Class Incremental Learning"},{"task_slug":"incremental-learning","task_name":"Incremental Learning"},{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"video-recognition","task_name":"Video Recognition"},{"task_slug":"class-incremental-learning-1","task_name":"class-incremental learning"}],"methods":[{"method_slug":"fail","method_name":"fail"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2308.11073","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2308.11073"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/weiguopian/av-cil_iccv2023","reach":null}],"summary":{"ran_draft_wrong":4,"ran":3,"unverified":3},"by_repo_kind":{"official":{"samples":10,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":10,"samples":[{"code_sha256_prefix":"f9ae6e33c3aadd62","entry":"CE_loss","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f9ae6e33c3aadd62"}},{"code_sha256_prefix":"b704e32d33e265ca","entry":"LSCLinear","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b704e32d33e265ca"}},{"code_sha256_prefix":"ecf63b2240c893c5","entry":"cal_contrastive_loss","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ecf63b2240c893c5"}},{"code_sha256_prefix":"ad9bb1a102549e3b","entry":"class_contrastive_loss","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ad9bb1a102549e3b"}},{"code_sha256_prefix":"b8de6a8bb6a96884","entry":"reduce_proxies","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b8de6a8bb6a96884"}},{"code_sha256_prefix":"1f5d558e108778a2","entry":"stable_cosine_distance","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1f5d558e108778a2"}},{"code_sha256_prefix":"1421246901b640f4","entry":"top_1_acc","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1421246901b640f4"}},{"code_sha256_prefix":"98d8d0bd73341c28","entry":"IncreAudioVisualNet","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"98d8d0bd73341c28"}},{"code_sha256_prefix":"a47019e4e5c5ff88","entry":"adjust_learning_rate","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a47019e4e5c5ff88"}},{"code_sha256_prefix":"211bcfcce2e3dfcc","entry":"train","repo":"weiguopian/av-cil_iccv2023","repo_kind":"official","path":"ours/train_incremental_ours.py","file_url":"https://github.com/weiguopian/av-cil_iccv2023/blob/HEAD/ours/train_incremental_ours.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"211bcfcce2e3dfcc"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}