{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rethinking-transfer-and-auxiliary-learning","title":"Rethinking Transfer and Auxiliary Learning for Improving Audio Captioning Transformer","arxiv_id":null,"date":"2023-08-20","proceeding":"Interspeech 2023 8","authors":["WooSeok Shin","Hyun Joon Park","Jin Sob Kim","Dongwon Kim","Seungjin Lee","Sung Won Han"],"abstract":"The performance of automated audio captioning (AAC) has been improved considerably through a transformer-based encoder and transfer learning. However, their performance improvement is constrained by the following problems: (1) discrepancy in the input patch size between pretraining and fine-tuning steps. (2) lack of local-level relations between inputs and captions. In this paper, we propose a simple transfer learning scheme that maintains input patch sizes, unlike previous methods, to avoid input discrepancies. Furthermore, we propose a patch-wise keyword estimation branch that utilizes an attention pooling method to effectively represent both global- and local-level information. The results on the AudioCaps dataset reveal that the proposed learning scheme and method considerably contribute to performance gain. Finally, the visualization results demonstrate that the proposed attention-pooling method effectively detects local-level information in the AAC system.","url_abs":"https://www.isca-speech.org/archive/interspeech_2023/shin23_interspeech.html","url_pdf":"https://www.isca-speech.org/archive/pdfs/interspeech_2023/shin23_interspeech.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"audio-captioning","task_name":"Audio captioning"},{"task_slug":null,"task_name":"AudioCaps"},{"task_slug":"auxiliary-learning","task_name":"Auxiliary Learning"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[{"method_slug":"attention-pooling","method_name":"Attention Pooling"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-captioning-on-audiocaps","task":"Audio captioning","dataset":"AudioCaps","model":"Rethink-ACT (AST + TF + MIL)","rank_in_archive_order":12,"of":18,"metrics":{"BLEU-4":"0.285","CIDEr":"0.764","METEOR":"0.242","ROUGE-L":"0.504","SPICE":"0.180","SPIDEr":"0.472"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}