{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/watch-listen-and-describe-globally-and","title":"Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning","arxiv_id":"1804.05448","date":"2018-04-15","proceeding":"NAACL 2018 6","authors":["Xin Wang","Yuan-Fang Wang","William Yang Wang"],"abstract":"A major challenge for video captioning is to combine audio and visual cues.\nExisting multi-modal fusion methods have shown encouraging results in video\nunderstanding. However, the temporal structures of multiple modalities at\ndifferent granularities are rarely explored, and how to selectively fuse the\nmulti-modal representations at different levels of details remains uncharted.\nIn this paper, we propose a novel hierarchically aligned cross-modal attention\n(HACA) framework to learn and selectively fuse both global and local temporal\ndynamics of different modalities. Furthermore, for the first time, we validate\nthe superior performance of the deep audio features on the video captioning\ntask. Finally, our HACA model significantly outperforms the previous best\nsystems and achieves new state-of-the-art results on the widely used MSR-VTT\ndataset.","url_abs":"http://arxiv.org/abs/1804.05448v1","url_pdf":"http://arxiv.org/pdf/1804.05448v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"watch-listen-and-describe-globally-and","repo_url":"https://github.com/chitwansaharia/HACAModel","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.05448","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}