{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gpt2mvs-generative-pre-trained-transformer-2","title":"GPT2MVS: Generative Pre-trained Transformer-2 for Multi-modal Video Summarization","arxiv_id":"2104.12465","date":"2021-04-26","proceeding":null,"authors":["Jia-Hong Huang","Luka Murn","Marta Mrak","Marcel Worring"],"abstract":"Traditional video summarization methods generate fixed video representations regardless of user interest. Therefore such methods limit users' expectations in content search and exploration scenarios. Multi-modal video summarization is one of the methods utilized to address this problem. When multi-modal video summarization is used to help video exploration, a text-based query is considered as one of the main drivers of video summary generation, as it is user-defined. Thus, encoding the text-based query and the video effectively are both important for the task of multi-modal video summarization. In this work, a new method is proposed that uses a specialized attention network and contextualized word representations to tackle this task. The proposed model consists of a contextualized video summary controller, multi-modal attention mechanisms, an interactive attention network, and a video summary generator. Based on the evaluation of the existing multi-modal video summarization benchmark, experimental results show that the proposed model is effective with the increase of +5.88% in accuracy and +4.06% increase of F1-score, compared with the state-of-the-art method.","url_abs":"https://arxiv.org/abs/2104.12465v1","url_pdf":"https://arxiv.org/pdf/2104.12465v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gpt2mvs-generative-pre-trained-transformer-2","repo_url":"https://github.com/2023-MindSpore-1/ms-code-5/tree/main/gpt2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gpt2mvs-generative-pre-trained-transformer-2","repo_url":"https://github.com/2023-MindSpore-4/Code12/tree/main/MindFormers/gpt2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gpt2mvs-generative-pre-trained-transformer-2","repo_url":"https://github.com/pwc-1/Paper-5/tree/main/megatron_gpt2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"gpt2mvs-generative-pre-trained-transformer-2","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/2/megatron_gpt2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"video-summarization","task_name":"Video Summarization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}