{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-video-summarization-via","title":"Unsupervised Video Summarization via Attention-Driven Adversarial Learning","arxiv_id":null,"date":"2019-12-24","proceeding":"MultiMedia Modeling (MMM) 2019 12","authors":["Evlampios Apostolidis","Eleni Adamantidou","Alexandros I. Metsai","Vasileios Mezaris","Ioannis Patras"],"abstract":"This paper presents a new video summarization approach that integrates an attention mechanism to identify the significant parts of the video, and is trained unsupervisingly via generative adversarial learning. Starting from the SUM-GAN model, we first develop an improved version of it (called SUM-GAN-sl) that has a significantly reduced number of learned parameters, performs incremental training of the model’s components, and applies a stepwise label-based strategy for updating the adversarial part. Subsequently, we introduce an attention mechanism to SUM-GAN-sl in two ways: (i) by integrating an attention layer within the variational auto-encoder (VAE) of the architecture (SUM-GAN-VAAE), and (ii) by replacing the VAE with a deterministic attention auto-encoder (SUM-GAN-AAE). Experimental evaluation on two datasets (SumMe and TVSum) documents the contribution of the attention auto-encoder to faster and more stable training of the model, resulting in a significant performance improvement with respect to the original model and demonstrating the competitiveness of the proposed SUM-GAN-AAE against the state of the art.","url_abs":"https://link.springer.com/chapter/10.1007/978-3-030-37731-1_40","url_pdf":"https://link.springer.com/chapter/10.1007/978-3-030-37731-1_40","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-video-summarization-via","repo_url":"https://github.com/e-apostolidis/SUM-GAN-AAE","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"unsupervised-video-summarization","task_name":"Unsupervised Video Summarization"},{"task_slug":"video-summarization","task_name":"Video Summarization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-video-summarization-on-summe","task":"Unsupervised Video Summarization","dataset":"SumMe","model":"SUM-GAN-AAE","rank_in_archive_order":7,"of":10,"metrics":{"F1-score":"48.9","Parameters (M)":"24.31","training time (s)":"1639"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-video-summarization-on-tvsum","task":"Unsupervised Video Summarization","dataset":"TvSum","model":"SUM-GAN-AAE","rank_in_archive_order":6,"of":8,"metrics":{"F1-score":"58.3","Parameters (M)":"24.31","training time (s)":"5423"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}