{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cycle-sum-cycle-consistent-adversarial-lstm","title":"Cycle-SUM: Cycle-consistent Adversarial LSTM Networks for Unsupervised Video Summarization","arxiv_id":"1904.08265","date":"2019-04-17","proceeding":null,"authors":["Li Yuan","Francis EH Tay","Ping Li","Li Zhou","Jiashi Feng"],"abstract":"In this paper, we present a novel unsupervised video summarization model that\nrequires no manual annotation. The proposed model termed Cycle-SUM adopts a new\ncycle-consistent adversarial LSTM architecture that can effectively maximize\nthe information preserving and compactness of the summary video. It consists of\na frame selector and a cycle-consistent learning based evaluator. The selector\nis a bi-direction LSTM network that learns video representations that embed the\nlong-range relationships among video frames. The evaluator defines a learnable\ninformation preserving metric between original video and summary video and\n\"supervises\" the selector to identify the most informative frames to form the\nsummary video. In particular, the evaluator is composed of two generative\nadversarial networks (GANs), in which the forward GAN is learned to reconstruct\noriginal video from summary video while the backward GAN learns to invert the\nprocessing. The consistency between the output of such cycle learning is\nadopted as the information preserving metric for video summarization. We\ndemonstrate the close relation between mutual information maximization and such\ncycle learning procedure. Experiments on two video summarization benchmark\ndatasets validate the state-of-the-art performance and superiority of the\nCycle-SUM model over previous baselines.","url_abs":"http://arxiv.org/abs/1904.08265v1","url_pdf":"http://arxiv.org/pdf/1904.08265v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"unsupervised-video-summarization","task_name":"Unsupervised Video Summarization"},{"task_slug":"video-summarization","task_name":"Video Summarization"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-video-summarization-on-summe","task":"Unsupervised Video Summarization","dataset":"SumMe","model":"Cycle-SUM","rank_in_archive_order":9,"of":10,"metrics":{"F1-score":"41.9"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-video-summarization-on-tvsum","task":"Unsupervised Video Summarization","dataset":"TvSum","model":"Cycle-SUM","rank_in_archive_order":8,"of":8,"metrics":{"F1-score":"57.6"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.08265","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}