{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-summarization-using-deep-semantic","title":"Video Summarization using Deep Semantic Features","arxiv_id":"1609.08758","date":"2016-09-28","proceeding":null,"authors":["Mayu Otani","Yuta Nakashima","Esa Rahtu","Janne Heikkilä","Naokazu Yokoya"],"abstract":"This paper presents a video summarization technique for an Internet video to\nprovide a quick way to overview its content. This is a challenging problem\nbecause finding important or informative parts of the original video requires\nto understand its content. Furthermore the content of Internet videos is very\ndiverse, ranging from home videos to documentaries, which makes video\nsummarization much more tough as prior knowledge is almost not available. To\ntackle this problem, we propose to use deep video features that can encode\nvarious levels of content semantics, including objects, actions, and scenes,\nimproving the efficiency of standard video summarization techniques. For this,\nwe design a deep neural network that maps videos as well as descriptions to a\ncommon semantic space and jointly trained it with associated pairs of videos\nand descriptions. To generate a video summary, we extract the deep features\nfrom each segment of the original video and apply a clustering-based\nsummarization technique to them. We evaluate our video summaries using the\nSumMe dataset as well as baseline approaches. The results demonstrated the\nadvantages of incorporating our deep semantic features in a video summarization\ntechnique.","url_abs":"http://arxiv.org/abs/1609.08758v1","url_pdf":"http://arxiv.org/pdf/1609.08758v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-summarization-using-deep-semantic","repo_url":"https://github.com/590shun/video_summarize_dsf","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"video-summarization-using-deep-semantic","repo_url":"https://github.com/590shun/vsum_dsf","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"video-summarization","task_name":"Video Summarization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}