{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/query-focused-video-summarization-dataset","title":"Query-Focused Video Summarization: Dataset, Evaluation, and A Memory Network Based Approach","arxiv_id":"1707.04960","date":"2017-07-16","proceeding":"CVPR 2017 7","authors":["Aidean Sharghi","Jacob S. Laurel","Boqing Gong"],"abstract":"Recent years have witnessed a resurgence of interest in video summarization.\nHowever, one of the main obstacles to the research on video summarization is\nthe user subjectivity - users have various preferences over the summaries. The\nsubjectiveness causes at least two problems. First, no single video summarizer\nfits all users unless it interacts with and adapts to the individual users.\nSecond, it is very challenging to evaluate the performance of a video\nsummarizer.\n  To tackle the first problem, we explore the recently proposed query-focused\nvideo summarization which introduces user preferences in the form of text\nqueries about the video into the summarization process. We propose a memory\nnetwork parameterized sequential determinantal point process in order to attend\nthe user query onto different video frames and shots. To address the second\nchallenge, we contend that a good evaluation metric for video summarization\nshould focus on the semantic information that humans can perceive rather than\nthe visual features or temporal overlaps. To this end, we collect dense\nper-video-shot concept annotations, compile a new dataset, and suggest an\nefficient evaluation method defined upon the concept annotations. We conduct\nextensive experiments contrasting our video summarizer to existing ones and\npresent detailed analyses about the dataset and the new evaluation method.","url_abs":"http://arxiv.org/abs/1707.04960v1","url_pdf":"http://arxiv.org/pdf/1707.04960v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"query-focused-video-summarization","task_name":"Query focused video summarization"},{"task_slug":"video-summarization","task_name":"Video Summarization"}],"methods":[],"datasets_introduced":[{"slug":"query-focused-video-summarization-dataset","name":"Query-Focused Video Summarization Dataset","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.04960","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}