{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-re-localization","title":"Video Re-localization","arxiv_id":"1808.01575","date":"2018-08-05","proceeding":"ECCV 2018 9","authors":["Yang Feng","Lin Ma","Wei Liu","Tong Zhang","Jiebo Luo"],"abstract":"Many methods have been developed to help people find the video contents they\nwant efficiently. However, there are still some unsolved problems in this area.\nFor example, given a query video and a reference video, how to accurately\nlocalize a segment in the reference video such that the segment semantically\ncorresponds to the query video? We define a distinctively new task, namely\n\\textbf{video re-localization}, to address this scenario. Video re-localization\nis an important emerging technology implicating many applications, such as fast\nseeking in videos, video copy detection, video surveillance, etc. Meanwhile, it\nis also a challenging research task because the visual appearance of a semantic\nconcept in videos can have large variations. The first hurdle to clear for the\nvideo re-localization task is the lack of existing datasets. It is labor\nexpensive to collect pairs of videos with semantic coherence or correspondence\nand label the corresponding segments. We first exploit and reorganize the\nvideos in ActivityNet to form a new dataset for video re-localization research,\nwhich consists of about 10,000 videos of diverse visual appearances associated\nwith localized boundary information. Subsequently, we propose an innovative\ncross gated bilinear matching model such that every time-step in the reference\nvideo is matched against the attentively weighted query video. Consequently,\nthe prediction of the starting and ending time is formulated as a\nclassification problem based on the matching results. Extensive experimental\nresults show that the proposed method outperforms the competing methods. Our\ncode is available at: https://github.com/fengyang0317/video_reloc.","url_abs":"http://arxiv.org/abs/1808.01575v1","url_pdf":"http://arxiv.org/pdf/1808.01575v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-re-localization","repo_url":"https://github.com/fengyang0317/video_reloc","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"copy-detection","task_name":"Copy Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.01575","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}