{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/localizing-moments-in-video-with-temporal","title":"Localizing Moments in Video with Temporal Language","arxiv_id":"1809.01337","date":"2018-09-05","proceeding":"EMNLP 2018 10","authors":["Lisa Anne Hendricks","Oliver Wang","Eli Shechtman","Josef Sivic","Trevor Darrell","Bryan Russell"],"abstract":"Localizing moments in a longer video via natural language queries is a new,\nchallenging task at the intersection of language and video understanding.\nThough moment localization with natural language is similar to other language\nand vision tasks like natural language object retrieval in images, moment\nlocalization offers an interesting opportunity to model temporal dependencies\nand reasoning in text. We propose a new model that explicitly reasons about\ndifferent temporal segments in a video, and shows that temporal context is\nimportant for localizing phrases which include temporal language. To benchmark\nwhether our model, and other recent video localization models, can effectively\nreason about temporal language, we collect the novel TEMPOral reasoning in\nvideo and language (TEMPO) dataset. Our dataset consists of two parts: a\ndataset with real videos and template sentences (TEMPO - Template Language)\nwhich allows for controlled studies on temporal language, and a human language\ndataset which consists of temporal sentences annotated by humans (TEMPO - Human\nLanguage).","url_abs":"http://arxiv.org/abs/1809.01337v1","url_pdf":"http://arxiv.org/pdf/1809.01337v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"localizing-moments-in-video-with-temporal","repo_url":"https://github.com/LisaAnne/TemporalLanguageRelease","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"caffe2","reach":{"status":"ok"}}],"tasks":[{"task_slug":"natural-language-queries","task_name":"Natural Language Queries"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[{"slug":"tempo","name":"TEMPO","full_name":"Localizing Moments in Video with Temporal Language"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.01337","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}