{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/when-did-it-happen-duration-informed-temporal","title":"When Did It Happen? Duration-informed Temporal Localization of Narrated Actions in Vlogs","arxiv_id":"2202.08138","date":"2022-02-16","proceeding":null,"authors":["Oana Ignat","Santiago Castro","YuHang Zhou","Jiajun Bao","Dandan Shan","Rada Mihalcea"],"abstract":"We consider the task of temporal human action localization in lifestyle vlogs. We introduce a novel dataset consisting of manual annotations of temporal localization for 13,000 narrated actions in 1,200 video clips. We present an extensive analysis of this data, which allows us to better understand how the language and visual modalities interact throughout the videos. We propose a simple yet effective method to localize the narrated actions based on their expected duration. Through several experiments and analyses, we show that our method brings complementary information with respect to previous methods, and leads to improvements over previous work for the task of temporal action localization.","url_abs":"https://arxiv.org/abs/2202.08138v2","url_pdf":"https://arxiv.org/pdf/2202.08138v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"when-did-it-happen-duration-informed-temporal","repo_url":"https://github.com/michigannlp/vlog_action_localization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"temporal-localization","task_name":"Temporal Localization"}],"methods":[],"datasets_introduced":[{"slug":"whenact","name":"WhenAct","full_name":"Temporal Human Action Localization in Lifestyle Vlogs"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}