{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporal-localization-of-fine-grained-actions","title":"Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images","arxiv_id":"1504.00983","date":"2015-04-04","proceeding":null,"authors":["Chen Sun","Sanketh Shetty","Rahul Sukthankar","Ram Nevatia"],"abstract":"We address the problem of fine-grained action localization from temporally\nuntrimmed web videos. We assume that only weak video-level annotations are\navailable for training. The goal is to use these weak labels to identify\ntemporal segments corresponding to the actions, and learn models that\ngeneralize to unconstrained web videos. We find that web images queried by\naction names serve as well-localized highlights for many actions, but are\nnoisily labeled. To solve this problem, we propose a simple yet effective\nmethod that takes weak video labels and noisy image labels as input, and\ngenerates localized action frames as output. This is achieved by cross-domain\ntransfer between video frames and web images, using pre-trained deep\nconvolutional neural networks. We then use the localized action frames to train\naction recognition models with long short-term memory networks. We collect a\nfine-grained sports action data set FGA-240 of more than 130,000 YouTube\nvideos. It has 240 fine-grained actions under 85 sports activities. Convincing\nresults are shown on the FGA-240 data set, as well as the THUMOS 2014\nlocalization data set with untrimmed training videos.","url_abs":"http://arxiv.org/abs/1504.00983v2","url_pdf":"http://arxiv.org/pdf/1504.00983v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"temporal-localization-of-fine-grained-actions","repo_url":"https://github.com/zhengshou/AutoLoc","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"temporal-localization","task_name":"Temporal Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1504.00983","atlas_url":"https://app.syntology.ai/?focus=1504.00983","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}