{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-recognition-from-single-timestamp","title":"Action Recognition from Single Timestamp Supervision in Untrimmed Videos","arxiv_id":"1904.04689","date":"2019-04-09","proceeding":"CVPR 2019 6","authors":["Davide Moltisanti","Sanja Fidler","Dima Damen"],"abstract":"Recognising actions in videos relies on labelled supervision during training,\ntypically the start and end times of each action instance. This supervision is\nnot only subjective, but also expensive to acquire. Weak video-level\nsupervision has been successfully exploited for recognition in untrimmed\nvideos, however it is challenged when the number of different actions in\ntraining videos increases. We propose a method that is supervised by single\ntimestamps located around each action instance, in untrimmed videos. We replace\nexpensive action bounds with sampling distributions initialised from these\ntimestamps. We then use the classifier's response to iteratively update the\nsampling distributions. We demonstrate that these distributions converge to the\nlocation and extent of discriminative action segments. We evaluate our method\non three datasets for fine-grained recognition, with increasing number of\ndifferent actions per video, and show that single timestamps offer a reasonable\ncompromise between recognition performance and labelling effort, performing\ncomparably to full temporal supervision. Our update method improves top-1 test\naccuracy by up to 5.4%. across the evaluated datasets.","url_abs":"http://arxiv.org/abs/1904.04689v1","url_pdf":"http://arxiv.org/pdf/1904.04689v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"action-recognition-from-single-timestamp","repo_url":"https://bitbucket.org/dmoltisanti/action_recognition_single_timestamps","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.04689","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}