{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-flexible-model-for-training-action","title":"A flexible model for training action localization with varying levels of supervision","arxiv_id":"1806.11328","date":"2018-06-29","proceeding":"NeurIPS 2018 12","authors":["Guilhem Chéron","Jean-Baptiste Alayrac","Ivan Laptev","Cordelia Schmid"],"abstract":"Spatio-temporal action detection in videos is typically addressed in a\nfully-supervised setup with manual annotation of training videos required at\nevery frame. Since such annotation is extremely tedious and prohibits\nscalability, there is a clear need to minimize the amount of manual\nsupervision. In this work we propose a unifying framework that can handle and\ncombine varying types of less-demanding weak supervision. Our model is based on\ndiscriminative clustering and integrates different types of supervision as\nconstraints on the optimization. We investigate applications of such a model to\ntraining setups with alternative supervisory signals ranging from video-level\nclass labels to the full per-frame annotation of action bounding boxes.\nExperiments on the challenging UCF101-24 and DALY datasets demonstrate\ncompetitive performance of our method at a fraction of supervision used by\nprevious methods. The flexibility of our model enables joint learning from data\nwith different levels of annotation. Experimental results demonstrate a\nsignificant gain by adding a few fully supervised examples to otherwise weakly\nlabeled videos.","url_abs":"http://arxiv.org/abs/1806.11328v2","url_pdf":"http://arxiv.org/pdf/1806.11328v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-flexible-model-for-training-action","repo_url":"https://github.com/jalayrac/weakactionloc","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.11328","atlas_url":"https://app.syntology.ai/?focus=1806.11328","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}