{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-task-weakly-supervised-learning-from","title":"Cross-task weakly supervised learning from instructional videos","arxiv_id":"1903.08225","date":"2019-03-19","proceeding":"CVPR 2019 6","authors":["Dimitri Zhukov","Jean-Baptiste Alayrac","Ramazan Gokberk Cinbis","David Fouhey","Ivan Laptev","Josef Sivic"],"abstract":"In this paper we investigate learning visual models for the steps of ordinary\ntasks using weak supervision via instructional narrations and an ordered list\nof steps instead of strong supervision via temporal annotations. At the heart\nof our approach is the observation that weakly supervised learning may be\neasier if a model shares components while learning different steps: `pour egg'\nshould be trained jointly with other tasks involving `pour' and `egg'. We\nformalize this in a component model for recognizing steps and a weakly\nsupervised learning framework that can learn this model under temporal\nconstraints from narration and the list of steps. Past data does not permit\nsystematic studying of sharing and so we also gather a new dataset, CrossTask,\naimed at assessing cross-task sharing. Our experiments demonstrate that sharing\nacross tasks improves performance, especially when done at the component level\nand that our component model can parse previously unseen tasks by virtue of its\ncompositionality.","url_abs":"http://arxiv.org/abs/1903.08225v2","url_pdf":"http://arxiv.org/pdf/1903.08225v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cross-task-weakly-supervised-learning-from","repo_url":"https://github.com/DmZhukov/CrossTask","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"cross-task-weakly-supervised-learning-from","repo_url":"https://github.com/dpfried/action-segmentation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"weakly-supervised-learning","task_name":"Weakly-supervised Learning"}],"methods":[],"datasets_introduced":[{"slug":"crosstask","name":"CrossTask","full_name":"CrossTask"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/temporal-action-localization-on-crosstask","task":"Temporal Action Localization","dataset":"CrossTask","model":"Fully-supervised upper-bound","rank_in_archive_order":5,"of":7,"metrics":{"Recall":"31.6"},"uses_additional_data":false},{"leaderboard":"/sota/temporal-action-localization-on-crosstask","task":"Temporal Action Localization","dataset":"CrossTask","model":"Zhukov","rank_in_archive_order":6,"of":7,"metrics":{"Recall":"22.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.08225","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}