{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-from-video-and-text-via-large-scale","title":"Learning from Video and Text via Large-Scale Discriminative Clustering","arxiv_id":"1707.09074","date":"2017-07-27","proceeding":"ICCV 2017 10","authors":["Antoine Miech","Jean-Baptiste Alayrac","Piotr Bojanowski","Ivan Laptev","Josef Sivic"],"abstract":"Discriminative clustering has been successfully applied to a number of\nweakly-supervised learning tasks. Such applications include person and action\nrecognition, text-to-video alignment, object co-segmentation and colocalization\nin videos and images. One drawback of discriminative clustering, however, is\nits limited scalability. We address this issue and propose an online\noptimization algorithm based on the Block-Coordinate Frank-Wolfe algorithm. We\napply the proposed method to the problem of weakly supervised learning of\nactions and actors from movies together with corresponding movie scripts. The\nscaling up of the learning problem to 66 feature length movies enables us to\nsignificantly improve weakly supervised action recognition.","url_abs":"http://arxiv.org/abs/1707.09074v1","url_pdf":"http://arxiv.org/pdf/1707.09074v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-from-video-and-text-via-large-scale","repo_url":"https://github.com/jpeyre/unrel","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"learning-from-video-and-text-via-large-scale","repo_url":"https://github.com/antoine77340/iccv17learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-alignment","task_name":"Video Alignment"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"},{"task_slug":"weakly-supervised-action-recognition","task_name":"Weakly-Supervised Action Recognition"},{"task_slug":"weakly-supervised-learning","task_name":"Weakly-supervised Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-retrieval-on-lsmdc","task":"Video Retrieval","dataset":"LSMDC","model":"Large-Scale Discriminative Clustering","rank_in_archive_order":35,"of":38,"metrics":{"text-to-video Median Rank":"52","text-to-video R@1":"7.3","text-to-video R@10":"27.1","text-to-video R@5":"19.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.09074","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}