{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-visual-concept-learning-with","title":"Multimodal Visual Concept Learning with Weakly Supervised Techniques","arxiv_id":"1712.00796","date":"2017-12-03","proceeding":"CVPR 2018 6","authors":["Giorgos Bouritsas","Petros Koutras","Athanasia Zlatintsi","Petros Maragos"],"abstract":"Despite the availability of a huge amount of video data accompanied by\ndescriptive texts, it is not always easy to exploit the information contained\nin natural language in order to automatically recognize video concepts. Towards\nthis goal, in this paper we use textual cues as means of supervision,\nintroducing two weakly supervised techniques that extend the Multiple Instance\nLearning (MIL) framework: the Fuzzy Sets Multiple Instance Learning (FSMIL) and\nthe Probabilistic Labels Multiple Instance Learning (PLMIL). The former encodes\nthe spatio-temporal imprecision of the linguistic descriptions with Fuzzy Sets,\nwhile the latter models different interpretations of each description's\nsemantics with Probabilistic Labels, both formulated through a convex\noptimization algorithm. In addition, we provide a novel technique to extract\nweak labels in the presence of complex semantics, that consists of semantic\nsimilarity computations. We evaluate our methods on two distinct problems,\nnamely face and action recognition, in the challenging and realistic setting of\nmovies accompanied by their screenplays, contained in the COGNIMUSE database.\nWe show that, on both tasks, our method considerably outperforms a\nstate-of-the-art weakly supervised approach, as well as other baselines.","url_abs":"http://arxiv.org/abs/1712.00796v3","url_pdf":"http://arxiv.org/pdf/1712.00796v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-visual-concept-learning-with","repo_url":"https://github.com/gbouritsas/cvpr18_multimodal_weakly_supervised_learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"descriptive","task_name":"Descriptive"},{"task_slug":"multiple-instance-learning","task_name":"Multiple Instance Learning"},{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}