{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploiting-multi-modal-curriculum-in-noisy","title":"Exploiting Multi-modal Curriculum in Noisy Web Data for Large-scale Concept Learning","arxiv_id":"1607.04780","date":"2016-07-16","proceeding":null,"authors":["Junwei Liang","Lu Jiang","Deyu Meng","Alexander Hauptmann"],"abstract":"Learning video concept detectors automatically from the big but noisy web\ndata with no additional manual annotations is a novel but challenging area in\nthe multimedia and the machine learning community. A considerable amount of\nvideos on the web are associated with rich but noisy contextual information,\nsuch as the title, which provides weak annotations or labels about the video\ncontent. To leverage the big noisy web labels, this paper proposes a novel\nmethod called WEbly-Labeled Learning (WELL), which is established on the\nstate-of-the-art machine learning algorithm inspired by the learning process of\nhuman. WELL introduces a number of novel multi-modal approaches to incorporate\nmeaningful prior knowledge called curriculum from the noisy web videos. To\ninvestigate this problem, we empirically study the curriculum constructed from\nthe multi-modal features of the videos collected from YouTube and Flickr. The\nefficacy and the scalability of WELL have been extensively demonstrated on two\npublic benchmarks, including the largest multimedia dataset and the largest\nmanually-labeled video set. The comprehensive experimental results demonstrate\nthat WELL outperforms state-of-the-art studies by a statically significant\nmargin on learning concepts from noisy web video data. In addition, the results\nalso verify that WELL is robust to the level of noisiness in the video data.\nNotably, WELL trained on sufficient noisy web labels is able to achieve a\ncomparable accuracy to supervised learning methods trained on the clean\nmanually-labeled data.","url_abs":"http://arxiv.org/abs/1607.04780v1","url_pdf":"http://arxiv.org/pdf/1607.04780v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exploiting-multi-modal-curriculum-in-noisy","repo_url":"https://github.com/JunweiLiang/Semantic_Features","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}