{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-source-weak-supervision-for-saliency","title":"Multi-source weak supervision for saliency detection","arxiv_id":"1904.00566","date":"2019-04-01","proceeding":"CVPR 2019 6","authors":["Yu Zeng","Yunzhi Zhuge","Huchuan Lu","Lihe Zhang","Mingyang Qian","Yizhou Yu"],"abstract":"The high cost of pixel-level annotations makes it appealing to train saliency\ndetection models with weak supervision. However, a single weak supervision\nsource usually does not contain enough information to train a well-performing\nmodel. To this end, we propose a unified framework to train saliency detection\nmodels with diverse weak supervision sources. In this paper, we use category\nlabels, captions, and unlabelled data for training, yet other supervision\nsources can also be plugged into this flexible framework. We design a\nclassification network (CNet) and a caption generation network (PNet), which\nlearn to predict object categories and generate captions, respectively,\nmeanwhile highlight the most important regions for corresponding tasks. An\nattention transfer loss is designed to transmit supervision signal between\nnetworks, such that the network designed to be trained with one supervision\nsource can benefit from another. An attention coherence loss is defined on\nunlabelled data to encourage the networks to detect generally salient regions\ninstead of task-specific regions. We use CNet and PNet to generate pixel-level\npseudo labels to train a saliency prediction network (SNet). During the testing\nphases, we only need SNet to predict saliency maps. Experiments demonstrate the\nperformance of our method compares favourably against unsupervised and weakly\nsupervised methods and even some supervised methods.","url_abs":"http://arxiv.org/abs/1904.00566v1","url_pdf":"http://arxiv.org/pdf/1904.00566v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-source-weak-supervision-for-saliency","repo_url":"https://github.com/zengxianyu/mws","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"caption-generation","task_name":"Caption Generation"},{"task_slug":"saliency-detection","task_name":"Saliency Detection"},{"task_slug":"saliency-prediction","task_name":"Saliency Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.00566","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}