{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/uc-net-uncertainty-inspired-rgb-d-saliency","title":"UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders","arxiv_id":"2004.05763","date":"2020-04-13","proceeding":"CVPR 2020 6","authors":["Jing Zhang","Deng-Ping Fan","Yuchao Dai","Saeed Anwar","Fatemeh Sadat Saleh","Tong Zhang","Nick Barnes"],"abstract":"In this paper, we propose the first framework (UCNet) to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection methods treat the saliency detection task as a point estimation problem, and produce a single saliency map following a deterministic learning pipeline. Inspired by the saliency data labeling process, we propose probabilistic RGB-D saliency detection network via conditional variational autoencoders to model human annotation uncertainty and generate multiple saliency maps for each input image by sampling in the latent space. With the proposed saliency consensus process, we are able to generate an accurate saliency map based on these multiple predictions. Quantitative and qualitative evaluations on six challenging benchmark datasets against 18 competing algorithms demonstrate the effectiveness of our approach in learning the distribution of saliency maps, leading to a new state-of-the-art in RGB-D saliency detection.","url_abs":"https://arxiv.org/abs/2004.05763v1","url_pdf":"https://arxiv.org/pdf/2004.05763v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"uc-net-uncertainty-inspired-rgb-d-saliency","repo_url":"https://github.com/JingZhang617/UCNet","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"rgb-d-salient-object-detection","task_name":"RGB-D Salient Object Detection"},{"task_slug":"saliency-detection","task_name":"Saliency Detection"},{"task_slug":"thermal-image-segmentation","task_name":"Thermal Image Segmentation"}],"methods":[{"method_slug":"ucnet","method_name":"UCNet"}],"datasets_introduced":[],"methods_introduced":[{"slug":"ucnet","name":"UCNet","full_name":"UCNet"}],"results":[{"leaderboard":"/sota/rgb-d-salient-object-detection-on-des","task":"RGB-D Salient Object Detection","dataset":"DES","model":"UC-Net","rank_in_archive_order":8,"of":13,"metrics":{"Average MAE":"0.019","S-Measure":"93.4"},"uses_additional_data":false},{"leaderboard":"/sota/rgb-d-salient-object-detection-on-lfsd","task":"RGB-D Salient Object Detection","dataset":"LFSD","model":"UC-Net","rank_in_archive_order":4,"of":8,"metrics":{"Average MAE":"0.066","S-Measure":"86.4"},"uses_additional_data":false},{"leaderboard":"/sota/rgb-d-salient-object-detection-on-nju2k","task":"RGB-D Salient Object Detection","dataset":"NJU2K","model":"UC-Net","rank_in_archive_order":17,"of":27,"metrics":{"Average MAE":"0.043","S-Measure":"89.7"},"uses_additional_data":false},{"leaderboard":"/sota/rgb-d-salient-object-detection-on-nlpr","task":"RGB-D Salient Object Detection","dataset":"NLPR","model":"UC-Net","rank_in_archive_order":9,"of":14,"metrics":{"Average MAE":"0.025","S-Measure":"92.0"},"uses_additional_data":false},{"leaderboard":"/sota/rgb-d-salient-object-detection-on-sip","task":"RGB-D Salient Object Detection","dataset":"SIP","model":"UC-Net","rank_in_archive_order":12,"of":16,"metrics":{"Average MAE":"0.051","S-Measure":"87.5"},"uses_additional_data":false},{"leaderboard":"/sota/rgb-d-salient-object-detection-on-stere","task":"RGB-D Salient Object Detection","dataset":"STERE","model":"UC-Net","rank_in_archive_order":11,"of":14,"metrics":{"Average MAE":"0.039","S-Measure":"90.3"},"uses_additional_data":false},{"leaderboard":"/sota/thermal-image-segmentation-on-rgb-t-glass","task":"Thermal Image Segmentation","dataset":"RGB-T-Glass-Segmentation","model":"UCNet","rank_in_archive_order":16,"of":22,"metrics":{"MAE":"0.071"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2004.05763","atlas_url":"https://app.syntology.ai/?focus=2004.05763","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}