{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-visual-attention-prediction","title":"Deep Visual Attention Prediction","arxiv_id":"1705.02544","date":"2017-05-07","proceeding":"journal 2017 11","authors":["Wenguan Wang","Jianbing Shen"],"abstract":"In this work, we aim to predict human eye fixation with view-free scenes\nbased on an end-to-end deep learning architecture. Although Convolutional\nNeural Networks (CNNs) have made substantial improvement on human attention\nprediction, it is still needed to improve CNN based attention models by\nefficiently leveraging multi-scale features. Our visual attention network is\nproposed to capture hierarchical saliency information from deep, coarse layers\nwith global saliency information to shallow, fine layers with local saliency\nresponse. Our model is based on a skip-layer network structure, which predicts\nhuman attention from multiple convolutional layers with various reception\nfields. Final saliency prediction is achieved via the cooperation of those\nglobal and local predictions. Our model is learned in a deep supervision\nmanner, where supervision is directly fed into multi-level layers, instead of\nprevious approaches of providing supervision only at the output layer and\npropagating this supervision back to earlier layers. Our model thus\nincorporates multi-level saliency predictions within a single network, which\nsignificantly decreases the redundancy of previous approaches of learning\nmultiple network streams with different input scales. Extensive experimental\nanalysis on various challenging benchmark datasets demonstrate our method\nyields state-of-the-art performance with competitive inference time.","url_abs":"http://arxiv.org/abs/1705.02544v3","url_pdf":"http://arxiv.org/pdf/1705.02544v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-visual-attention-prediction","repo_url":"https://github.com/wenguanwang/deepattention","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"saliency-prediction","task_name":"Saliency Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1705.02544","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}