{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pixel-wise-attentional-gating-for","title":"Pixel-wise Attentional Gating for Parsimonious Pixel Labeling","arxiv_id":"1805.01556","date":"2018-05-03","proceeding":null,"authors":["Shu Kong","Charless Fowlkes"],"abstract":"To achieve parsimonious inference in per-pixel labeling tasks with a limited\ncomputational budget, we propose a \\emph{Pixel-wise Attentional Gating} unit\n(\\emph{PAG}) that learns to selectively process a subset of spatial locations\nat each layer of a deep convolutional network. PAG is a generic,\narchitecture-independent, problem-agnostic mechanism that can be readily\n\"plugged in\" to an existing model with fine-tuning. We utilize PAG in two ways:\n1) learning spatially varying pooling fields that improve model performance\nwithout the extra computation cost associated with multi-scale pooling, and 2)\nlearning a dynamic computation policy for each pixel to decrease total\ncomputation while maintaining accuracy.\n  We extensively evaluate PAG on a variety of per-pixel labeling tasks,\nincluding semantic segmentation, boundary detection, monocular depth and\nsurface normal estimation. We demonstrate that PAG allows competitive or\nstate-of-the-art performance on these tasks. Our experiments show that PAG\nlearns dynamic spatial allocation of computation over the input image which\nprovides better performance trade-offs compared to related approaches (e.g.,\ntruncating deep models or dynamically skipping whole layers). Generally, we\nobserve PAG can reduce computation by $10\\%$ without noticeable loss in\naccuracy and performance degrades gracefully when imposing stronger\ncomputational constraints.","url_abs":"http://arxiv.org/abs/1805.01556v2","url_pdf":"http://arxiv.org/pdf/1805.01556v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pixel-wise-attentional-gating-for","repo_url":"https://github.com/aimerykong/Pixel-Attentional-Gating","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"boundary-detection","task_name":"Boundary Detection"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"surface-normal-estimation","task_name":"Surface Normal Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-kitti-semantic","task":"Semantic Segmentation","dataset":"KITTI Semantic Segmentation","model":"APMoE_seg","rank_in_archive_order":7,"of":7,"metrics":{"Mean IoU (class)":"47.96"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.01556","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}