{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pixelnet-representation-of-the-pixels-by-the","title":"PixelNet: Representation of the pixels, by the pixels, and for the pixels","arxiv_id":"1702.06506","date":"2017-02-21","proceeding":null,"authors":["Aayush Bansal","Xinlei Chen","Bryan Russell","Abhinav Gupta","Deva Ramanan"],"abstract":"We explore design principles for general pixel-level prediction problems,\nfrom low-level edge detection to mid-level surface normal estimation to\nhigh-level semantic segmentation. Convolutional predictors, such as the\nfully-convolutional network (FCN), have achieved remarkable success by\nexploiting the spatial redundancy of neighboring pixels through convolutional\nprocessing. Though computationally efficient, we point out that such approaches\nare not statistically efficient during learning precisely because spatial\nredundancy limits the information learned from neighboring pixels. We\ndemonstrate that stratified sampling of pixels allows one to (1) add diversity\nduring batch updates, speeding up learning; (2) explore complex nonlinear\npredictors, improving accuracy; and (3) efficiently train state-of-the-art\nmodels tabula rasa (i.e., \"from scratch\") for diverse pixel-labeling tasks. Our\nsingle architecture produces state-of-the-art results for semantic segmentation\non PASCAL-Context dataset, surface normal estimation on NYUDv2 depth dataset,\nand edge detection on BSDS.","url_abs":"http://arxiv.org/abs/1702.06506v1","url_pdf":"http://arxiv.org/pdf/1702.06506v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pixelnet-representation-of-the-pixels-by-the","repo_url":"https://github.com/bdecost/pixelnet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"edge-detection","task_name":"Edge Detection"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"surface-normal-estimation","task_name":"Surface Normal Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.06506","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}