{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-iterative-refinement-of-densely","title":"On the iterative refinement of densely connected representation levels for semantic segmentation","arxiv_id":"1804.11332","date":"2018-04-30","proceeding":null,"authors":["Arantxa Casanova","Guillem Cucurull","Michal Drozdzal","Adriana Romero","Yoshua Bengio"],"abstract":"State-of-the-art semantic segmentation approaches increase the receptive\nfield of their models by using either a downsampling path composed of\npoolings/strided convolutions or successive dilated convolutions. However, it\nis not clear which operation leads to best results. In this paper, we\nsystematically study the differences introduced by distinct receptive field\nenlargement methods and their impact on the performance of a novel\narchitecture, called Fully Convolutional DenseResNet (FC-DRN). FC-DRN has a\ndensely connected backbone composed of residual networks. Following standard\nimage segmentation architectures, receptive field enlargement operations that\nchange the representation level are interleaved among residual networks. This\nallows the model to exploit the benefits of both residual and dense\nconnectivity patterns, namely: gradient flow, iterative refinement of\nrepresentations, multi-scale feature combination and deep supervision. In order\nto highlight the potential of our model, we test it on the challenging CamVid\nurban scene understanding benchmark and make the following observations: 1)\ndownsampling operations outperform dilations when the model is trained from\nscratch, 2) dilations are useful during the finetuning step of the model, 3)\ncoarser representations require less refinement steps, and 4) ResNets (by model\nconstruction) are good regularizers, since they can reduce the model capacity\nwhen needed. Finally, we compare our architecture to alternative methods and\nreport state-of-the-art result on the Camvid dataset, with at least twice fewer\nparameters.","url_abs":"http://arxiv.org/abs/1804.11332v1","url_pdf":"http://arxiv.org/pdf/1804.11332v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-iterative-refinement-of-densely","repo_url":"https://github.com/ArantxaCasanova/fc-drn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1804.11332","atlas_url":"https://app.syntology.ai/?focus=1804.11332","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}