{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recurrent-transformer-networks-for-semantic","title":"Recurrent Transformer Networks for Semantic Correspondence","arxiv_id":"1810.12155","date":"2018-10-29","proceeding":"NeurIPS 2018 12","authors":["Seungryong Kim","Stephen Lin","Sangryul Jeon","Dongbo Min","Kwanghoon Sohn"],"abstract":"We present recurrent transformer networks (RTNs) for obtaining dense\ncorrespondences between semantically similar images. Our networks accomplish\nthis through an iterative process of estimating spatial transformations between\nthe input images and using these transformations to generate aligned\nconvolutional activations. By directly estimating the transformations between\nan image pair, rather than employing spatial transformer networks to\nindependently normalize each individual image, we show that greater accuracy\ncan be achieved. This process is conducted in a recursive manner to refine both\nthe transformation estimates and the feature representations. In addition, a\ntechnique is presented for weakly-supervised training of RTNs that is based on\na proposed classification loss. With RTNs, state-of-the-art performance is\nattained on several benchmarks for semantic correspondence.","url_abs":"http://arxiv.org/abs/1810.12155v1","url_pdf":"http://arxiv.org/pdf/1810.12155v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"recurrent-transformer-networks-for-semantic","repo_url":"https://github.com/seungryong/RTNs","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"semantic-correspondence","task_name":"Semantic correspondence"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"spatial-transformer","method_name":"Spatial Transformer"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.12155","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}