{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/polar-transformer-networks","title":"Polar Transformer Networks","arxiv_id":"1709.01889","date":"2017-09-06","proceeding":"ICLR 2018 1","authors":["Carlos Esteves","Christine Allen-Blanchette","Xiaowei Zhou","Kostas Daniilidis"],"abstract":"Convolutional neural networks (CNNs) are inherently equivariant to\ntranslation. Efforts to embed other forms of equivariance have concentrated\nsolely on rotation. We expand the notion of equivariance in CNNs through the\nPolar Transformer Network (PTN). PTN combines ideas from the Spatial\nTransformer Network (STN) and canonical coordinate representations. The result\nis a network invariant to translation and equivariant to both rotation and\nscale. PTN is trained end-to-end and composed of three distinct stages: a polar\norigin predictor, the newly introduced polar transformer module and a\nclassifier. PTN achieves state-of-the-art on rotated MNIST and the newly\nintroduced SIM2MNIST dataset, an MNIST variation obtained by adding clutter and\nperturbing digits with translation, rotation and scaling. The ideas of PTN are\nextensible to 3D which we demonstrate through the Cylindrical Transformer\nNetwork.","url_abs":"http://arxiv.org/abs/1709.01889v3","url_pdf":"http://arxiv.org/pdf/1709.01889v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"polar-transformer-networks","repo_url":"https://github.com/daniilidis-group/polar-transformer-networks","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"rotated-mnist","task_name":"Rotated MNIST"},{"task_slug":"translation","task_name":"Translation"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.01889","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}