{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/iterative-reorganization-with-weak-spatial","title":"Iterative Reorganization with Weak Spatial Constraints: Solving Arbitrary Jigsaw Puzzles for Unsupervised Representation Learning","arxiv_id":"1812.00329","date":"2018-12-02","proceeding":"CVPR 2019 6","authors":["Chen Wei","Lingxi Xie","Xutong Ren","Yingda Xia","Chi Su","Jiaying Liu","Qi Tian","Alan L. Yuille"],"abstract":"Learning visual features from unlabeled image data is an important yet\nchallenging task, which is often achieved by training a model on some\nannotation-free information. We consider spatial contexts, for which we solve\nso-called jigsaw puzzles, i.e., each image is cut into grids and then\ndisordered, and the goal is to recover the correct configuration. Existing\napproaches formulated it as a classification task by defining a fixed mapping\nfrom a small subset of configurations to a class set, but these approaches\nignore the underlying relationship between different configurations and also\nlimit their application to more complex scenarios. This paper presents a novel\napproach which applies to jigsaw puzzles with an arbitrary grid size and\ndimensionality. We provide a fundamental and generalized principle, that weaker\ncues are easier to be learned in an unsupervised manner and also transfer\nbetter. In the context of puzzle recognition, we use an iterative manner which,\ninstead of solving the puzzle all at once, adjusts the order of the patches in\neach step until convergence. In each step, we combine both unary and binary\nfeatures on each patch into a cost function judging the correctness of the\ncurrent configuration. Our approach, by taking similarity between puzzles into\nconsideration, enjoys a more reasonable way of learning visual knowledge. We\nverify the effectiveness of our approach in two aspects. First, it is able to\nsolve arbitrarily complex puzzles, including high-dimensional puzzles, that\nprior methods are difficult to handle. Second, it serves as a reliable way of\nnetwork initialization, which leads to better transfer performance in a few\nvisual recognition tasks including image classification, object detection, and\nsemantic segmentation.","url_abs":"http://arxiv.org/abs/1812.00329v1","url_pdf":"http://arxiv.org/pdf/1812.00329v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"iterative-reorganization-with-weak-spatial","repo_url":"https://github.com/weichen582/Unsupervised-Visual-Recognition-by-Solving-Arbitrary-Puzzles","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"image-classification","task_name":"image-classification"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"jigsaw","method_name":"Jigsaw"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.00329","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}