{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mat-mask-aware-transformer-for-large-hole","title":"MAT: Mask-Aware Transformer for Large Hole Image Inpainting","arxiv_id":"2203.15270","date":"2022-03-29","proceeding":"CVPR 2022 1","authors":["Wenbo Li","Zhe Lin","Kun Zhou","Lu Qi","Yi Wang","Jiaya Jia"],"abstract":"Recent studies have shown the importance of modeling long-range interactions in the inpainting problem. To achieve this goal, existing approaches exploit either standalone attention techniques or transformers, but usually under a low resolution in consideration of computational cost. In this paper, we present a novel transformer-based model for large hole inpainting, which unifies the merits of transformers and convolutions to efficiently process high-resolution images. We carefully design each component of our framework to guarantee the high fidelity and diversity of recovered images. Specifically, we customize an inpainting-oriented transformer block, where the attention module aggregates non-local information only from partial valid tokens, indicated by a dynamic mask. Extensive experiments demonstrate the state-of-the-art performance of the new model on multiple benchmark datasets. Code is released at https://github.com/fenglinglwb/MAT.","url_abs":"https://arxiv.org/abs/2203.15270v3","url_pdf":"https://arxiv.org/pdf/2203.15270v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mat-mask-aware-transformer-for-large-hole","repo_url":"https://github.com/fenglinglwb/mat","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"image-inpainting","task_name":"Image Inpainting"},{"task_slug":null,"task_name":"valid"}],"methods":[{"method_slug":"pixel-prediction","method_name":"Inpainting"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-inpainting-on-celeba-hq","task":"Image Inpainting","dataset":"CelebA-HQ","model":"MAT","rank_in_archive_order":1,"of":6,"metrics":{"FID":"4.86","P-IDS":"13.83","U-IDS":"25.33"},"uses_additional_data":false},{"leaderboard":"/sota/image-inpainting-on-places2-1","task":"Image Inpainting","dataset":"Places2","model":"MAT","rank_in_archive_order":3,"of":14,"metrics":{"FID":"1.96","P-IDS":"23.42","U-IDS":"38.34"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2203.15270","atlas_url":"https://app.syntology.ai/?focus=2203.15270","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}