{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fine-grained-image-style-transfer-with-visual","title":"Fine-Grained Image Style Transfer with Visual Transformers","arxiv_id":"2210.05176","date":"2022-10-11","proceeding":null,"authors":["Jianbo Wang","Huan Yang","Jianlong Fu","Toshihiko Yamasaki","Baining Guo"],"abstract":"With the development of the convolutional neural network, image style transfer has drawn increasing attention. However, most existing approaches adopt a global feature transformation to transfer style patterns into content images (e.g., AdaIN and WCT). Such a design usually destroys the spatial information of the input images and fails to transfer fine-grained style patterns into style transfer results. To solve this problem, we propose a novel STyle TRansformer (STTR) network which breaks both content and style images into visual tokens to achieve a fine-grained style transformation. Specifically, two attention mechanisms are adopted in our STTR. We first propose to use self-attention to encode content and style tokens such that similar tokens can be grouped and learned together. We then adopt cross-attention between content and style tokens that encourages fine-grained style transformations. To compare STTR with existing approaches, we conduct user studies on Amazon Mechanical Turk (AMT), which are carried out with 50 human subjects with 1,000 votes in total. Extensive evaluations demonstrate the effectiveness and efficiency of the proposed STTR in generating visually pleasing style transfer results.","url_abs":"https://arxiv.org/abs/2210.05176v1","url_pdf":"https://arxiv.org/pdf/2210.05176v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fine-grained-image-style-transfer-with-visual","repo_url":"https://github.com/researchmm/sttr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"style-transfer","task_name":"Style Transfer"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2210.05176","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2210.05176"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/researchmm/sttr","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":2,"ran_violates":1,"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"029ef88513e67c12","entry":"conv1x1","repo":"researchmm/sttr","repo_kind":"official","path":"models_istt/backbone.py","file_url":"https://github.com/researchmm/sttr/blob/HEAD/models_istt/backbone.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"029ef88513e67c12"}},{"code_sha256_prefix":"ef02d64125a1bf3d","entry":"conv3x3","repo":"researchmm/sttr","repo_kind":"official","path":"models_istt/backbone.py","file_url":"https://github.com/researchmm/sttr/blob/HEAD/models_istt/backbone.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ef02d64125a1bf3d"}},{"code_sha256_prefix":"ac8fe530cdad4d8c","entry":"dice_loss","repo":"researchmm/sttr","repo_kind":"official","path":"models_istt/segmentation.py","file_url":"https://github.com/researchmm/sttr/blob/HEAD/models_istt/segmentation.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ac8fe530cdad4d8c"}},{"code_sha256_prefix":"5c0711aada67957e","entry":"sigmoid_focal_loss","repo":"researchmm/sttr","repo_kind":"official","path":"models_istt/segmentation.py","file_url":"https://github.com/researchmm/sttr/blob/HEAD/models_istt/segmentation.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5c0711aada67957e"}},{"code_sha256_prefix":"a7bb87542c72ca8e","entry":"build_transformer","repo":"researchmm/sttr","repo_kind":"official","path":"models_istt/transformer_nonorm_flx.py","file_url":"https://github.com/researchmm/sttr/blob/HEAD/models_istt/transformer_nonorm_flx.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a7bb87542c72ca8e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}