{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/infmae-a-foundation-model-in-infrared","title":"InfMAE: A Foundation Model in the Infrared Modality","arxiv_id":"2402.00407","date":"2024-02-01","proceeding":null,"authors":["Fangcen Liu","Chenqiang Gao","Yaming Zhang","Junjie Guo","Jinhao Wang","Deyu Meng"],"abstract":"In recent years, the foundation models have swept the computer vision field and facilitated the development of various tasks within different modalities. However, it remains an open question on how to design an infrared foundation model. In this paper, we propose InfMAE, a foundation model in infrared modality. We release an infrared dataset, called Inf30 to address the problem of lacking large-scale data for self-supervised learning in the infrared vision community. Besides, we design an information-aware masking strategy, which is suitable for infrared images. This masking strategy allows for a greater emphasis on the regions with richer information in infrared images during the self-supervised learning process, which is conducive to learning the generalized representation. In addition, we adopt a multi-scale encoder to enhance the performance of the pre-trained encoders in downstream tasks. Finally, based on the fact that infrared images do not have a lot of details and texture information, we design an infrared decoder module, which further improves the performance of downstream tasks. Extensive experiments show that our proposed method InfMAE outperforms other supervised methods and self-supervised learning methods in three downstream tasks.","url_abs":"https://arxiv.org/abs/2402.00407v2","url_pdf":"https://arxiv.org/pdf/2402.00407v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"infmae-a-foundation-model-in-infrared","repo_url":"https://github.com/liufangcen/infmae","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.00407","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.00407"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/liufangcen/infmae","reach":null}],"summary":{"ran":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"186d4bf946b6f0ae","entry":"CBlock","repo":"liufangcen/infmae","repo_kind":"official","path":"models_infmae_skip4.py","file_url":"https://github.com/liufangcen/infmae/blob/HEAD/models_infmae_skip4.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"186d4bf946b6f0ae"}},{"code_sha256_prefix":"ece1e5ecfc0c52e1","entry":"PatchEmbed","repo":"liufangcen/infmae","repo_kind":"official","path":"models_infmae_skip4.py","file_url":"https://github.com/liufangcen/infmae/blob/HEAD/models_infmae_skip4.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ece1e5ecfc0c52e1"}},{"code_sha256_prefix":"7de4964e32c74039","entry":"PatchEmbed_F","repo":"liufangcen/infmae","repo_kind":"official","path":"models_infmae_skip4.py","file_url":"https://github.com/liufangcen/infmae/blob/HEAD/models_infmae_skip4.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7de4964e32c74039"}},{"code_sha256_prefix":"4c03a6b8d6c59dd5","entry":"MaskedAutoencoderInfMAE","repo":"liufangcen/infmae","repo_kind":"official","path":"models_infmae_skip4.py","file_url":"https://github.com/liufangcen/infmae/blob/HEAD/models_infmae_skip4.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4c03a6b8d6c59dd5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}