{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/understanding-and-improving-feature-learning-1","title":"Understanding and Improving Feature Learning for Out-of-Distribution Generalization","arxiv_id":"2304.11327","date":"2023-04-22","proceeding":"NeurIPS 2023 11","authors":["Yongqiang Chen","Wei Huang","Kaiwen Zhou","Yatao Bian","Bo Han","James Cheng"],"abstract":"A common explanation for the failure of out-of-distribution (OOD) generalization is that the model trained with empirical risk minimization (ERM) learns spurious features instead of invariant features. However, several recent studies challenged this explanation and found that deep networks may have already learned sufficiently good features for OOD generalization. Despite the contradictions at first glance, we theoretically show that ERM essentially learns both spurious and invariant features, while ERM tends to learn spurious features faster if the spurious correlation is stronger. Moreover, when fed the ERM learned features to the OOD objectives, the invariant feature learning quality significantly affects the final OOD performance, as OOD objectives rarely learn new features. Therefore, ERM feature learning can be a bottleneck to OOD generalization. To alleviate the reliance, we propose Feature Augmented Training (FeAT), to enforce the model to learn richer features ready for OOD generalization. FeAT iteratively augments the model to learn new features while retaining the already learned features. In each round, the retention and augmentation operations are performed on different subsets of the training data that capture distinct features. Extensive experiments show that FeAT effectively learns richer features thus boosting the performance of various OOD objectives.","url_abs":"https://arxiv.org/abs/2304.11327v2","url_pdf":"https://arxiv.org/pdf/2304.11327v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"understanding-and-improving-feature-learning-1","repo_url":"https://github.com/lfhase/feat","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"out-of-distribution-generalization","task_name":"Out-of-Distribution Generalization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2304.11327","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2304.11327"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/pytorch/captum","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lfhase/feat","reach":{"status":"ok"}}],"summary":{"ran":3,"unverified":3},"by_repo_kind":{"found_in_text":{"samples":6,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"5a2b42df4b995db6","entry":"apply_gradient_requirements","repo":"pytorch/captum","repo_kind":"found_in_text","path":"captum/_utils/gradient.py","file_url":"https://github.com/pytorch/captum/blob/HEAD/captum/_utils/gradient.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"5a2b42df4b995db6"}},{"code_sha256_prefix":"25d60f9ffaab0b57","entry":"parse_version","repo":"pytorch/captum","repo_kind":"found_in_text","path":"captum/_utils/common.py","file_url":"https://github.com/pytorch/captum/blob/HEAD/captum/_utils/common.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"25d60f9ffaab0b57"}},{"code_sha256_prefix":"0e5cf30aeea28b18","entry":"safe_div","repo":"pytorch/captum","repo_kind":"found_in_text","path":"captum/_utils/common.py","file_url":"https://github.com/pytorch/captum/blob/HEAD/captum/_utils/common.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"0e5cf30aeea28b18"}},{"code_sha256_prefix":"3102f2fd6c8bff3d","entry":"compute_gradients","repo":"pytorch/captum","repo_kind":"found_in_text","path":"captum/_utils/gradient.py","file_url":"https://github.com/pytorch/captum/blob/HEAD/captum/_utils/gradient.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"3102f2fd6c8bff3d"}},{"code_sha256_prefix":"0841c88ad916c6ee","entry":"compute_layer_gradients_and_eval","repo":"pytorch/captum","repo_kind":"found_in_text","path":"captum/_utils/gradient.py","file_url":"https://github.com/pytorch/captum/blob/HEAD/captum/_utils/gradient.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"0841c88ad916c6ee"}},{"code_sha256_prefix":"7dde11be93a92ca2","entry":"progress","repo":"pytorch/captum","repo_kind":"found_in_text","path":"captum/_utils/progress.py","file_url":"https://github.com/pytorch/captum/blob/HEAD/captum/_utils/progress.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"7dde11be93a92ca2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}