{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-robust-representations-by-projecting-1","title":"Learning Robust Representations by Projecting Superficial Statistics Out","arxiv_id":"1903.06256","date":"2019-03-02","proceeding":"ICLR 2019 5","authors":["Haohan Wang","Zexue He","Zachary C. Lipton","Eric P. Xing"],"abstract":"Despite impressive performance as evaluated on i.i.d. holdout data, deep\nneural networks depend heavily on superficial statistics of the training data\nand are liable to break under distribution shift. For example, subtle changes\nto the background or texture of an image can break a seemingly powerful\nclassifier. Building on previous work on domain generalization, we hope to\nproduce a classifier that will generalize to previously unseen domains, even\nwhen domain identifiers are not available during training. This setting is\nchallenging because the model may extract many distribution-specific\n(superficial) signals together with distribution-agnostic (semantic) signals.\nTo overcome this challenge, we incorporate the gray-level co-occurrence matrix\n(GLCM) to extract patterns that our prior knowledge suggests are superficial:\nthey are sensitive to the texture but unable to capture the gestalt of an\nimage. Then we introduce two techniques for improving our networks'\nout-of-sample performance. The first method is built on the reverse gradient\nmethod that pushes our model to learn representations from which the GLCM\nrepresentation is not predictable. The second method is built on the\nindependence introduced by projecting the model's representation onto the\nsubspace orthogonal to GLCM representation's. We test our method on the battery\nof standard domain generalization data sets and, interestingly, achieve\ncomparable or better performance as compared to other domain generalization\nmethods that explicitly require samples from the target distribution for\ntraining.","url_abs":"http://arxiv.org/abs/1903.06256v1","url_pdf":"http://arxiv.org/pdf/1903.06256v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"domain-generalization","task_name":"Domain Generalization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/domain-generalization-on-pacs-2","task":"Domain Generalization","dataset":"PACS","model":"Hex (Alexnet)","rank_in_archive_order":123,"of":133,"metrics":{"Average Accuracy":"70.20"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1903.06256","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}