{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/convergent-learning-do-different-neural","title":"Convergent Learning: Do different neural networks learn the same representations?","arxiv_id":"1511.07543","date":"2015-11-24","proceeding":null,"authors":["Yixuan Li","Jason Yosinski","Jeff Clune","Hod Lipson","John Hopcroft"],"abstract":"Recent success in training deep neural networks have prompted active\ninvestigation into the features learned on their intermediate layers. Such\nresearch is difficult because it requires making sense of non-linear\ncomputations performed by millions of parameters, but valuable because it\nincreases our ability to understand current models and create improved versions\nof them. In this paper we investigate the extent to which neural networks\nexhibit what we call convergent learning, which is when the representations\nlearned by multiple nets converge to a set of features which are either\nindividually similar between networks or where subsets of features span similar\nlow-dimensional spaces. We propose a specific method of probing\nrepresentations: training multiple networks and then comparing and contrasting\ntheir individual, learned representations at the level of neurons or groups of\nneurons. We begin research into this question using three techniques to\napproximately align different neural networks on a feature level: a bipartite\nmatching approach that makes one-to-one assignments between neurons, a sparse\nprediction approach that finds one-to-many mappings, and a spectral clustering\napproach that finds many-to-many mappings. This initial investigation reveals a\nfew previously unknown properties of neural networks, and we argue that future\nresearch into the question of convergent learning will yield many more. The\ninsights described here include (1) that some features are learned reliably in\nmultiple networks, yet other features are not consistently learned; (2) that\nunits learn to span low-dimensional subspaces and, while these subspaces are\ncommon to multiple networks, the specific basis vectors learned are not; (3)\nthat the representation codes show evidence of being a mix between a local code\nand slightly, but not fully, distributed codes across multiple units.","url_abs":"http://arxiv.org/abs/1511.07543v3","url_pdf":"http://arxiv.org/pdf/1511.07543v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"convergent-learning-do-different-neural","repo_url":"https://github.com/yixuanli/convergent_learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1511.07543","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}