{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/analyzing-vision-transformers-for-image-1","title":"Analyzing Vision Transformers for Image Classification in Class Embedding Space","arxiv_id":"2310.18969","date":"2023-10-29","proceeding":"NeurIPS 2023 11","authors":["Martina G. Vilas","Timothy Schaumlöffel","Gemma Roig"],"abstract":"Despite the growing use of transformer models in computer vision, a mechanistic understanding of these networks is still needed. This work introduces a method to reverse-engineer Vision Transformers trained to solve image classification tasks. Inspired by previous research in NLP, we demonstrate how the inner representations at any level of the hierarchy can be projected onto the learned class embedding space to uncover how these networks build categorical representations for their predictions. We use our framework to show how image tokens develop class-specific representations that depend on attention mechanisms and contextual information, and give insights on how self-attention and MLP layers differentially contribute to this categorical composition. We additionally demonstrate that this method (1) can be used to determine the parts of an image that would be important for detecting the class of interest, and (2) exhibits significant advantages over traditional linear probing approaches. Taken together, our results position our proposed framework as a powerful tool for mechanistic interpretability and explainability research.","url_abs":"https://arxiv.org/abs/2310.18969v1","url_pdf":"https://arxiv.org/pdf/2310.18969v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"analyzing-vision-transformers-for-image-1","repo_url":"https://github.com/martinagvilas/vit-cls_emb","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.18969","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.18969"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/martinagvilas/vit-cls_emb","reach":null}],"summary":{"ran":1,"unverified":2},"by_repo_kind":{"official":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"a01d578faa9dac32","entry":"MyCIFAR100","repo":"martinagvilas/vit-cls_emb","repo_kind":"official","path":"src/identifiability.py","file_url":"https://github.com/martinagvilas/vit-cls_emb/blob/HEAD/src/identifiability.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a01d578faa9dac32"}},{"code_sha256_prefix":"836b7ba9e527c4df","entry":"ImagenetDatasetS","repo":"martinagvilas/vit-cls_emb","repo_kind":"official","path":"src/identifiability.py","file_url":"https://github.com/martinagvilas/vit-cls_emb/blob/HEAD/src/identifiability.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"836b7ba9e527c4df"}},{"code_sha256_prefix":"cc7e4c9b32b0b7f9","entry":"get_class_embed","repo":"martinagvilas/vit-cls_emb","repo_kind":"official","path":"src/identifiability.py","file_url":"https://github.com/martinagvilas/vit-cls_emb/blob/HEAD/src/identifiability.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cc7e4c9b32b0b7f9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}