{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/token-wise-decomposition-of-autoregressive","title":"Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model Predictions","arxiv_id":"2305.10614","date":"2023-05-17","proceeding":null,"authors":["Byung-Doh Oh","William Schuler"],"abstract":"While there is much recent interest in studying why Transformer-based large language models make predictions the way they do, the complex computations performed within each layer have made their behavior somewhat opaque. To mitigate this opacity, this work presents a linear decomposition of final hidden states from autoregressive language models based on each initial input token, which is exact for virtually all contemporary Transformer architectures. This decomposition allows the definition of probability distributions that ablate the contribution of specific input tokens, which can be used to analyze their influence on model probabilities over a sequence of upcoming words with only one forward pass from the model. Using the change in next-word probability as a measure of importance, this work first examines which context words make the biggest contribution to language model predictions. Regression experiments suggest that Transformer-based language models rely primarily on collocational associations, followed by linguistic factors such as syntactic dependencies and coreference relationships in making next-word predictions. Additionally, analyses using these measures to predict syntactic dependencies and coreferent mention spans show that collocational association and repetitions of the same token largely explain the language models' predictions on these tasks.","url_abs":"https://arxiv.org/abs/2305.10614v2","url_pdf":"https://arxiv.org/pdf/2305.10614v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"token-wise-decomposition-of-autoregressive","repo_url":"https://github.com/byungdoh/llm_decomposition","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"model","task_name":"model"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2305.10614","atlas_url":"https://app.syntology.ai/?focus=2305.10614","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2305.10614"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/byungdoh/llm_decomposition","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":3,"unverified":9},"by_repo_kind":{"official":{"samples":12,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a727df5fd2f5599a","entry":"check_number_comma","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/convert_slow_tokenizer.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/convert_slow_tokenizer.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a727df5fd2f5599a"}},{"code_sha256_prefix":"b220e701be05a184","entry":"ensure_valid_input","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/convert_graph_to_onnx.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/convert_graph_to_onnx.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b220e701be05a184"}},{"code_sha256_prefix":"7efe3db5c9e2f94d","entry":"generate_identified_filename","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/convert_graph_to_onnx.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/convert_graph_to_onnx.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"7efe3db5c9e2f94d"}},{"code_sha256_prefix":"9d93436a1b858d0e","entry":"convert_slow_tokenizer","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/convert_slow_tokenizer.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/convert_slow_tokenizer.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9d93436a1b858d0e"}},{"code_sha256_prefix":"37a5eed2dbd663ca","entry":"gelu_fast","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/activations_tf.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/activations_tf.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"37a5eed2dbd663ca"}},{"code_sha256_prefix":"a00c425fa9fa3aed","entry":"get_architectures_from_config_class","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/utils/create_dummy_models.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/utils/create_dummy_models.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a00c425fa9fa3aed"}},{"code_sha256_prefix":"8291db316b0e343a","entry":"get_config_class_from_processor_class","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/utils/create_dummy_models.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/utils/create_dummy_models.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8291db316b0e343a"}},{"code_sha256_prefix":"dd3166a87bc972ff","entry":"get_configuration_file","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/configuration_utils.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/configuration_utils.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"dd3166a87bc972ff"}},{"code_sha256_prefix":"5503523a9bbdff58","entry":"get_processor_types_from_config_class","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/utils/create_dummy_models.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/utils/create_dummy_models.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"5503523a9bbdff58"}},{"code_sha256_prefix":"0da835817b8a57bc","entry":"infer_shapes","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/convert_graph_to_onnx.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/convert_graph_to_onnx.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"0da835817b8a57bc"}},{"code_sha256_prefix":"cc8c8c3ebf0c343f","entry":"mish","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/activations_tf.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/activations_tf.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"cc8c8c3ebf0c343f"}},{"code_sha256_prefix":"e2d56cb91f999bef","entry":"quick_gelu","repo":"byungdoh/llm_decomposition","repo_kind":"official","path":"huggingface/src/transformers/activations_tf.py","file_url":"https://github.com/byungdoh/llm_decomposition/blob/HEAD/huggingface/src/transformers/activations_tf.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e2d56cb91f999bef"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}