{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/beyond-prompt-engineering-robust-behavior","title":"Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms","arxiv_id":"2505.20322","date":"2025-05-23","proceeding":null,"authors":["Mengru Wang","Ziwen Xu","Shengyu Mao","Shumin Deng","Zhaopeng Tu","Huajun Chen","Ningyu Zhang"],"abstract":"Precise control over language model generation is vital for ensuring both safety and reliability. Although prompt engineering and steering are commonly used to intervene in model behaviors, the vast number of parameters in models often results in highly intertwined internal representations. This interdependency can limit control precision and sometimes lead to unintended side effects. Recent research has explored the use of sparse autoencoders (SAE) to disentangle knowledge in high-dimensional spaces for steering. However, these applications have been limited to toy tasks owing to the nontrivial issue of locating atomic knowledge components. In this paper, we propose Steering Target Atoms (STA), a novel method that isolates and manipulates disentangled knowledge components to enhance safety. Comprehensive experiments demonstrate the effectiveness of our approach. Further analysis reveals that steering exhibits superior robustness and flexibility, particularly in adversarial scenarios. We also apply the steering strategy to the large reasoning model, confirming its effectiveness in precise reasoning control.","url_abs":"https://arxiv.org/abs/2505.20322v1","url_pdf":"https://arxiv.org/pdf/2505.20322v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"beyond-prompt-engineering-robust-behavior","repo_url":"https://github.com/zjunlp/steer-target-atoms","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"jax","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"prompt-engineering","task_name":"Prompt Engineering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2505.20322","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.20322"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zjunlp/steer-target-atoms","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":5},"by_repo_kind":{"official":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1c6c704aa0ae1ae2","entry":"extend_attention_mask_by_fixed_len","repo":"zjunlp/steer-target-atoms","repo_kind":"official","path":"sae_utils.py","file_url":"https://github.com/zjunlp/steer-target-atoms/blob/HEAD/sae_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1c6c704aa0ae1ae2"}},{"code_sha256_prefix":"61c4fdbffae9387f","entry":"extend_attention_mask_by_lens","repo":"zjunlp/steer-target-atoms","repo_kind":"official","path":"sae_utils.py","file_url":"https://github.com/zjunlp/steer-target-atoms/blob/HEAD/sae_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"61c4fdbffae9387f"}},{"code_sha256_prefix":"4e409608d73ba8ed","entry":"get_pred","repo":"zjunlp/steer-target-atoms","repo_kind":"official","path":"dataloader.py","file_url":"https://github.com/zjunlp/steer-target-atoms/blob/HEAD/dataloader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4e409608d73ba8ed"}},{"code_sha256_prefix":"e6f2de940dfd784f","entry":"prepare_input","repo":"zjunlp/steer-target-atoms","repo_kind":"official","path":"baseline/generate_vectors.py","file_url":"https://github.com/zjunlp/steer-target-atoms/blob/HEAD/baseline/generate_vectors.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e6f2de940dfd784f"}},{"code_sha256_prefix":"dee4be385625fb2b","entry":"signed_min_max_normalize","repo":"zjunlp/steer-target-atoms","repo_kind":"official","path":"generate_sae_caa_vector.py","file_url":"https://github.com/zjunlp/steer-target-atoms/blob/HEAD/generate_sae_caa_vector.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"dee4be385625fb2b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}