{"about":{"non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","site":"https://codewithpapers.app","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page","syntology":{"site":"https://syntology.ai","developers":"https://syntology.ai/developers","mcp":{"server":"https://syntology.ai/mcp","transport":"streamable-http","server_card":"https://syntology.ai/.well-known/mcp/server-card.json","auth":{"type":"trial token, no account","trial_token":"https://syntology.ai/api/oauth/trial/token","method":"POST","docs":"https://syntology.ai/developers"}},"have":"https://syntology.ai/api/graph/have?x=<method, arXiv id or title> (free, answers coverage only)","paper_base":"https://syntology.ai/paper/","atlas_base":"https://app.syntology.ai/?focus="},"machine_readable":[{"url":"https://codewithpapers.app/llms.txt","what":"the machine catalog: every machine-readable file, counted"},{"url":"https://codewithpapers.app/index/manifest.json","what":"paper-to-code index by arXiv id, with Syntology's counts"},{"url":"https://codewithpapers.app/search/manifest.json","what":"site search index (titles, authors) and its files"},{"url":"https://codewithpapers.app/download","what":"bulk files: Syntology's layer, described there"},{"url":"https://codewithpapers.app/build_manifest.json","what":"the build record: inputs, counts, exclusions, probes"}]},"url":"/paper/arxiv-2609-03887","title":"Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness","arxiv_id":"2609.03887","date":"2026-09-03","proceeding":null,"authors":["Hoang Cuong Nguyen","Mark Dras","Usman Naseem"],"abstract":"How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no method we study achieves all three properties we would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits. We caution against treating current post-training methods as a solved, reliable defense, especially for security-critical use. Code and models are available in https://github.com/hoangcuongnguyen2001/Beyond-Shallow-Alignment.","url_abs":"https://arxiv.org/abs/2609.03887","url_pdf":"https://arxiv.org/pdf/2609.03887","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2609.03887","atlas_url":"https://app.syntology.ai/?focus=2609.03887","mcp":null,"developers":"https://syntology.ai/developers","held":true,"reason":"in_graph_layer","checked_against":null,"agent_calls":[{"tool":"get_citation_path","arguments":{"paper_1":"2609.03887"},"arguments_in":{"paper_2":"another paper's arXiv id or title"},"call":"get_citation_path(paper_1=\"2609.03887\", paper_2=\"…\")"},{"tool":"get_concepts_for_paper","arguments":{"arxiv_id":"2609.03887"},"call":"get_concepts_for_paper(arxiv_id=\"2609.03887\")"}],"by_repo":[{"repo":"hoangcuongnguyen2001/Be","kind":"found_in_text","samples":0,"ran":0,"instrument":0,"unverified":0}],"by_repo_is":"one row per repository in the archive's code links, Syntology's graph links or the samples; samples 0 means no sample is linked to this paper, a statement about Syntology's coverage, not about the repository: with harvested_for_other_papers present, Syntology harvested the repository for that many other papers, and without it, Syntology harvested nothing from it; placed_by_identical_code counts samples whose link records no repository, placed here because identical code was harvested from this repository for another paper; ran counts executions on a synthesized input and instrument counts failures of Syntology's instrument, not of the code","read_at":"2026-09-28T10:30:06+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","ran_record":{"url":"https://syntology.ai/api/ran/2609.03887.json","schema":"syntology.ran_record/1","counts_crosswalk":{"n_samples":"syntology.counts.lifted","n_ran":"syntology.counts.executed","n_ran_checked":"syntology.counts.checked","n_instrument":"syntology.counts.instrument_failures","n_unverified":"syntology.counts.no_recorded_run"},"note":"live record, measured at request time; may be newer than this page's graph read","badge":"https://syntology.ai/api/ran/2609.03887.svg"},"claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/hoangcuongnguyen2001/Be","reach":null}],"summary":{},"n_samples":0,"n_ran":0,"n_constructed":0,"n_ran_checked":0,"n_instrument":0,"n_unverified":0,"n_instrument_is":"failures of Syntology's instrument, not of the code","every_run_is_an_instrument_failure":false,"key_notes":{"n_ran_checked":"legacy name, kept unchanged so existing readers do not break: it counts the samples that ran with no instrument failure (honoured, violated, and ran with no contract checked); it does not mean a contract was checked, and the pages print it as 'K with no instrument failure', not 'K checked'","n_constructed":"a sub-count of the samples that ran, never subtracted from them and never a failure: an executed sample whose run returned an instance of its own class (fixture_out_type equals the entry name): the run built an object and did not compute a result (Syntology's RAN record, counts.constructed)"},"by_repo_kind":{},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[]},"code_links_note":"this paper is newer than the archive, which holds no code links for it; its repositories, from Syntology's graph, are under syntology.repos","arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.CL","source":"arxiv_daily_20260904.json"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 8,886 of the 9,662 papers on this site that are newer than the archive (for 3,717 of them no archive leaderboard matched the paper's tables, so there was nothing further to check); 733 were read and have nothing to place (arXiv has no HTML version of the paper, or that version has no tables), 41 could not be read (the extractor's reply could not be parsed), and results from the other 2 appear after they are checked.","papers_newer_than_archive":9662,"papers_checked":8886,"papers_read_nothing_to_place":733,"papers_could_not_be_read":41},"entries":[],"judged_to_report_on":{"model":"Claude Sonnet 4.5","model_as_recorded":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","note":"a model's judgement from the paper's own tables, not a result; outcome is what Syntology's checks did with it; reason is the rules' plain words, only for refused_by_rule","boards":[]},"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}},"record_sha256":"8967455877616109fdafec91aa589b0bbd339326c332b9b340beab5455d8a3da","record_changed_at":"2026-09-28","record_changed_at_basis":"first_hashed"}