{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2605-13043","title":"Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models","arxiv_id":"2605.13043","date":"2026-05-13","proceeding":null,"authors":["Yejin Lee","Yo-Sub Han"],"abstract":"Diffusion Language Models (DLMs) provide a promising alternative to autoregressive language models by generating text through iterative denoising and bidirectional refinement. However, this iterative generation paradigm also introduces unique safety vulnerabilities when harmful tokens generated at intermediate denoising steps propagate through subsequent refinement processes and eventually induce unsafe outputs. While there are a few attempts to remedy this issue, they either fail to generate safe outputs or generate safe yet low-quality outputs. This motivates us to propose an inference-time defense framework based on the step-wise intervention during the denoising process, which then improves the safety without compromising the output quality. The key component of our framework is a contrastive safety direction (SGD), a latent direction that captures the semantic boundary between harmful and safe generations. We leverage SGD to assess the alignment of generated tokens with harmful semantics at each denoising step. When harmful alignment is detected, our method remasks the corresponding tokens and resumes the denoising process with adaptive steering, where the steering strength is modulated according to the estimated degree of harmfulness. As a plug-and-play module, our method circumvents the need for additional fine-tuning and can be directly incorporated into off-the-shelf diffusion models. The experimental results show that our approaches reduce jailbreak success rates to 0.64% while preserving generation quality close to the original model performance. This confirms the effectiveness of step-wise intervention for safe diffusion language model generation. Our code is available at https://github.com/leeyejin1231/DLM_Steering_Remasking.","url_abs":"https://arxiv.org/abs/2605.13043","url_pdf":"https://arxiv.org/pdf/2605.13043","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2605.13043","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2605.13043"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/leeyejin1231/DLM_Steering_Remasking","reach":null}],"summary":{"ran":9},"by_repo_kind":{"found_in_text":{"samples":9,"ran":9,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":9,"samples":[{"code_sha256_prefix":"c169bd06be5b3321","entry":"build_sequence","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/make_csd_llada.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/make_csd_llada.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c169bd06be5b3321"}},{"code_sha256_prefix":"56f597e42de94f88","entry":"evaluate","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/mmlu_eval.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/mmlu_eval.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"56f597e42de94f88"}},{"code_sha256_prefix":"2e86f3b1851c2105","entry":"extract_pred","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/MATH-500_eval.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/MATH-500_eval.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2e86f3b1851c2105"}},{"code_sha256_prefix":"4f2590062677868c","entry":"find_transformer_blocks","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/make_csd_dream.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/make_csd_dream.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4f2590062677868c"}},{"code_sha256_prefix":"7b6c6fcf52ac97a7","entry":"last_boxed","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/MATH-500_eval.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/MATH-500_eval.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7b6c6fcf52ac97a7"}},{"code_sha256_prefix":"874f1c16ed0a11ed","entry":"load_harmful_data","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/make_csd_dream.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/make_csd_dream.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"874f1c16ed0a11ed"}},{"code_sha256_prefix":"45332b8855ee51fa","entry":"load_refusals","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/make_csd_dream.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/make_csd_dream.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"45332b8855ee51fa"}},{"code_sha256_prefix":"bae3b10ba5297700","entry":"parse_response","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/mmlu_eval.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/mmlu_eval.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bae3b10ba5297700"}},{"code_sha256_prefix":"fbfcafd9f1e99a1f","entry":"strip_boxed","repo":"leeyejin1231/DLM_Steering_Remasking","repo_kind":"found_in_text","path":"utils/MATH-500_eval.py","file_url":"https://github.com/leeyejin1231/DLM_Steering_Remasking/blob/HEAD/utils/MATH-500_eval.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fbfcafd9f1e99a1f"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.CL","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}