{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2601-12890","title":"Efficient Code Analysis via Graph Representation Learning-Guided Large Language Models","arxiv_id":"2601.12890","date":"2026-01-19","proceeding":"ICML","authors":["Hang Gao","Tao Peng","Baoquan Cui","Hong Huang","Fengge Wu","Junsuo Zhao","Jian Zhang"],"abstract":"Large Language Models (LLMs) have significantly advanced code analysis tasks, yet they struggle to detect malicious behaviors fragmented across files, whose intricate dependencies easily get lost in the vast amount of benign code. We therefore propose a graph-centric attention acquisition pipeline that enhances LLMs' ability to localize malicious behavior. The approach parses a project into a code graph, uses an LLM to encode nodes with semantic and structural signals, and trains a Graph Neural Network (GNN) under sparse supervision. The GNN performs an initial detection, and by interpreting these predictions, identifies key code sections that are most likely to contain malicious behavior. These influential regions are then used to guide the LLM's attention for in-depth analysis. This strategy significantly reduces interference from irrelevant context while maintaining low annotation costs. Extensive experiments show that the method consistently outperforms existing approaches on multiple public and custom datasets, highlighting its potential for practical deployment in software security scenarios.","url_abs":"https://arxiv.org/abs/2601.12890","url_pdf":"https://arxiv.org/pdf/2601.12890","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2601.12890","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2601.12890"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/Epiphaniespt/GMLLM","reach":null}],"summary":{"ran":2,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"cfded09dafcc228c","entry":"GcnEncoderGraph","repo":"Epiphaniespt/GMLLM","repo_kind":"found_in_text","path":"GMLLM/GNNExplainer/model.py","file_url":"https://github.com/Epiphaniespt/GMLLM/blob/HEAD/GMLLM/GNNExplainer/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cfded09dafcc228c"}},{"code_sha256_prefix":"970d2c464403c7fa","entry":"GraphConv","repo":"Epiphaniespt/GMLLM","repo_kind":"found_in_text","path":"GMLLM/GNNExplainer/model.py","file_url":"https://github.com/Epiphaniespt/GMLLM/blob/HEAD/GMLLM/GNNExplainer/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"970d2c464403c7fa"}},{"code_sha256_prefix":"d568e30efe0c57e5","entry":"SoftPoolingGcnEncoder","repo":"Epiphaniespt/GMLLM","repo_kind":"found_in_text","path":"GMLLM/GNNExplainer/model.py","file_url":"https://github.com/Epiphaniespt/GMLLM/blob/HEAD/GMLLM/GNNExplainer/model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d568e30efe0c57e5"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.SE","source":"arxiv_api"},"syntology_extracted_results":null}