{"about":{"non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","site":"https://codewithpapers.app","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page","syntology":{"site":"https://syntology.ai","developers":"https://syntology.ai/developers","mcp":{"server":"https://syntology.ai/mcp","transport":"streamable-http","server_card":"https://syntology.ai/.well-known/mcp/server-card.json","auth":{"type":"trial token, no account","trial_token":"https://syntology.ai/api/oauth/trial/token","method":"POST","docs":"https://syntology.ai/developers"}},"have":"https://syntology.ai/api/graph/have?x=<method, arXiv id or title> (free, answers coverage only)","paper_base":"https://syntology.ai/paper/","atlas_base":"https://app.syntology.ai/?focus="},"machine_readable":[{"url":"https://codewithpapers.app/llms.txt","what":"the machine catalog: every machine-readable file, counted"},{"url":"https://codewithpapers.app/index/manifest.json","what":"paper-to-code index by arXiv id, with Syntology's counts"},{"url":"https://codewithpapers.app/search/manifest.json","what":"site search index (titles, authors) and its files"},{"url":"https://codewithpapers.app/download","what":"bulk files: Syntology's layer, described there"},{"url":"https://codewithpapers.app/build_manifest.json","what":"the build record: inputs, counts, exclusions, probes"}]},"url":"/paper/arxiv-2609-28963","title":"Back to the Definition: Estimating Step-Level Advantages via Trajectory Graphs for Agentic Reinforcement Learning","arxiv_id":"2609.28963","date":"2026-09-24","proceeding":null,"authors":["Xincheng Yao","Haobo Fu","Weiming Liu","Chongyang Zhang"],"abstract":"Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value estimate. Extending the faithful estimation to step-level would in principle demand sampling multiple actions from each intermediate state, which is too costly on a per-state basis. To mitigate this issue, we propose a Graph-based Faithful sTep-level credit-assignment framework (GRAFT) that grafts all rollout trajectories into a trajectory graph, recovering node state-values via Bellman iteration on the graph, and assigning credit to each edge by the node value difference. Theoretically, the estimated step-level advantage faithfully adheres to the basic advantage definition in RL. To further ensure the reliability of step-level advantage estimation, we further propose Graph GAE, which extends GAE to the trajectory graph for reducing the impact of state-value estimation bias. Experiments across a range of multi-turn agentic benchmarks show consistent gains over GRPO and superior performance compared to recent agentic RL algorithms. Code will be available at https://github.com/xcyao00/GRAFT.","url_abs":"https://arxiv.org/abs/2609.28963","url_pdf":"https://arxiv.org/pdf/2609.28963","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2609.28963","atlas_url":"https://app.syntology.ai/?focus=2609.28963","mcp":null,"developers":"https://syntology.ai/developers","held":true,"reason":"in_graph_layer","checked_against":null,"agent_calls":[{"tool":"get_citation_path","arguments":{"paper_1":"2609.28963"},"arguments_in":{"paper_2":"another paper's arXiv id or title"},"call":"get_citation_path(paper_1=\"2609.28963\", paper_2=\"…\")"},{"tool":"get_concepts_for_paper","arguments":{"arxiv_id":"2609.28963"},"call":"get_concepts_for_paper(arxiv_id=\"2609.28963\")"}],"by_repo":[{"repo":"xcyao00/GRAFT","kind":"found_in_text","samples":0,"ran":0,"instrument":0,"unverified":0}],"by_repo_is":"one row per repository in the archive's code links, Syntology's graph links or the samples; samples 0 means no sample is linked to this paper, a statement about Syntology's coverage, not about the repository: with harvested_for_other_papers present, Syntology harvested the repository for that many other papers, and without it, Syntology harvested nothing from it; placed_by_identical_code counts samples whose link records no repository, placed here because identical code was harvested from this repository for another paper; ran counts executions on a synthesized input and instrument counts failures of Syntology's instrument, not of the code","read_at":"2026-09-28T10:30:06+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","ran_record":{"url":"https://syntology.ai/api/ran/2609.28963.json","schema":"syntology.ran_record/1","counts_crosswalk":{"n_samples":"syntology.counts.lifted","n_ran":"syntology.counts.executed","n_ran_checked":"syntology.counts.checked","n_instrument":"syntology.counts.instrument_failures","n_unverified":"syntology.counts.no_recorded_run"},"note":"live record, measured at request time; may be newer than this page's graph read","badge":"https://syntology.ai/api/ran/2609.28963.svg"},"claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/xcyao00/GRAFT","reach":null}],"summary":{},"n_samples":0,"n_ran":0,"n_constructed":0,"n_ran_checked":0,"n_instrument":0,"n_unverified":0,"n_instrument_is":"failures of Syntology's instrument, not of the code","every_run_is_an_instrument_failure":false,"key_notes":{"n_ran_checked":"legacy name, kept unchanged so existing readers do not break: it counts the samples that ran with no instrument failure (honoured, violated, and ran with no contract checked); it does not mean a contract was checked, and the pages print it as 'K with no instrument failure', not 'K checked'","n_constructed":"a sub-count of the samples that ran, never subtracted from them and never a failure: an executed sample whose run returned an instance of its own class (fixture_out_type equals the entry name): the run built an object and did not compute a result (Syntology's RAN record, counts.constructed)"},"by_repo_kind":{},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[]},"code_links_note":"this paper is newer than the archive, which holds no code links for it; its repositories, from Syntology's graph, are under syntology.repos","arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.AI","source":"arxiv_daily_20260925.json"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 8,886 of the 9,662 papers on this site that are newer than the archive (for 3,717 of them no archive leaderboard matched the paper's tables, so there was nothing further to check); 733 were read and have nothing to place (arXiv has no HTML version of the paper, or that version has no tables), 41 could not be read (the extractor's reply could not be parsed), and results from the other 2 appear after they are checked.","papers_newer_than_archive":9662,"papers_checked":8886,"papers_read_nothing_to_place":733,"papers_could_not_be_read":41},"entries":[],"judged_to_report_on":{"model":"Claude Sonnet 4.5","model_as_recorded":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","note":"a model's judgement from the paper's own tables, not a result; outcome is what Syntology's checks did with it; reason is the rules' plain words, only for refused_by_rule","boards":[]},"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}},"record_sha256":"0e294584bfe8808fbaa179c35d38653ccca3f4afa54994f29d8a285f2a9f4729","record_changed_at":"2026-09-28","record_changed_at_basis":"first_hashed"}