{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2602-03554","title":"When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs","arxiv_id":"2602.03554","date":"2026-02-03","proceeding":"ICML","authors":["Bogdan Zagribelnyy","Ivan Ilin","Maksim Kuznetsov","Nikita Bondarev","Mathieu Reymond","Roman Schutski","Thomas MacDougall","Rim Shayakhmetov","Zulfat Miftakhutdinov","Mikolaj Mizera","Vladimir Aladinskiy","Alex Aliper","Alex Zhavoronkov"],"abstract":"Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning. However, objective evaluation of retrosynthesis performance remains limited. Existing benchmarks and metrics typically rely on published synthetic procedures and Top-K accuracy based on single ground-truth, which does not capture the open-ended nature of real-world synthesis planning. We propose a new benchmarking framework for single-step retrosynthesis that evaluates both general-purpose and chemistry-specialized LLMs using ChemCensor, a novel metric for chemical plausibility. By emphasizing plausibility over exact match, this approach better aligns with human synthesis planning practices. We also introduce CREED, a novel dataset comprising millions of ChemCensor-validated reaction records for LLM training, and use it to train a model that improves over the LLM baselines under this benchmark.","url_abs":"https://arxiv.org/abs/2602.03554","url_pdf":"https://arxiv.org/pdf/2602.03554","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2602.03554","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2602.03554"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/pandegroup/reaction_prediction_seq2seq","reach":null},{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/pandegroup/reaction_prediction_seq2se","reach":{"status":"gone","observed_at":"2026-09-16","how":"tree_404+repo_404"}}],"summary":{"unverified":5},"by_repo_kind":{"found_in_text":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"181e72bda062f028","entry":"att_sum_bahdanau","repo":"pandegroup/reaction_prediction_seq2seq","repo_kind":"found_in_text","path":"seq2seq/decoders/attention.py","file_url":"https://github.com/pandegroup/reaction_prediction_seq2seq/blob/HEAD/seq2seq/decoders/attention.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"181e72bda062f028"}},{"code_sha256_prefix":"af4e61cae8ea2e6d","entry":"att_sum_dot","repo":"pandegroup/reaction_prediction_seq2seq","repo_kind":"found_in_text","path":"seq2seq/decoders/attention.py","file_url":"https://github.com/pandegroup/reaction_prediction_seq2seq/blob/HEAD/seq2seq/decoders/attention.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"af4e61cae8ea2e6d"}},{"code_sha256_prefix":"358434439625669b","entry":"cross_entropy_sequence_loss","repo":"pandegroup/reaction_prediction_seq2seq","repo_kind":"found_in_text","path":"seq2seq/losses.py","file_url":"https://github.com/pandegroup/reaction_prediction_seq2seq/blob/HEAD/seq2seq/losses.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"358434439625669b"}},{"code_sha256_prefix":"66b405ecdb68e352","entry":"get_dict_from_collection","repo":"pandegroup/reaction_prediction_seq2seq","repo_kind":"found_in_text","path":"seq2seq/graph_utils.py","file_url":"https://github.com/pandegroup/reaction_prediction_seq2seq/blob/HEAD/seq2seq/graph_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"66b405ecdb68e352"}},{"code_sha256_prefix":"c73e2ed8b047237f","entry":"templatemethod","repo":"pandegroup/reaction_prediction_seq2seq","repo_kind":"found_in_text","path":"seq2seq/graph_utils.py","file_url":"https://github.com/pandegroup/reaction_prediction_seq2seq/blob/HEAD/seq2seq/graph_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c73e2ed8b047237f"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}