{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-ideation-execution-gap-execution-outcomes","title":"The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas","arxiv_id":"2506.20803","date":"2025-06-25","proceeding":null,"authors":["Chenglei Si","Tatsunori Hashimoto","Diyi Yang"],"abstract":"Large Language Models (LLMs) have shown promise in accelerating the scientific research pipeline. A key capability for this process is the ability to generate novel research ideas, and prior studies have found settings in which LLM-generated research ideas were judged as more novel than human-expert ideas. However, a good idea should not simply appear to be novel, it should also result in better research after being executed. To test whether AI-generated ideas lead to better research outcomes, we conduct an execution study by recruiting 43 expert researchers to execute randomly-assigned ideas, either written by experts or generated by an LLM. Each expert spent over 100 hours implementing the idea and wrote a 4-page short paper to document the experiments. All the executed projects are then reviewed blindly by expert NLP researchers. Comparing the review scores of the same ideas before and after execution, the scores of the LLM-generated ideas decrease significantly more than expert-written ideas on all evaluation metrics (novelty, excitement, effectiveness, and overall; p < 0.05), closing the gap between LLM and human ideas observed at the ideation stage. When comparing the aggregated review scores from the execution study, we even observe that for many metrics there is a flip in rankings where human ideas score higher than LLM ideas. This ideation-execution gap highlights the limitations of current LLMs in generating truly effective research ideas and the challenge of evaluating research ideas in the absence of execution outcomes.","url_abs":"https://arxiv.org/abs/2506.20803v1","url_pdf":"https://arxiv.org/pdf/2506.20803v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-ideation-execution-gap-execution-outcomes","repo_url":"https://github.com/NoviScl/AI-Researcher","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[{"method_slug":"flip","method_name":"FLIP"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2506.20803","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2506.20803"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/NoviScl/AI-Researcher","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":5},"by_repo_kind":{"official":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"128a74f95150bb65","entry":"count_ideas_in_directory","repo":"NoviScl/AI-Researcher","repo_kind":"official","path":"ai_researcher/src/count_ideas.py","file_url":"https://github.com/NoviScl/AI-Researcher/blob/HEAD/ai_researcher/src/count_ideas.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"128a74f95150bb65"}},{"code_sha256_prefix":"8572718b11af8061","entry":"find_representative_paper","repo":"NoviScl/AI-Researcher","repo_kind":"official","path":"ai_researcher/src/analyze_experiment_plans_semantic_similarity.py","file_url":"https://github.com/NoviScl/AI-Researcher/blob/HEAD/ai_researcher/src/analyze_experiment_plans_semantic_similarity.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8572718b11af8061"}},{"code_sha256_prefix":"8841bc0d03a60e22","entry":"get_top_n_and_lowest_n_papers","repo":"NoviScl/AI-Researcher","repo_kind":"official","path":"ai_researcher/src/analyze_scores.py","file_url":"https://github.com/NoviScl/AI-Researcher/blob/HEAD/ai_researcher/src/analyze_scores.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8841bc0d03a60e22"}},{"code_sha256_prefix":"aa1065b49e195e70","entry":"jaccard_similarity","repo":"NoviScl/AI-Researcher","repo_kind":"official","path":"ai_researcher/src/analyze_experiment_plans_semantic_similarity.py","file_url":"https://github.com/NoviScl/AI-Researcher/blob/HEAD/ai_researcher/src/analyze_experiment_plans_semantic_similarity.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"aa1065b49e195e70"}},{"code_sha256_prefix":"2098ab2bf0e96cf0","entry":"process_text","repo":"NoviScl/AI-Researcher","repo_kind":"official","path":"ai_researcher/src/analyze_experiment_plans_semantic_similarity.py","file_url":"https://github.com/NoviScl/AI-Researcher/blob/HEAD/ai_researcher/src/analyze_experiment_plans_semantic_similarity.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2098ab2bf0e96cf0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}