{"url":"/sota/code-generation-on-webapp1k-react","task":{"name":"Code Generation","url":"/task/code-generation","note":null},"dataset":{"name":"WebApp1K-React","url":"/dataset/webapp1k-react"},"category":"Natural Language Processing","categories":["Computer Code","Natural Language Processing","Reasoning"],"category_note":null,"description":"**Code Generation** is an important field to predict explicit code or program structure from multimodal data sources such as incomplete code, programs in another programming language, natural language descriptions or execution examples. Code Generation tools can assist the development of automatic programming tools to improve programming productivity.\r\n\r\n\r\n<span class=\"description-source\">Source: [Deep Learning for Source Code Modeling and Generation ](https://arxiv.org/abs/2002.05442)</span>\r\n\r\nImage source: [Measuring Coding Challenge Competence With APPS](https://paperswithcode.com/paper/measuring-coding-challenge-competence-with)","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["pass@1"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"pass@1":null}},"counts":{"rows":8,"rows_with_code":8,"rows_with_paper_page":8,"rows_dated":8,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"o1-preview","metrics":{"pass@1":"0.952"},"uses_additional_data":false,"paper_date":"2024-09-19","paper":"/paper/a-case-study-of-web-app-coding-with-openai","paper_url":"https://arxiv.org/abs/2409.13773v1","paper_title":"A Case Study of Web App Coding with OpenAI Reasoning Models","code":"https://github.com/onekq/webapp1k","n_code_links":1,"syntology":null},{"rank_in_archive_order":2,"model":"o1-mini","metrics":{"pass@1":"0.939"},"uses_additional_data":false,"paper_date":"2024-09-19","paper":"/paper/a-case-study-of-web-app-coding-with-openai","paper_url":"https://arxiv.org/abs/2409.13773v1","paper_title":"A Case Study of Web App Coding with OpenAI Reasoning Models","code":"https://github.com/onekq/webapp1k","n_code_links":1,"syntology":null},{"rank_in_archive_order":3,"model":"gpt-4o-2024-08-06","metrics":{"pass@1":"0.885"},"uses_additional_data":false,"paper_date":"2024-09-08","paper":"/paper/insights-from-benchmarking-frontier-language","paper_url":"https://arxiv.org/abs/2409.05177v1","paper_title":"Insights from Benchmarking Frontier Language Models on Web App Code Generation","code":"https://github.com/onekq/webapp1k","n_code_links":1,"syntology":null},{"rank_in_archive_order":4,"model":"claude-3.5-sonnet","metrics":{"pass@1":"0.8808"},"uses_additional_data":false,"paper_date":"2024-09-08","paper":"/paper/insights-from-benchmarking-frontier-language","paper_url":"https://arxiv.org/abs/2409.05177v1","paper_title":"Insights from Benchmarking Frontier Language Models on Web App Code Generation","code":"https://github.com/onekq/webapp1k","n_code_links":1,"syntology":null},{"rank_in_archive_order":5,"model":"deepseek-v2.5","metrics":{"pass@1":"0.834"},"uses_additional_data":false,"paper_date":"2024-09-19","paper":"/paper/a-case-study-of-web-app-coding-with-openai","paper_url":"https://arxiv.org/abs/2409.13773v1","paper_title":"A Case Study of Web App Coding with OpenAI Reasoning Models","code":"https://github.com/onekq/webapp1k","n_code_links":1,"syntology":null},{"rank_in_archive_order":6,"model":"mistral-large-2","metrics":{"pass@1":"0.7804"},"uses_additional_data":false,"paper_date":"2024-09-08","paper":"/paper/insights-from-benchmarking-frontier-language","paper_url":"https://arxiv.org/abs/2409.05177v1","paper_title":"Insights from Benchmarking Frontier Language Models on Web App Code Generation","code":"https://github.com/onekq/webapp1k","n_code_links":1,"syntology":null},{"rank_in_archive_order":7,"model":"deepseek-coder-v2-instruct","metrics":{"pass@1":"0.7002"},"uses_additional_data":false,"paper_date":"2024-09-08","paper":"/paper/insights-from-benchmarking-frontier-language","paper_url":"https://arxiv.org/abs/2409.05177v1","paper_title":"Insights from Benchmarking Frontier Language Models on Web App Code Generation","code":"https://github.com/onekq/webapp1k","n_code_links":1,"syntology":null},{"rank_in_archive_order":8,"model":"llama-v3p1-405b-instruct","metrics":{"pass@1":"0.302"},"uses_additional_data":false,"paper_date":"2024-09-08","paper":"/paper/insights-from-benchmarking-frontier-language","paper_url":"https://arxiv.org/abs/2409.05177v1","paper_title":"Insights from Benchmarking Frontier Language Models on Web App Code Generation","code":"https://github.com/onekq/webapp1k","n_code_links":1,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":0,"rows_with_any_sample_ran":0,"distinct_papers_with_graph_line":0,"distinct_papers_with_any_sample_ran":0,"samples_over_distinct_papers":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":0,"n_unverified":0,"n_samples":0,"n_pointer_only_licence":0,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}