{"url":"/dataset/devbench","name":"DevBench","full_name":null,"description_markdown":"DevBench is a comprehensive benchmark designed to evaluate Large Language Models (LLMs) across various stages of the software development lifecycle. It covers critical steps such as software design, environment setup, implementation, acceptance testing, and unit testing. By integrating these interconnected tasks under a single framework, DevBench offers a holistic perspective on the potential of LLMs for automated software development1.\r\n\r\nHere are some key details about DevBench:\r\n\r\nPurpose: DevBench aims to assess LLMs’ capabilities in automating software development tasks.\r\nDataset: The DevBench dataset comprises 22 curated repositories across 4 programming languages (Python, C/C++, Java, JavaScript). These repositories cover diverse domains, including machine learning, databases, web services, and command-line utilities.\r\nEvaluation Suite:\r\nImplementation Task: DevBench provides extensive acceptance and unit test cases for the implementation task.\r\nSoftware Design Task: For evaluating the software design task, DevBench utilizes LLM-as-a-Judge.\r\nBaseline Agent System: DevBench includes a baseline agent system based on the popular multi-agent software development system, ChatDev.","description_withheld":null,"homepage":"https://github.com/open-compass/devbench","introduced_date":"2024-03-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/devbench-a-comprehensive-benchmark-for","title":"Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study","first_author":"Bowen Li","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["DevBench"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}