{"url":"/dataset/bigcodebench","name":"BigCodeBench","full_name":null,"description_markdown":"**BigCodeBench** is an easy-to-use benchmark for code generation with practical and challenging programming tasks¹. It aims to evaluate the true programming capabilities of large language models (LLMs) in a more realistic setting¹. The benchmark is designed for HumanEval-like function-level code generation tasks, but with much more complex instructions and diverse function calls¹.\r\n\r\nHere are some key features of BigCodeBench:\r\n- **Precise evaluation & ranking**: It provides a leaderboard for latest LLM rankings before & after rigorous evaluation¹.\r\n- **Pre-generated samples**: BigCodeBench accelerates code intelligence research by open-sourcing LLM-generated samples for various models¹.\r\n- **Execution Environment**: The execution environment in BigCodeBench is less bounded than EvalPlus to support tasks with diverse library dependencies¹.\r\n- **Test Evaluation**: BigCodeBench relies on unittest for evaluating the generated code¹.\r\n\r\n(1) GitHub - bigcode-project/bigcodebench: BigCodeBench: The Next .... https://github.com/bigcode-project/bigcodebench/.","description_withheld":null,"homepage":"https://github.com/bigcode-project/bigcodebench","introduced_date":"2024-06-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/bigcodebench-benchmarking-code-generation","title":"BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions","first_author":"Terry Yue Zhuo","url":null},"license":{"name":"Apache License 2.0","url":null},"modalities":[],"tasks":[{"name":"Code Generation","url":"/task/code-generation","datasets_with_task":"/datasets/task/code-generation"}],"languages":[],"variants":["BigCodeBench","BigCodeBench-Complete","BigCodeBench-Instruct"],"data_loaders":[{"repo":"https://github.com/bigcode-project/bigcodebench","url":"https://huggingface.co/datasets/bigcode/bigcodebench","frameworks":[]},{"repo":"https://github.com/bigcode-project/bigcodebench","url":"https://github.com/bigcode-project/bigcodebench","frameworks":[]}],"num_papers_in_archive":38,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/code-generation-on-bigcodebench-complete","task":"Code Generation","dataset_variant":"BigCodeBench-Complete","rows":2,"metrics":["Pass@1"],"first_row_in_archive_order":{"model":"GPT-4o-2024-05-13","paper":"/paper/bigcodebench-benchmarking-code-generation","metrics":{"Pass@1":"61.1"},"code_links":[{"title":"mlfoundations/Evalchemy","url":"https://github.com/mlfoundations/Evalchemy"},{"title":"bigcode-project/bigcodebench","url":"https://github.com/bigcode-project/bigcodebench"},{"title":"bigcode-project/bigcodebench-annotation","url":"https://github.com/bigcode-project/bigcodebench-annotation"},{"title":"stovecat/convcodeworld","url":"https://github.com/stovecat/convcodeworld"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/code-generation-on-bigcodebench-instruct","task":"Code Generation","dataset_variant":"BigCodeBench-Instruct","rows":2,"metrics":["Pass@1"],"first_row_in_archive_order":{"model":"GPT-4o-2024-05-13","paper":"/paper/bigcodebench-benchmarking-code-generation","metrics":{"Pass@1":"51.1"},"code_links":[{"title":"mlfoundations/Evalchemy","url":"https://github.com/mlfoundations/Evalchemy"},{"title":"bigcode-project/bigcodebench","url":"https://github.com/bigcode-project/bigcodebench"},{"title":"bigcode-project/bigcodebench-annotation","url":"https://github.com/bigcode-project/bigcodebench-annotation"},{"title":"stovecat/convcodeworld","url":"https://github.com/stovecat/convcodeworld"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/bigcodebench-benchmarking-code-generation","title":"BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions","date":"2024-06-22","rows_on_this_dataset":4,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":4,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":6,"samples_ran":4,"samples_unverified":2,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}