{"url":"/dataset/codefuseeval","name":"CodeFuseEval","full_name":null,"description_markdown":"CodeFuseEval is a Code Generation benchmark that combines the multi-tasking scenarios of CodeFuse Model with the benchmarks of HumanEval-x and MBPP. This benchmark is designed to evaluate the performance of models in various multi-tasking tasks, including code completion, code generation from natural language, test case generation, cross-language code translation, and code generation from Chinese commands, among others.\r\n\r\nThe evaluation of the generated codes involves compiling and running in multiple programming languages. The versions of the programming language environments and packages we use are as follows:\r\n\r\n| Dependency | Version  |\r\n| ---------- |----------|\r\n| Python     | 3.10.9   |\r\n| JDK        | 18.0.2.1 |\r\n| Node.js    | 16.14.0  |\r\n| js-md5     | 0.7.3    |\r\n| C++        | 11       |\r\n| g++        | 7.5.0    |\r\n| Boost      | 1.75.0   |\r\n| OpenSSL    | 3.0.0    |\r\n| go         | 1.18.4   |\r\n| cargo      | 1.71.1   |","description_withheld":null,"homepage":"https://github.com/codefuse-ai/codefuse-evaluation","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["CodeFuseEval"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}