{"url":"/dataset/evoeval","name":"EvoEval","full_name":null,"description_markdown":"EvoEval is a holistic benchmark suite created by evolving HumanEval problems¹. It contains 828 new problems across 5 semantic-altering and 2 semantic-preserving benchmarks¹. EvoEval allows evaluation and comparison across different dimensions and problem types, such as Difficult, Creative, or Tool Use problems¹.\r\n\r\nThe goal of EvoEval is to provide a comprehensive evaluation of Language Learning Models' (LLMs) coding abilities². It was introduced to address the limitations of existing benchmarks, which contain only a very limited set of problems, both in quantity and variety². \r\n\r\nEvoEval can be used to further evolve arbitrary problems to keep up with advances and the ever-changing landscape of LLMs for code². It comes complete with a leaderboard, ground truth solutions, robust test cases, and evaluation scripts to easily fit into your evaluation pipeline¹.\r\n\r\n(1) evo-eval/evoeval: EvoEval: Evolving Coding Benchmarks via LLM - GitHub. https://github.com/evo-eval/evoeval.\r\n(2) Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval .... https://arxiv.org/abs/2403.19114.\r\n(3) evoeval (EvoEval) - Hugging Face. https://huggingface.co/evoeval.\r\n(4) EvoEval: Evolving Coding Benchmarks via LLM. https://evo-eval.github.io/.\r\n(5) Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval .... https://paperswithcode.com/paper/top-leaderboard-ranking-top-coding.\r\n(6) undefined. https://doi.org/10.48550/arXiv.2403.19114.","description_withheld":null,"homepage":"https://evo-eval.github.io","introduced_date":"2024-03-28","introduced_date_note":null,"introduced_by":{"paper":"/paper/top-leaderboard-ranking-top-coding","title":"Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM","first_author":"Chunqiu Steven Xia","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["EvoEval"],"data_loaders":[],"num_papers_in_archive":7,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}