{"url":"/dataset/envbench","name":"EnvBench","full_name":null,"description_markdown":"EnvBench is a comprehensive benchmark for **automating environment setup** - an important task in software engineering. We have collected the largest dataset to date for this task and introduced a robust framework for developing and evaluating LLM-based agents that tackle environment setup challenges.\r\n\r\nOur benchmark includes:\r\n\r\n* **994 repositories**: 329 Python and 665 JVM-based (Java, Kotlin) projects\r\n* **Genuine configuration challenges**: Carefully selected repositories that cannot be configured with simple deterministic scripts\r\n* **Evaluation metrics**: Static analysis for missing imports in Python and compilation checks for JVM languages\r\n* **Baselines**: Zero-shot baselines and agentic workflows tested with GPT-4o and GPT-4o-mini","description_withheld":null,"homepage":"https://huggingface.co/collections/JetBrains-Research/envbench-67d40fc30023c728e49427ad","introduced_date":"2025-03-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/envbench-a-benchmark-for-automated","title":"EnvBench: A Benchmark for Automated Environment Setup","first_author":"Aleksandra Eliseeva","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["EnvBench"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}