{"url":"/dataset/webgen-bench","name":"WebGen-Bench","full_name":null,"description_markdown":"# WebGen-Bench\r\n\r\nWebGen-Bench is created to benchmark LLM-based agent's ability to generate websites from scratch. The dataset is introduced in [WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch](https://arxiv.org/abs/2505.03733). It contains 101 instructions and 647 test cases. It also has a training set of 6667 instructions, named WebGen-Instruct.\r\n\r\nThe code for evaluation as well as the training code and data are released at [WebGen-Bench (Github)](https://github.com/mnluzimu/WebGen-Bench)\r\n\r\n## Citation\r\n\r\nIf you find our project useful, please cite:\r\n\r\n```\r\n@misc{lu2025webgenbenchevaluatingllmsgenerating,\r\n      title={WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch}, \r\n      author={Zimu Lu and Yunqiao Yang and Houxing Ren and Haotian Hou and Han Xiao and Ke Wang and Weikang Shi and Aojun Zhou and Mingjie Zhan and Hongsheng Li},\r\n      year={2025},\r\n      eprint={2505.03733},\r\n      archivePrefix={arXiv},\r\n      primaryClass={cs.CL},\r\n      url={https://arxiv.org/abs/2505.03733}, \r\n}\r\n```","description_withheld":null,"homepage":"https://huggingface.co/datasets/luzimu/WebGen-Bench","introduced_date":"2025-05-06","introduced_date_note":null,"introduced_by":{"paper":"/paper/webgen-bench-evaluating-llms-on-generating","title":"WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch","first_author":"Zimu Lu","url":null},"license":{"name":"MIT","url":"https://mit-license.org/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Code Generation","url":"/task/code-generation","datasets_with_task":"/datasets/task/code-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["WebGen-Bench"],"data_loaders":[{"repo":"https://github.com/mnluzimu/webgen-bench","url":"https://github.com/mnluzimu/webgen-bench","frameworks":["pytorch"]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}