Papers › Web-Bench: A LLM Code Benchmark Based on Web Standards and Frameworks
Web-Bench: A LLM Code Benchmark Based on Web Standards and Frameworks
Kai Xu, YiWei Mao, Xinyi Guan, ZiLong Feng
The application of large language models (LLMs) in the field of coding is evolving rapidly: from code assistants, to autonomous coding agents, and then to generating complete projects through natural language. Early LLM code benchmarks primarily focused on code generation accuracy, but these benchmarks have gradually become saturated. Benchmark saturation weakens their guiding role for LLMs. For example, HumanEval Pass@1 has reached 99.4% and MBPP 94.2%. Among various attempts to address benchmark saturation, approaches based on software engineering have stood out, but the saturation of existing software engineering benchmarks is rapidly increasing. To address this, we propose a new benchmark, Web-Bench, which contains 50 projects, each consisting of 20 tasks with sequential dependencies. The tasks implement project features in sequence, simulating real-world human development workflows. When designing Web-Bench, we aim to cover the foundational elements of Web development: Web Standards and Web Frameworks. Given the scale and complexity of these projects, which were designed by engineers with 5 to 10 years of experience, each presents a significant challenge. On average, a single project takes 4 to 8 hours for a senior engineer to complete. On our given benchmark agent (Web-Agent), SOTA (Claude 3.7 Sonnet) achieves only 25.1% Pass@1, significantly lower (better) than SWE-Bench's Verified (65.4%) and Full (33.8%) scores. Finally, we discuss that in any development field, Standards and Frameworks represent foundational knowledge and efficiency tools, respectively, and LLMs require optimization tailored to them.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Web-Bench | claude-3-7-sonnet-20250219-thinking | Pass@1 | 25.11% | #1 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-7-sonnet-20250219-thinking | Pass@2 | 35.33% | #1 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-5-sonnet-20241022 | Pass@1 | 23.04% | #2 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-5-sonnet-20241022 | Pass@2 | 32.39% | #2 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-7-sonnet-20250219 | Pass@1 | 22.50% | #3 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-7-sonnet-20250219 | Pass@2 | 30.98% | #3 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-5-sonnet-20240620 | Pass@1 | 21.96% | #4 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-5-sonnet-20240620 | Pass@2 | 30.33% | #4 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4.1 | Pass@1 | 21.09% | #5 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4.1 | Pass@2 | 25.11% | #5 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4.1-mini | Pass@1 | 20.76% | #6 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4.1-mini | Pass@2 | 23.70% | #6 of 40 | Archive leaderboard | report | |
| Web-Bench | doubao-pro-1.5-thinking | Pass@1 | 20.11% | #7 of 40 | Archive leaderboard | report | |
| Web-Bench | doubao-pro-1.5-thinking | Pass@2 | 30.22% | #7 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4o | Pass@1 | 17.17% | #8 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4o | Pass@2 | 23.80% | #8 of 40 | Archive leaderboard | report | |
| Web-Bench | deepseek-v3-0324 | Pass@1 | 17.07% | #9 of 40 | Archive leaderboard | report | |
| Web-Bench | deepseek-v3-0324 | Pass@2 | 23.59% | #9 of 40 | Archive leaderboard | report | |
| Web-Bench | deepseek-coder-v2 | Pass@1 | 16.74% | #10 of 40 | Archive leaderboard | report | |
| Web-Bench | deepseek-coder-v2 | Pass@2 | 23.15% | #10 of 40 | Archive leaderboard | report | |
| Web-Bench | doubao-pro-1.5-32k | Pass@1 | 16.63% | #11 of 40 | Archive leaderboard | report | |
| Web-Bench | doubao-pro-1.5-32k | Pass@2 | 22.93% | #11 of 40 | Archive leaderboard | report | |
| Web-Bench | llama-4 Maverick | Pass@1 | 15.98% | #12 of 40 | Archive leaderboard | report | |
| Web-Bench | llama-4 Maverick | Pass@2 | 20.87% | #12 of 40 | Archive leaderboard | report | |
| Web-Bench | qwen-max-2025-01-25 | Pass@1 | 15.87% | #13 of 40 | Archive leaderboard | report | |
| Web-Bench | qwen-max-2025-01-25 | Pass@2 | 19.02% | #13 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-2.5-pro | Pass@1 | 15.67% | #14 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-2.5-pro | Pass@2 | 24.02% | #14 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-5-haiku-20241022 | Pass@1 | 15.43% | #15 of 40 | Archive leaderboard | report | |
| Web-Bench | claude-3-5-haiku-20241022 | Pass@2 | 21.74% | #15 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-2.0-flash | Pass@1 | 15.33% | #16 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-2.0-flash | Pass@2 | 20.87% | #16 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-2.0-flash-thinking | Pass@1 | 14.89% | #17 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-2.0-flash-thinking | Pass@2 | 19.24% | #17 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-pro-1.5 | Pass@1 | 14.78% | #18 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-pro-1.5 | Pass@2 | 20.87% | #18 of 40 | Archive leaderboard | report | |
| Web-Bench | deepseek-r1 | Pass@1 | 14.46% | #19 of 40 | Archive leaderboard | report | |
| Web-Bench | deepseek-r1 | Pass@2 | 26.20% | #19 of 40 | Archive leaderboard | report | |
| Web-Bench | step-fun-2-16k | Pass@1 | 13.70% | #20 of 40 | Archive leaderboard | report | |
| Web-Bench | step-fun-2-16k | Pass@2 | 15.87% | #20 of 40 | Archive leaderboard | report | |
| Web-Bench | o4-mini | Pass@1 | 13.26% | #21 of 40 | Archive leaderboard | report | |
| Web-Bench | o4-mini | Pass@2 | 22.93% | #21 of 40 | Archive leaderboard | report | |
| Web-Bench | mistral-large-2411 | Pass@1 | 13.04% | #22 of 40 | Archive leaderboard | report | |
| Web-Bench | mistral-large-2411 | Pass@2 | 18.70% | #22 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-flash-1.5 | Pass@1 | 12.83% | #23 of 40 | Archive leaderboard | report | |
| Web-Bench | gemini-flash-1.5 | Pass@2 | 17.07% | #23 of 40 | Archive leaderboard | report | |
| Web-Bench | qwen-plus-2025-01-25 | Pass@1 | 11.85% | #24 of 40 | Archive leaderboard | report | |
| Web-Bench | qwen-plus-2025-01-25 | Pass@2 | 15.11% | #24 of 40 | Archive leaderboard | report | |
| Web-Bench | grok-2-1212 | Pass@1 | 11.30% | #25 of 40 | Archive leaderboard | report | |
| Web-Bench | grok-2-1212 | Pass@2 | 17.17% | #25 of 40 | Archive leaderboard | report | |
| Web-Bench | qwen-2.5-72b-instruct | Pass@1 | 10.54% | #26 of 40 | Archive leaderboard | report | |
| Web-Bench | qwen-2.5-72b-instruct | Pass@2 | 13.70% | #26 of 40 | Archive leaderboard | report | |
| Web-Bench | o1 | Pass@1 | 10.43% | #27 of 40 | Archive leaderboard | report | |
| Web-Bench | o1 | Pass@2 | 12.39% | #27 of 40 | Archive leaderboard | report | |
| Web-Bench | gemma-3-27b | Pass@1 | 9.89% | #28 of 40 | Archive leaderboard | report | |
| Web-Bench | gemma-3-27b | Pass@2 | 11.85% | #28 of 40 | Archive leaderboard | report | |
| Web-Bench | o3-mini | Pass@1 | 9.13% | #29 of 40 | Archive leaderboard | report | |
| Web-Bench | o3-mini | Pass@2 | 14.24% | #29 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4o-mini | Pass@1 | 8.48% | #30 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4o-mini | Pass@2 | 13.04% | #30 of 40 | Archive leaderboard | report | |
| Web-Bench | sense-chat-5 | Pass@1 | 8.48% | #31 of 40 | Archive leaderboard | report | |
| Web-Bench | sense-chat-5 | Pass@2 | 12.72% | #31 of 40 | Archive leaderboard | report | |
| Web-Bench | minimax-text | Pass@1 | 8.48% | #32 of 40 | Archive leaderboard | report | |
| Web-Bench | minimax-text | Pass@2 | 10.76% | #32 of 40 | Archive leaderboard | report | |
| Web-Bench | 360-gpt2-o1 | Pass@1 | 8.26% | #33 of 40 | Archive leaderboard | report | |
| Web-Bench | 360-gpt2-o1 | Pass@2 | 14.46% | #33 of 40 | Archive leaderboard | report | |
| Web-Bench | GLM-4-0414 | Pass@1 | 7.50% | #34 of 40 | Archive leaderboard | report | |
| Web-Bench | GLM-4-0414 | Pass@2 | 9.02% | #34 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4.1-nano | Pass@1 | 7.07% | #35 of 40 | Archive leaderboard | report | |
| Web-Bench | gpt-4.1-nano | Pass@2 | 12.28% | #35 of 40 | Archive leaderboard | report | |
| Web-Bench | llama-3.3 | Pass@1 | 6.63% | #36 of 40 | Archive leaderboard | report | |
| Web-Bench | llama-3.3 | Pass@2 | 9.57% | #36 of 40 | Archive leaderboard | report | |
| Web-Bench | moonshot-kimi-latest | Pass@1 | 5.22% | #37 of 40 | Archive leaderboard | report | |
| Web-Bench | moonshot-kimi-latest | Pass@2 | 11.85% | #37 of 40 | Archive leaderboard | report | |
| Web-Bench | llama-4 Scout | Pass@1 | 5.00% | #38 of 40 | Archive leaderboard | report | |
| Web-Bench | llama-4 Scout | Pass@2 | 7.72% | #38 of 40 | Archive leaderboard | report | |
| Web-Bench | doubao-pro-1.5-32k-lite | Pass@1 | 3.48% | #39 of 40 | Archive leaderboard | report | |
| Web-Bench | doubao-pro-1.5-32k-lite | Pass@2 | 5.98% | #39 of 40 | Archive leaderboard | report | |
| Web-Bench | qwen-turbo-2024-11-01 | Pass@1 | 2.61% | #40 of 40 | Archive leaderboard | report | |
| Web-Bench | qwen-turbo-2024-11-01 | Pass@2 | 5.11% | #40 of 40 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections