{"url":"/dataset/arena-hard","name":"Arena-Hard","full_name":null,"description_markdown":"The **Arena-Hard benchmark** is a high-quality benchmarking tool for Language Learning Models (LLMs) developed by LMSYS Org¹. It was designed to address the limitations of traditional benchmarks, which are often static or close-ended¹.\r\n\r\nKey features of the Arena-Hard benchmark include¹²:\r\n- **Robustly separates model capability**: It can differentiate the capabilities of various models.\r\n- **Reflects human preference in real-world use cases**: The benchmark score has a high agreement with human preference.\r\n- **Frequently updates to avoid over-fitting or test set leakage**: It uses new, unseen prompts to ensure the benchmark remains challenging and relevant.\r\n\r\nThe Arena-Hard benchmark is built from live data in the Chatbot Arena, a crowd-sourced platform for LLM evaluations¹. It contains 500 challenging user queries². The benchmark uses GPT-4-Turbo as a judge to compare the responses of different models against a baseline model².\r\n\r\nThe Arena-Hard benchmark has been found to offer significantly stronger separability against other benchmarks, with tighter confidence intervals¹. It also has a higher agreement (89.1%) with the human preference ranking by Chatbot Arena¹.\r\n\r\n(1) From Live Data to High-Quality Benchmarks: The Arena-Hard Pipeline .... https://lmsys.org/blog/2024-04-19-arena-hard/.\r\n(2) GitHub - lm-sys/arena-hard: Arena-Hard benchmark. https://github.com/lm-sys/arena-hard.\r\n(3) From Live Data to High-Quality Benchmarks: The Arena-Hard Pipeline .... https://lmsys.org/blog/2024-04-19-arena-hard/.\r\n(4) GitHub - lm-sys/arena-hard: Arena-Hard benchmark. https://github.com/lm-sys/arena-hard.\r\n(5) undefined. https://github.com/lm-sys/arena-hard.git.\r\n(6) undefined. https://huggingface.co/spaces/lmsys/arena-hard-browser.","description_withheld":null,"homepage":"https://github.com/lm-sys/arena-hard","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Arena-Hard"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}