{"url":"/dataset/arena-hard-auto","name":"Arena-Hard-Auto","full_name":null,"description_markdown":"The **Arena-Hard-Auto** benchmark is an automatic evaluation tool for instruction-tuned Language Learning Models (LLMs)¹. It was developed to provide a cheaper and faster approximation to human preference¹.\r\n\r\nHere are some key features of the Arena-Hard-Auto benchmark:\r\n- It contains 500 challenging user queries¹.\r\n- It uses GPT-4-Turbo as a judge to compare the models' responses against a baseline model (default: GPT-4-0314)¹.\r\n- It employs an automatic judge as a cheaper and faster approximator to human preference¹.\r\n- It has the highest correlation and separability to Chatbot Arena among popular open-ended LLM benchmarks¹.\r\n- If you are curious to see how well your model might perform on Chatbot Arena, Arena-Hard-Auto is recommended¹.\r\n\r\nThe benchmark is built from live data in Chatbot Arena, which is a popular crowd-sourced platform for LLM evaluations⁴. It offers significantly stronger separability against other benchmarks with tighter confidence intervals². \r\n\r\n(1) GitHub - lm-sys/arena-hard-auto: Arena-Hard-Auto: An automatic LLM .... https://github.com/lm-sys/arena-hard-auto.\r\n(2) Arena Hard – UC Berkeley Sky Computing. https://sky.cs.berkeley.edu/project/arena-hard/.\r\n(3) From Live Data to High-Quality Benchmarks: The Arena-Hard Pipeline .... https://lmsys.org/blog/2024-04-19-arena-hard/.\r\n(4) GitHub - lm-sys/arena-hard-auto: Arena-Hard-Auto: An automatic LLM .... https://github.com/lm-sys/arena-hard-auto.\r\n(5) Arena-Hard：开源高质量大模型评估基准-CSDN博客. https://blog.csdn.net/weixin_57291105/article/details/138132998.","description_withheld":null,"homepage":"https://github.com/lm-sys/arena-hard-auto","introduced_date":"2024-06-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/from-crowdsourced-data-to-high-quality","title":"From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline","first_author":"Tianle Li","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Arena-Hard-Auto"],"data_loaders":[],"num_papers_in_archive":8,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}