{"url":"/dataset/olympic-2024","name":"Olympic 2024","full_name":null,"description_markdown":"Olympic 2024 is a human-annotated dataset that contains 220 high-quality instance. Each instance consists of an input tuple (user instruction, response 1, response confidence of response 1, response 2, response confidence of response 2) and an output tuple (evaluation explanation, evaluation result).\r\nThe evaluation result would be either ‘1’ or ‘2’, indicating that response 1 or response 2 is better. To ensure the quality of human annotations, we involve three experts to concurrently annotate the same data point during the annotation process.","description_withheld":null,"homepage":"https://huggingface.co/datasets/XieQJ123/Olympic-2024","introduced_date":"2025-02-15","introduced_date_note":null,"introduced_by":{"paper":"/paper/an-empirical-analysis-of-uncertainty-in-large","title":"An Empirical Analysis of Uncertainty in Large Language Model Evaluations","first_author":"Qiujie Xie","url":null},"license":{"name":"Creative Commons Attribution Non Commercial Share Alike 4.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Olympic 2024"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}