{"url":"/dataset/llm-evaluation-scores","name":"LLM evaluation scores","full_name":"Scores given by LLM according to preassigned score","description_markdown":"Dataset is a CSV file, that contains evaluation scores given by a panel of LLMs to responses produced by other LLMs . Responses regard a forecasting task assigned to multiple LLMs. The evaluation of the individual forecasts are performed according to 9 criteria indicated in the prompt. (see for details https://arxiv.org/abs/2412.09385). \r\n\r\nData are organised as follows. Each row in the dataset represents a forecast evaluation. In columns: the forecaster number; marks for each criterium; mean of the marks; Arena score of the evaluator LLM (as per July 12, 2024).\r\n\r\nScores can be further elaborated to evaluate the ability of LLM panels to assess forecasts. For example you can compute Intraclass Correlation Coefficients (ICC) to evaluate consistency and coherence of the panel evaluation. You can also process marks according to classical statistics or ranking evaluation algorithms (e.g. Kendall distance) or formation of optimized subsets.\r\n\r\nWe also provide apps to upload the dataset, filter the data, compute ICC values, and visualize the results through heatmaps.","description_withheld":null,"homepage":"https://github.com/LeonardoErcolani/AGILab-Peer-Review/blob/main/data/merged_llm_AGI_evaluation_scores.csv","introduced_date":"2024-12-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/ai-predicts-agi-leveraging-agi-forecasting","title":"AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities","first_author":"Fabrizio Davide","url":null},"license":{"name":"CC BY-SA","url":"https://creativecommons.org/licenses/by-sa/4.0/"},"modalities":[{"name":"Tabular","url":"/datasets/modality/tabular"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["LLM evaluation scores"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}