{"url":"/dataset/superbench","name":"SuperBench","full_name":null,"description_markdown":"SuperBench is a comprehensive evaluation system for large language models that includes five benchmark datasets: ExtremeGLUE for semantics, CodeBench for code, AlignBench for alignment, AgentBench for intelligent agents, and SafetyBench for safety. These benchmarks cover a wide range of tasks and dimensions to assess the overall capabilities of large language models. The SuperBench team aims to provide objective and scientific evaluation standards for large models to promote their healthy development in terms of technology, applications, and ecosystem.","description_withheld":null,"homepage":"https://fm.ai.tsinghua.edu.cn/superbench","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["SuperBench"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}