{"url":"/dataset/z-bench","name":"Z-Bench","full_name":null,"description_markdown":"**Z-Bench** is a fascinating **Chinese language model prompt dataset** developed by an enthusiastic AI-focused team at **Zhenfund**. Let me share some intriguing details about it:\r\n\r\n1. **Purpose and Origin**:\r\n   - Z-Bench was created to **qualitatively test large language models** (LLMs) for non-technical users, particularly those similar to ChatGPT products.\r\n   - The team behind Z-Bench recognized that existing NLP task datasets had limitations, such as being unsuitable for dialogue systems or lacking good Chinese versions.\r\n   - Their goal was to provide a **practical and user-friendly benchmark** for assessing the capabilities of LLMs.\r\n\r\n2. **Dataset Overview**:\r\n   - Z-Bench v1.0 covers three dimensions:\r\n     - **Basic Abilities**: Includes 100 prompts.\r\n     - **Advanced Abilities**: Contains 100 prompts.\r\n     - **Vertical Abilities**: Provides 100 prompts.\r\n   - In total, there are **300 prompts** designed to cover a wide range of NLP tasks.\r\n   - The intention is not to create an academically rigorous dataset but rather to offer a useful tool for non-technical users.\r\n\r\nSource: Conversation with Bing, 3/19/2024\r\n(1) Z-Bench 1.0 by 真格基金 - GitHub. https://github.com/zhenbench/z-bench.\r\n(2) ZhenBench · GitHub. https://github.com/zhenbench/.\r\n(3) Z.bench as an estimate of sigma capability - Minitab. https://support.minitab.com/minitab/21/help-and-how-to/quality-and-process-improvement/capability-analysis/supporting-topics/capability-metrics/z-bench-as-an-estimate-of-sigma-capability/.\r\n(4) undefined. https://docs.qq.com/sheet/DTEFsdkNERVVtR3BX.","description_withheld":null,"homepage":"https://github.com/zhenbench/z-bench","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["Z-Bench"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}