{"url":"/dataset/cvalues","name":"CValues","full_name":null,"description_markdown":"**CValues** is a **Chinese human values evaluation benchmark** designed to assess the alignment of **Chinese Large Language Models (LLMs)** with human values. Let me provide you with more details:\r\n\r\n1. **Purpose and Context**:\r\n   - With the rapid evolution of large language models, there is a growing concern that they may pose risks or have negative social impacts.\r\n   - **CValues** focuses on evaluating the alignment ability of Chinese LLMs in terms of both **safety** and **responsibility** criteria.\r\n   - Previous work mainly assessed LLMs based on knowledge and reasoning abilities, but **CValues** specifically targets human values alignment, especially in a Chinese context.\r\n\r\n2. **Data Collection**:\r\n   - The benchmark involves **manually collecting** adversarial safety prompts across **10 scenarios** and inducing responsibility prompts from **8 domains** using input from professional experts.\r\n\r\n3. **Evaluation Methods**:\r\n   - **Human Evaluation**: Experts assess the alignment of Chinese LLMs with human values.\r\n   - **Automatic Evaluation**: Multi-choice prompts are constructed for automatic assessment.\r\n\r\n4. **Findings**:\r\n   - Most Chinese LLMs perform well in terms of **safety**.\r\n   - However, there is **room for improvement** in terms of **responsibility**.\r\n   - Both automatic and human evaluations are crucial for assessing human values alignment.\r\n\r\nSource: Conversation with Bing, 3/18/2024\r\n(1) [2307.09705] CValues: Measuring the Values of Chinese Large Language .... https://arxiv.org/abs/2307.09705.\r\n(2) VALUE - GitHub Pages. https://value-benchmark.github.io/.\r\n(3) Benchmarking 101: Definition, Types, Benefits and How to Use Them - Databox. https://databox.com/what-are-benchmarks.\r\n(4) Compare and Conquer: 12 Types of Benchmarking for Measuring ... - Databox. https://databox.com/benchmarking-types.\r\n(5) [2307.09705] CValues: Measuring the Values of Chinese Large Language .... https://ar5iv.labs.arxiv.org/html/2307.09705.\r\n(6) undefined. https://doi.org/10.48550/arXiv.2307.09705.","description_withheld":null,"homepage":"https://github.com/x-plug/cvalues","introduced_date":"2023-07-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/cvalues-measuring-the-values-of-chinese-large","title":"CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility","first_author":"Guohai Xu","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["CValues"],"data_loaders":[],"num_papers_in_archive":19,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}