Datasets › CValues

CValues

Introduced by Guohai Xu et al. in CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility19 Jul 2023 archive 2025-07-28

CValues is a Chinese human values evaluation benchmark designed to assess the alignment of Chinese Large Language Models (LLMs) with human values. Let me provide you with more details:

  1. Purpose and Context:
  2. With the rapid evolution of large language models, there is a growing concern that they may pose risks or have negative social impacts.
  3. CValues focuses on evaluating the alignment ability of Chinese LLMs in terms of both safety and responsibility criteria.
  4. Previous work mainly assessed LLMs based on knowledge and reasoning abilities, but CValues specifically targets human values alignment, especially in a Chinese context.

  5. Data Collection:

  6. The benchmark involves manually collecting adversarial safety prompts across 10 scenarios and inducing responsibility prompts from 8 domains using input from professional experts.

  7. Evaluation Methods:

  8. Human Evaluation: Experts assess the alignment of Chinese LLMs with human values.
  9. Automatic Evaluation: Multi-choice prompts are constructed for automatic assessment.

  10. Findings:

  11. Most Chinese LLMs perform well in terms of safety.
  12. However, there is room for improvement in terms of responsibility.
  13. Both automatic and human evaluations are crucial for assessing human values alignment.

Source: Conversation with Bing, 3/18/2024 (1) [2307.09705] CValues: Measuring the Values of Chinese Large Language .... https://arxiv.org/abs/2307.09705. (2) VALUE - GitHub Pages. https://value-benchmark.github.io/. (3) Benchmarking 101: Definition, Types, Benefits and How to Use Them - Databox. https://databox.com/what-are-benchmarks. (4) Compare and Conquer: 12 Types of Benchmarking for Measuring ... - Databox. https://databox.com/benchmarking-types. (5) [2307.09705] CValues: Measuring the Values of Chinese Large Language .... https://ar5iv.labs.arxiv.org/html/2307.09705. (6) undefined. https://doi.org/10.48550/arXiv.2307.09705.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 19 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • CValues

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections