{"url":"/dataset/rewardbench","name":"RewardBench","full_name":null,"description_markdown":"**RewardBench** is a benchmark designed to evaluate the capabilities and safety of reward models, including those trained with **Direct Preference Optimization (DPO)**. It serves as the first evaluation tool for reward models and provides valuable insights into their performance and reliability¹.\r\n\r\nHere are the key components of **RewardBench**:\r\n\r\n1. **Common Inference Code**: The repository includes common inference code for various reward models, such as **Starling**, **PairRM**, **OpenAssistant**, and more. These models can be evaluated using the provided tools¹.\r\n\r\n2. **Dataset and Evaluation**: The **RewardBench dataset** consists of prompt-win-lose trios spanning chat, reasoning, and safety scenarios. It allows benchmarking reward models on challenging, structured, and out-of-distribution queries. The goal is to enhance scientific understanding of reward models and their behavior².\r\n\r\n3. **Scripts for Evaluation**:\r\n    - `scripts/run_rm.py`: Used to evaluate individual reward models.\r\n    - `scripts/run_dpo.py`: Used to evaluate direct preference optimization (DPO) models.\r\n    - `scripts/train_rm.py`: A basic reward model training script built on TRL (Transformer Reinforcement Learning)¹.\r\n\r\n4. **Installation and Usage**:\r\n    - Install **PyTorch** on your system.\r\n    - Install the required dependencies using `pip install -e .`.\r\n    - Set the environment variable `HF_TOKEN` with your token.\r\n    - To contribute your model to the leaderboard, open an issue on **HuggingFace** with the model name. For local model evaluation, follow the instructions in the repository¹.\r\n\r\nRemember that **RewardBench** provides a standardized way to assess reward models, ensuring transparency and comparability across different approaches. 🌟🔍\r\n\r\n(1) GitHub - allenai/reward-bench: RewardBench: the first evaluation tool .... https://github.com/allenai/reward-bench.\r\n(2) RewardBench: Evaluating Reward Models for Language Modeling. https://arxiv.org/abs/2403.13787.\r\n(3) RewardBench: Evaluating Reward Models for Language Modeling. https://paperswithcode.com/paper/rewardbench-evaluating-reward-models-for.","description_withheld":null,"homepage":"https://github.com/allenai/reward-bench","introduced_date":"2024-03-20","introduced_date_note":null,"introduced_by":{"paper":"/paper/rewardbench-evaluating-reward-models-for","title":"RewardBench: Evaluating Reward Models for Language Modeling","first_author":"Nathan Lambert","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["RewardBench"],"data_loaders":[],"num_papers_in_archive":105,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}