{"url":"/dataset/b-pref","name":"B-Pref","full_name":null,"description_markdown":"**B-Pref** is a benchmark specially designed for preference-based RL. A key challenge with such a benchmark is providing the ability to evaluate candidate algorithms quickly, which makes relying on real human input for evaluation prohibitive. At the same time, simulating human input as giving perfect preferences for the ground truth reward function is unrealistic. B-Pref alleviates this by simulating teachers with a wide array of irrationalities, and proposes metrics not solely for performance but also for robustness to these potential irrationalities.","description_withheld":null,"homepage":"https://github.com/rll-research/BPref","introduced_date":"2021-11-04","introduced_date_note":null,"introduced_by":{"paper":"/paper/b-pref-benchmarking-preference-based","title":"B-Pref: Benchmarking Preference-Based Reinforcement Learning","first_author":"Kimin Lee","url":null},"license":{"name":"MIT License","url":"https://github.com/rll-research/BPref/blob/main/LICENSE"},"modalities":[],"tasks":[],"languages":[],"variants":["B-Pref"],"data_loaders":[],"num_papers_in_archive":9,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}