{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/batch-active-preference-based-learning-of","title":"Batch Active Preference-Based Learning of Reward Functions","arxiv_id":"1810.04303","date":"2018-10-10","proceeding":null,"authors":["Erdem Biyik","Dorsa Sadigh"],"abstract":"Data generation and labeling are usually an expensive part of learning for\nrobotics. While active learning methods are commonly used to tackle the former\nproblem, preference-based learning is a concept that attempts to solve the\nlatter by querying users with preference questions. In this paper, we will\ndevelop a new algorithm, batch active preference-based learning, that enables\nefficient learning of reward functions using as few data samples as possible\nwhile still having short query generation times. We introduce several\napproximations to the batch active learning problem, and provide theoretical\nguarantees for the convergence of our algorithms. Finally, we present our\nexperimental results for a variety of robotics tasks in simulation. Our results\nsuggest that our batch active learning algorithm requires only a few queries\nthat are computed in a short amount of time. We then showcase our algorithm in\na study to learn human users' preferences.","url_abs":"http://arxiv.org/abs/1810.04303v1","url_pdf":"http://arxiv.org/pdf/1810.04303v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"batch-active-preference-based-learning-of","repo_url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.04303","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1810.04303"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":7},"by_repo_kind":{"official":{"samples":7,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"7b31c8b445431412","entry":"feature","repo":"Stanford-ILIAD/batch-active-preference-based-learning","repo_kind":"official","path":"feature.py","file_url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning/blob/HEAD/feature.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7b31c8b445431412"}},{"code_sha256_prefix":"274e1610509d19a0","entry":"func","repo":"Stanford-ILIAD/batch-active-preference-based-learning","repo_kind":"official","path":"algos.py","file_url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning/blob/HEAD/algos.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"274e1610509d19a0"}},{"code_sha256_prefix":"2bcdf42f9515eb15","entry":"func_psi","repo":"Stanford-ILIAD/batch-active-preference-based-learning","repo_kind":"official","path":"algos.py","file_url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning/blob/HEAD/algos.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2bcdf42f9515eb15"}},{"code_sha256_prefix":"4497f69d2c6da5f1","entry":"generate_psi","repo":"Stanford-ILIAD/batch-active-preference-based-learning","repo_kind":"official","path":"algos.py","file_url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning/blob/HEAD/algos.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4497f69d2c6da5f1"}},{"code_sha256_prefix":"2a6e648f2f77d0f5","entry":"get_feedback","repo":"Stanford-ILIAD/batch-active-preference-based-learning","repo_kind":"official","path":"simulation_utils.py","file_url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning/blob/HEAD/simulation_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2a6e648f2f77d0f5"}},{"code_sha256_prefix":"418285ba39c6240b","entry":"kMedoids","repo":"Stanford-ILIAD/batch-active-preference-based-learning","repo_kind":"official","path":"kmedoids.py","file_url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning/blob/HEAD/kmedoids.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"418285ba39c6240b"}},{"code_sha256_prefix":"bbfd88d6d34e2002","entry":"speed","repo":"Stanford-ILIAD/batch-active-preference-based-learning","repo_kind":"official","path":"feature.py","file_url":"https://github.com/Stanford-ILIAD/batch-active-preference-based-learning/blob/HEAD/feature.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bbfd88d6d34e2002"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}