{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pragmatic-feature-preferences-learning-reward","title":"Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human Input","arxiv_id":"2405.14769","date":"2024-05-23","proceeding":null,"authors":["Andi Peng","Yuying Sun","Tianmin Shu","David Abel"],"abstract":"Humans use social context to specify preferences over behaviors, i.e. their reward functions. Yet, algorithms for inferring reward models from preference data do not take this social learning view into account. Inspired by pragmatic human communication, we study how to extract fine-grained data regarding why an example is preferred that is useful for learning more accurate reward models. We propose to enrich binary preference queries to ask both (1) which features of a given example are preferable in addition to (2) comparisons between examples themselves. We derive an approach for learning from these feature-level preferences, both for cases where users specify which features are reward-relevant, and when users do not. We evaluate our approach on linear bandit settings in both vision- and language-based domains. Results support the efficiency of our approach in quickly converging to accurate rewards with fewer comparisons vs. example-only labels. Finally, we validate the real-world applicability with a behavioral experiment on a mushroom foraging task. Our findings suggest that incorporating pragmatic feature preferences is a promising approach for more efficient user-aligned reward learning.","url_abs":"https://arxiv.org/abs/2405.14769v1","url_pdf":"https://arxiv.org/pdf/2405.14769v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2405.14769","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2405.14769"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/andipeng/feature-preference","reach":{"status":"gone","observed_at":"2026-09-17","how":"tree_404+repo_404"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/jlin816/rewards-from-language","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":4,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"489c96b934a1b5a6","entry":"expand_categorical_var","repo":"jlin816/rewards-from-language","repo_kind":"found_in_text","path":"rewards/utils.py","file_url":"https://github.com/jlin816/rewards-from-language/blob/HEAD/rewards/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"489c96b934a1b5a6"}},{"code_sha256_prefix":"11d2efcd245927eb","entry":"extremal_indicators","repo":"jlin816/rewards-from-language","repo_kind":"found_in_text","path":"rewards/utils.py","file_url":"https://github.com/jlin816/rewards-from-language/blob/HEAD/rewards/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"11d2efcd245927eb"}},{"code_sha256_prefix":"0161a3d0444d6c55","entry":"validate_embedding_speaker","repo":"jlin816/rewards-from-language","repo_kind":"found_in_text","path":"rewards/trainer.py","file_url":"https://github.com/jlin816/rewards-from-language/blob/HEAD/rewards/trainer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0161a3d0444d6c55"}},{"code_sha256_prefix":"726f82ecbfc24254","entry":"validate_listener","repo":"jlin816/rewards-from-language","repo_kind":"found_in_text","path":"rewards/trainer_listener.py","file_url":"https://github.com/jlin816/rewards-from-language/blob/HEAD/rewards/trainer_listener.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"726f82ecbfc24254"}},{"code_sha256_prefix":"5597528d7319f714","entry":"collate_rewards_and_options","repo":"jlin816/rewards-from-language","repo_kind":"found_in_text","path":"rewards/data.py","file_url":"https://github.com/jlin816/rewards-from-language/blob/HEAD/rewards/data.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5597528d7319f714"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}