{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2602-08835","title":"Learning the Value Systems of Societies with Preference-based Multi-objective Reinforcement Learning","arxiv_id":"2602.08835","date":"2026-02-09","proceeding":null,"authors":["Andrés Holgado-Sánchez","Peter Vamplew","Richard Dazeley","Sascha Ossowski","Holger Billhardt"],"abstract":"Value-aware AI should recognise human values and adapt to the value systems (value-based preferences) of different users. This requires operationalization of values, which can be prone to misspecification. The social nature of values demands their representation to adhere to multiple users while value systems are diverse, yet exhibit patterns among groups. In sequential decision making, efforts have been made towards personalization for different goals or values from demonstrations of diverse agents. However, these approaches demand manually designed features or lack value-based interpretability and/or adaptability to diverse user preferences. We propose algorithms for learning models of value alignment and value systems for a society of agents in Markov Decision Processes (MDPs), based on clustering and preference-based multi-objective reinforcement learning (PbMORL). We jointly learn socially-derived value alignment models (groundings) and a set of value systems that concisely represent different groups of users (clusters) in a society. Each cluster consists of a value system representing the value-based preferences of its members and an approximately Pareto-optimal policy that reflects behaviours aligned with this value system. We evaluate our method against a state-of-the-art PbMORL algorithm and baselines on two MDPs with human values.","url_abs":"https://arxiv.org/abs/2602.08835","url_pdf":"https://arxiv.org/pdf/2602.08835","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2602.08835","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2602.08835"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/andresh26-uam/ValueLearningInMOMDP","reach":null}],"summary":{"unverified":7},"by_repo_kind":{"found_in_text":{"samples":7,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e62a74f3a6f1c907","entry":"convert","repo":"andresh26-uam/ValueLearningInMOMDP","repo_kind":"found_in_text","path":"src/utils.py","file_url":"https://github.com/andresh26-uam/ValueLearningInMOMDP/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e62a74f3a6f1c907"}},{"code_sha256_prefix":"1295b6cac1b37be5","entry":"deconvert","repo":"andresh26-uam/ValueLearningInMOMDP","repo_kind":"found_in_text","path":"src/utils.py","file_url":"https://github.com/andresh26-uam/ValueLearningInMOMDP/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1295b6cac1b37be5"}},{"code_sha256_prefix":"7008f4604ba81757","entry":"most_recent_indices_to_ptr","repo":"andresh26-uam/ValueLearningInMOMDP","repo_kind":"found_in_text","path":"src/utils.py","file_url":"https://github.com/andresh26-uam/ValueLearningInMOMDP/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7008f4604ba81757"}},{"code_sha256_prefix":"9768bf7a2b1d1ef8","entry":"one_hot_encoding","repo":"andresh26-uam/ValueLearningInMOMDP","repo_kind":"found_in_text","path":"src/feature_extractors.py","file_url":"https://github.com/andresh26-uam/ValueLearningInMOMDP/blob/HEAD/src/feature_extractors.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9768bf7a2b1d1ef8"}},{"code_sha256_prefix":"462085f6717d61eb","entry":"one_hot_encoding_space","repo":"andresh26-uam/ValueLearningInMOMDP","repo_kind":"found_in_text","path":"src/feature_extractors.py","file_url":"https://github.com/andresh26-uam/ValueLearningInMOMDP/blob/HEAD/src/feature_extractors.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"462085f6717d61eb"}},{"code_sha256_prefix":"9f4cc5f33ee45036","entry":"select_output_shape","repo":"andresh26-uam/ValueLearningInMOMDP","repo_kind":"found_in_text","path":"utils.py","file_url":"https://github.com/andresh26-uam/ValueLearningInMOMDP/blob/HEAD/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9f4cc5f33ee45036"}},{"code_sha256_prefix":"d8f6f2287fc59e51","entry":"transform_weights_to_tuple","repo":"andresh26-uam/ValueLearningInMOMDP","repo_kind":"found_in_text","path":"defines.py","file_url":"https://github.com/andresh26-uam/ValueLearningInMOMDP/blob/HEAD/defines.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d8f6f2287fc59e51"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9581,"papers_checked":6264},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}