{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2602-10635","title":"OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization","arxiv_id":"2602.10635","date":"2026-02-11","proceeding":"ICML","authors":["Keane Ong","Sabri Boughorbel","Luwei Xiao","Chanakya Ekbote","Wei Dai","Ao Qu","Jingyao Wu","Rui Mao","Ehsan Hoque","Erik Cambria","Gianmarco Mengaldo","Paul Pu Liang"],"abstract":"Socially intelligent AI systems must reason across diverse human behavioral tasks and generalize to new social contexts. However, behavioral data is inherently heterogeneous, comprising diverse modalities and prediction targets that produce uneven training signals across samples, creating imbalanced learning dynamics that challenge existing AI models. To address this, we develop Omnisapiens-7B 2.0, a foundation model for social behavior processing that explicitly addresses learning from heterogeneous behavioral data. This is enabled through Heterogeneity-Aware Relative Policy Optimization, a new RL method that rebalances learning signals across samples by approximating each sample's contribution to the policy update and using these estimates to drive geometrically centered, inertially smoothed advantage modulation for stable training. Omnisapiens-7B 2.0 achieves the best and most consistent performance across 10 behavioral tasks, while also attaining the best performance on all five held-out benchmarks, with gains of up to +12.02% and +9.37% respectively. Furthermore, it demonstrates more consistent and interpretable reasoning traces, supporting reliable real-world behavioral applications. Our model is available at https://github.com/MIT-MI/human_behavior_atlas.","url_abs":"https://arxiv.org/abs/2602.10635","url_pdf":"https://arxiv.org/pdf/2602.10635","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2602.10635","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2602.10635"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/MIT-MI/human_behavior_atlas","reach":null}],"summary":{"unverified":12},"by_repo_kind":{"found_in_text":{"samples":12,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":12,"samples":[{"code_sha256_prefix":"5caf0cb66b1b555f","entry":"accuracy_reward","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/rl/reward_function/human_behaviour.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/rl/reward_function/human_behaviour.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5caf0cb66b1b555f"}},{"code_sha256_prefix":"f51f42f2c57d4d52","entry":"accuracy_reward","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/rl/reward_function/human_behaviour_harpo.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/rl/reward_function/human_behaviour_harpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f51f42f2c57d4d52"}},{"code_sha256_prefix":"a48b0946174e6730","entry":"build_domain_specs_from_labelscheme","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/sft/models/qwen2_5_omni_classifier_heads_decoder.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/sft/models/qwen2_5_omni_classifier_heads_decoder.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a48b0946174e6730"}},{"code_sha256_prefix":"c732f73509ec16cb","entry":"build_video_feat_single","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/sft/models/bam_utils.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/sft/models/bam_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c732f73509ec16cb"}},{"code_sha256_prefix":"5c942845e5ed2334","entry":"create_message","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"evaluation/llm_grader/llm_judge_eval.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/evaluation/llm_grader/llm_judge_eval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5c942845e5ed2334"}},{"code_sha256_prefix":"478ec7c6a2abcfff","entry":"extract_boxed_content","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/rl/reward_function/human_behaviour.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/rl/reward_function/human_behaviour.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"478ec7c6a2abcfff"}},{"code_sha256_prefix":"48ce7f7eaee7ec01","entry":"extract_boxed_content","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/rl/reward_function/human_behaviour_harpo.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/rl/reward_function/human_behaviour_harpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"48ce7f7eaee7ec01"}},{"code_sha256_prefix":"fd2112df3f149ab0","entry":"format_reward","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/rl/reward_function/human_behaviour.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/rl/reward_function/human_behaviour.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fd2112df3f149ab0"}},{"code_sha256_prefix":"b55511b458d41522","entry":"format_reward","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/rl/reward_function/human_behaviour_harpo.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/rl/reward_function/human_behaviour_harpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b55511b458d41522"}},{"code_sha256_prefix":"7380e01e3cd064c2","entry":"load_model","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/sft/train_sft.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/sft/train_sft.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7380e01e3cd064c2"}},{"code_sha256_prefix":"0bae42dc8a08d626","entry":"openpose_dict_to_framewise","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/sft/models/bam_utils.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/sft/models/bam_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0bae42dc8a08d626"}},{"code_sha256_prefix":"f142372d6c7bd301","entry":"pool_temporal","repo":"MIT-MI/human_behavior_atlas","repo_kind":"found_in_text","path":"training/sft/models/bam_utils.py","file_url":"https://github.com/MIT-MI/human_behavior_atlas/blob/HEAD/training/sft/models/bam_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f142372d6c7bd301"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}