{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/omg-llava-bridging-image-level-object-level","title":"OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding","arxiv_id":"2406.19389","date":"2024-06-27","proceeding":null,"authors":["Tao Zhang","Xiangtai Li","Hao Fei","Haobo Yuan","Shengqiong Wu","Shunping Ji","Chen Change Loy","Shuicheng Yan"],"abstract":"Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation and reasoning capabilities but lack pixel-level understanding and have difficulty accepting visual prompts for flexible user interaction. This paper proposes OMG-LLaVA, a new and elegant framework combining powerful pixel-level vision understanding with reasoning abilities. It can accept various visual and text prompts for flexible user interaction. Specifically, we use a universal segmentation method as the visual encoder, integrating image information, perception priors, and visual prompts into visual tokens provided to the LLM. The LLM is responsible for understanding the user's text instructions and providing text responses and pixel-level segmentation results based on the visual information. We propose perception prior embedding to better integrate perception priors with image features. OMG-LLaVA achieves image-level, object-level, and pixel-level reasoning and understanding in a single model, matching or surpassing the performance of specialized methods on multiple benchmarks. Rather than using LLM to connect each specialist, our work aims at end-to-end training on one encoder, one decoder, and one LLM. The code and model have been released for further research.","url_abs":"https://arxiv.org/abs/2406.19389v2","url_pdf":"https://arxiv.org/pdf/2406.19389v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"omg-llava-bridging-image-level-object-level","repo_url":"https://github.com/lxtgh/omg-seg","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"universal-segmentation","task_name":"Universal Segmentation"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2406.19389","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.19389"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lxtgh/omg-seg","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran":2,"unverified":2},"by_repo_kind":{"listed":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"34bdcec7a7a92c58","entry":"mask_pool","repo":"lxtgh/omg-seg","repo_kind":"listed","path":"omg_llava/omg_llava/model/omg_seg/utils.py","file_url":"https://github.com/lxtgh/omg-seg/blob/HEAD/omg_llava/omg_llava/model/omg_seg/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"34bdcec7a7a92c58"}},{"code_sha256_prefix":"c43421dd4af62490","entry":"prepare_seg_pretrain_data","repo":"lxtgh/omg-seg","repo_kind":"listed","path":"omg_llava/omg_llava/model/omg_llava.py","file_url":"https://github.com/lxtgh/omg-seg/blob/HEAD/omg_llava/omg_llava/model/omg_llava.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"c43421dd4af62490"}},{"code_sha256_prefix":"dcdc7ef233f3e1e8","entry":"build_model","repo":"lxtgh/omg-seg","repo_kind":"listed","path":"omg_llava/xtuner/apis/model.py","file_url":"https://github.com/lxtgh/omg-seg/blob/HEAD/omg_llava/xtuner/apis/model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dcdc7ef233f3e1e8"}},{"code_sha256_prefix":"5d262d036acd1009","entry":"get_center_coords","repo":"lxtgh/omg-seg","repo_kind":"listed","path":"omg_llava/omg_llava/model/omg_seg/mask2former_vid_semanticsam.py","file_url":"https://github.com/lxtgh/omg-seg/blob/HEAD/omg_llava/omg_llava/model/omg_seg/mask2former_vid_semanticsam.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"5d262d036acd1009"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}