{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sphinx-the-joint-mixing-of-weights-tasks-and","title":"SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models","arxiv_id":"2311.07575","date":"2023-11-13","proceeding":null,"authors":["Ziyi Lin","Chris Liu","Renrui Zhang","Peng Gao","Longtian Qiu","Han Xiao","Han Qiu","Chen Lin","Wenqi Shao","Keqin Chen","Jiaming Han","Siyuan Huang","Yichi Zhang","Xuming He","Hongsheng Li","Yu Qiao"],"abstract":"We present SPHINX, a versatile multi-modal large language model (MLLM) with a joint mixing of model weights, tuning tasks, and visual embeddings. First, for stronger vision-language alignment, we unfreeze the large language model (LLM) during pre-training, and introduce a weight mix strategy between LLMs trained by real-world and synthetic data. By directly integrating the weights from two domains, the mixed LLM can efficiently incorporate diverse semantics with favorable robustness. Then, to enable multi-purpose capabilities, we mix a variety of tasks for joint visual instruction tuning, and design task-specific instructions to avoid inter-task conflict. In addition to the basic visual question answering, we include more challenging tasks such as region-level understanding, caption grounding, document layout detection, and human pose estimation, contributing to mutual enhancement over different scenarios. Additionally, we propose to extract comprehensive visual embeddings from various network architectures, pre-training paradigms, and information granularity, providing language models with more robust image representations. Based on our proposed joint mixing, SPHINX exhibits superior multi-modal understanding capabilities on a wide range of applications. On top of this, we further propose an efficient strategy aiming to better capture fine-grained appearances of high-resolution images. With a mixing of different scales and high-resolution sub-images, SPHINX attains exceptional visual parsing and reasoning performance on existing evaluation benchmarks. We hope our work may cast a light on the exploration of joint mixing in future MLLM research. Code is released at https://github.com/Alpha-VLLM/LLaMA2-Accessory.","url_abs":"https://arxiv.org/abs/2311.07575v1","url_pdf":"https://arxiv.org/pdf/2311.07575v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sphinx-the-joint-mixing-of-weights-tasks-and","repo_url":"https://github.com/alpha-vllm/llama2-accessory","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"described-object-detection","task_name":"Described Object Detection"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[{"method_slug":"visual-parsing","method_name":"Visual Parsing"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/described-object-detection-on-description","task":"Described Object Detection","dataset":"Description Detection Dataset","model":"SPHINX-7B","rank_in_archive_order":6,"of":8,"metrics":{"Intra-scenario ABS mAP":"7.9","Intra-scenario FULL mAP":"10.6","Intra-scenario PRES mAP":"11.4"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-benchlmm","task":"Visual Question Answering","dataset":"BenchLMM","model":"Sphinx-V2-1K","rank_in_archive_order":2,"of":10,"metrics":{"GPT-3.5 score":"57.43"},"uses_additional_data":true},{"leaderboard":"/sota/visual-question-answering-on-mm-vet","task":"Visual Question Answering","dataset":"MM-Vet","model":"SPHINX-2k","rank_in_archive_order":109,"of":231,"metrics":{"GPT-4 score":"40.2"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-vqa-on-core-mm","task":"Visual Question Answering (VQA)","dataset":"InfiMM-Eval","model":"SPHINX v2","rank_in_archive_order":2,"of":14,"metrics":{"Abductive":"49.85","Analogical":"20.69","Deductive":"42.17","Overall score":"39.48","Params":"16B"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2311.07575","atlas_url":"https://app.syntology.ai/?focus=2311.07575","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2311.07575"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/alpha-vllm/llama2-accessory","reach":null}],"summary":{"ran_draft_wrong":1,"ran_violates":1,"ran_fixture":1},"by_repo_kind":{},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"dcae39c8928b5fd8","entry":"apply_rotary_emb","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"dcae39c8928b5fd8"}},{"code_sha256_prefix":"2c4423db8989ee05","entry":"precompute_freqs_cis","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"2c4423db8989ee05"}},{"code_sha256_prefix":"70bf6ebaafd266c4","entry":"reshape_for_broadcast","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"70bf6ebaafd266c4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}