{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/embedllm-learning-compact-representations-of","title":"EmbedLLM: Learning Compact Representations of Large Language Models","arxiv_id":"2410.02223","date":"2024-10-03","proceeding":null,"authors":["Richard Zhuang","Tianhao Wu","Zhaojin Wen","Andrew Li","Jiantao Jiao","Kannan Ramchandran"],"abstract":"With hundreds of thousands of language models available on Huggingface today, efficiently evaluating and utilizing these models across various downstream, tasks has become increasingly critical. Many existing methods repeatedly learn task-specific representations of Large Language Models (LLMs), which leads to inefficiencies in both time and computational resources. To address this, we propose EmbedLLM, a framework designed to learn compact vector representations, of LLMs that facilitate downstream applications involving many models, such as model routing. We introduce an encoder-decoder approach for learning such embeddings, along with a systematic framework to evaluate their effectiveness. Empirical results show that EmbedLLM outperforms prior methods in model routing both in accuracy and latency. Additionally, we demonstrate that our method can forecast a model's performance on multiple benchmarks, without incurring additional inference cost. Extensive probing experiments validate that the learned embeddings capture key model characteristics, e.g. whether the model is specialized for coding tasks, even without being explicitly trained on them. We open source our dataset, code and embedder to facilitate further research and application.","url_abs":"https://arxiv.org/abs/2410.02223v2","url_pdf":"https://arxiv.org/pdf/2410.02223v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"embedllm-learning-compact-representations-of","repo_url":"https://github.com/richardzhuang0412/embedllm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.02223","atlas_url":"https://app.syntology.ai/?focus=2410.02223","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.02223"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/richardzhuang0412/embedllm","reach":null}],"summary":{"ran_honours":2,"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"1057fe3ecfefaf71","entry":"correctness_prediction_evaluator","repo":"richardzhuang0412/embedllm","repo_kind":"official","path":"algorithm/mf.py","file_url":"https://github.com/richardzhuang0412/embedllm/blob/HEAD/algorithm/mf.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1057fe3ecfefaf71"}},{"code_sha256_prefix":"2adf956b4cc18778","entry":"evaluate","repo":"richardzhuang0412/embedllm","repo_kind":"official","path":"algorithm/mf.py","file_url":"https://github.com/richardzhuang0412/embedllm/blob/HEAD/algorithm/mf.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2adf956b4cc18778"}},{"code_sha256_prefix":"dd905b8c0143b74e","entry":"load_tensor_data","repo":"richardzhuang0412/embedllm","repo_kind":"official","path":"algorithm/knn.py","file_url":"https://github.com/richardzhuang0412/embedllm/blob/HEAD/algorithm/knn.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dd905b8c0143b74e"}},{"code_sha256_prefix":"b953000852e227f8","entry":"load_and_process_data","repo":"richardzhuang0412/embedllm","repo_kind":"official","path":"algorithm/mf.py","file_url":"https://github.com/richardzhuang0412/embedllm/blob/HEAD/algorithm/mf.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b953000852e227f8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}