{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fine-tuning-large-language-models-for-domain-1","title":"Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities","arxiv_id":"2409.03444","date":"2024-09-05","proceeding":null,"authors":["Wei Lu","Rachel K. Luu","Markus J. Buehler"],"abstract":"The advancement of Large Language Models (LLMs) for domain applications in fields such as materials science and engineering depends on the development of fine-tuning strategies that adapt models for specialized, technical capabilities. In this work, we explore the effects of Continued Pretraining (CPT), Supervised Fine-Tuning (SFT), and various preference-based optimization approaches, including Direct Preference Optimization (DPO) and Odds Ratio Preference Optimization (ORPO), on fine-tuned LLM performance. Our analysis shows how these strategies influence model outcomes and reveals that the merging of multiple fine-tuned models can lead to the emergence of capabilities that surpass the individual contributions of the parent models. We find that model merging leads to new functionalities that neither parent model could achieve alone, leading to improved performance in domain-specific assessments. Experiments with different model architectures are presented, including Llama 3.1 8B and Mistral 7B models, where similar behaviors are observed. Exploring whether the results hold also for much smaller models, we use a tiny LLM with 1.7 billion parameters and show that very small LLMs do not necessarily feature emergent capabilities under model merging, suggesting that model scaling may be a key component. In open-ended yet consistent chat conversations between a human and AI models, our assessment reveals detailed insights into how different model variants perform and show that the smallest model achieves a high intelligence score across key criteria including reasoning depth, creativity, clarity, and quantitative precision. Other experiments include the development of image generation prompts based on disparate biological material design concepts, to create new microstructures, architectural concepts, and urban design based on biological materials-inspired construction principles.","url_abs":"https://arxiv.org/abs/2409.03444v1","url_pdf":"https://arxiv.org/pdf/2409.03444v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fine-tuning-large-language-models-for-domain-1","repo_url":"https://github.com/lamm-mit/llm-finetuning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"fine-tuning-large-language-models-for-domain-1","repo_url":"https://huggingface.co/lamm-mit/Llama3.1-8b-Instruct-CPT-SFT-ORPO-SLERP-09022024","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"fine-tuning-large-language-models-for-domain-1","repo_url":"https://huggingface.co/lamm-mit/SmolLM-Base-1.7B-CPT-SFT-DPO-09022024","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"fine-tuning-large-language-models-for-domain-1","repo_url":"https://huggingface.co/lamm-mit/leaf-FLUX.1-dev","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"fine-tuning-large-language-models-for-domain-1","repo_url":"https://huggingface.co/lamm-mit/leaf-L-FLUX.1-dev","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"fine-tuning-large-language-models-for-domain-1","repo_url":"https://huggingface.co/lamm-mit/mistral-7B-v0.3-Base-CPT-SFT-DPO-09022024","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"fine-tuning-large-language-models-for-domain-1","repo_url":"https://huggingface.co/lamm-mit/mistral-7B-v0.3-Base-CPT-SFT-SLERP-09022024","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"domain-adaptation","task_name":"Domain Adaptation"},{"task_slug":"image-generation","task_name":"Image Generation"}],"methods":[{"method_slug":"llama","method_name":"LLaMA"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2409.03444","atlas_url":"https://app.syntology.ai/?focus=2409.03444","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2409.03444"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://huggingface.co/lamm-mit/leaf-L-FLUX.1-dev","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://huggingface.co/lamm-mit/mistral-7B-v0.3-Base-CPT-SFT-SLERP-09022024","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://huggingface.co/lamm-mit/leaf-FLUX.1-dev","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://huggingface.co/lamm-mit/mistral-7B-v0.3-Base-CPT-SFT-DPO-09022024","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lamm-mit/llm-finetuning","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://huggingface.co/lamm-mit/Llama3.1-8b-Instruct-CPT-SFT-ORPO-SLERP-09022024","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://huggingface.co/lamm-mit/SmolLM-Base-1.7B-CPT-SFT-DPO-09022024","reach":null}],"summary":{"ran":4,"ran_violates":1},"by_repo_kind":{"official":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"9a40c8da283bcd4a","entry":"apply_chat_template","repo":"lamm-mit/llm-finetuning","repo_kind":"official","path":"alignment-handbook/src/alignment/data.py","file_url":"https://github.com/lamm-mit/llm-finetuning/blob/HEAD/alignment-handbook/src/alignment/data.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9a40c8da283bcd4a"}},{"code_sha256_prefix":"0edeacd7a9034f53","entry":"extract_docstring","repo":"lamm-mit/llm-finetuning","repo_kind":"official","path":"alignment-handbook/src/alignment/decontaminate.py","file_url":"https://github.com/lamm-mit/llm-finetuning/blob/HEAD/alignment-handbook/src/alignment/decontaminate.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0edeacd7a9034f53"}},{"code_sha256_prefix":"4253aac010c4504f","entry":"is_openai_format","repo":"lamm-mit/llm-finetuning","repo_kind":"official","path":"alignment-handbook/src/alignment/data.py","file_url":"https://github.com/lamm-mit/llm-finetuning/blob/HEAD/alignment-handbook/src/alignment/data.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4253aac010c4504f"}},{"code_sha256_prefix":"a3791c5c2185e013","entry":"load_dataset_column","repo":"lamm-mit/llm-finetuning","repo_kind":"official","path":"alignment-handbook/src/alignment/decontaminate.py","file_url":"https://github.com/lamm-mit/llm-finetuning/blob/HEAD/alignment-handbook/src/alignment/decontaminate.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a3791c5c2185e013"}},{"code_sha256_prefix":"f43e59d239019e86","entry":"normalize_whitespace","repo":"lamm-mit/llm-finetuning","repo_kind":"official","path":"alignment-handbook/src/alignment/decontaminate.py","file_url":"https://github.com/lamm-mit/llm-finetuning/blob/HEAD/alignment-handbook/src/alignment/decontaminate.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f43e59d239019e86"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}