{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scaling-data-diversity-for-fine-tuning","title":"Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment","arxiv_id":"2403.11124","date":"2024-03-17","proceeding":null,"authors":["Feifan Song","Bowen Yu","Hao Lang","Haiyang Yu","Fei Huang","Houfeng Wang","Yongbin Li"],"abstract":"Alignment with human preference prevents large language models (LLMs) from generating misleading or toxic content while requiring high-cost human feedback. Assuming resources of human annotation are limited, there are two different ways of allocating considered: more diverse PROMPTS or more diverse RESPONSES to be labeled. Nonetheless, a straightforward comparison between their impact is absent. In this work, we first control the diversity of both sides according to the number of samples for fine-tuning, which can directly reflect their influence. We find that instead of numerous prompts, more responses but fewer prompts better trigger LLMs for human alignment. Additionally, the concept of diversity for prompts can be more complex than responses that are typically quantified by single digits. Consequently, a new formulation of prompt diversity is proposed, further implying a linear correlation with the final performance of LLMs after fine-tuning. We also leverage it on data augmentation and conduct experiments to show its effect on different algorithms.","url_abs":"https://arxiv.org/abs/2403.11124v2","url_pdf":"https://arxiv.org/pdf/2403.11124v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scaling-data-diversity-for-fine-tuning","repo_url":"https://github.com/f2-song/scalingalignment","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"diversity","task_name":"Diversity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2403.11124","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2403.11124"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/f2-song/scalingalignment","reach":{"status":"ok","spdx":"GPL-3.0"}}],"summary":{"ran":6},"by_repo_kind":{"official":{"samples":6,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"beaab6339f0946cc","entry":"generate_pipeline","repo":"f2-song/scalingalignment","repo_kind":"official","path":"eval_hh/infer_func.py","file_url":"https://github.com/f2-song/scalingalignment/blob/HEAD/eval_hh/infer_func.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"beaab6339f0946cc"}},{"code_sha256_prefix":"0d49b060dc9012dc","entry":"get_bleu","repo":"f2-song/scalingalignment","repo_kind":"official","path":"eval_hh/metrics.py","file_url":"https://github.com/f2-song/scalingalignment/blob/HEAD/eval_hh/metrics.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"0d49b060dc9012dc"}},{"code_sha256_prefix":"6ca582890a040098","entry":"get_info","repo":"f2-song/scalingalignment","repo_kind":"official","path":"data_preprocess/diversity.py","file_url":"https://github.com/f2-song/scalingalignment/blob/HEAD/data_preprocess/diversity.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"6ca582890a040098"}},{"code_sha256_prefix":"7735b404c6a4f365","entry":"get_ngrams","repo":"f2-song/scalingalignment","repo_kind":"official","path":"data_preprocess/diversity.py","file_url":"https://github.com/f2-song/scalingalignment/blob/HEAD/data_preprocess/diversity.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"7735b404c6a4f365"}},{"code_sha256_prefix":"814d25f588e5f6e9","entry":"load_raw_dataset","repo":"f2-song/scalingalignment","repo_kind":"official","path":"data_preprocess/split.py","file_url":"https://github.com/f2-song/scalingalignment/blob/HEAD/data_preprocess/split.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"814d25f588e5f6e9"}},{"code_sha256_prefix":"25009451201ba931","entry":"produce_tokenized_prompts","repo":"f2-song/scalingalignment","repo_kind":"official","path":"data_preprocess/diversity.py","file_url":"https://github.com/f2-song/scalingalignment/blob/HEAD/data_preprocess/diversity.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"25009451201ba931"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}