{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-glu-expansion-ratios-a-study-of","title":"Exploring GLU Expansion Ratios: A Study of Structured Pruning in LLaMA-3.2 Models","arxiv_id":null,"date":"2024-12-26","proceeding":"OSF Preprints 2024 12","authors":["Pere Martra."],"abstract":"Large language models with GLU architectures are typically designed with significant expansion ratios in their MLP layers, where output dimensions are several times larger than input dimensions. While various pruning techniques have been proposed to reduce model size, the relationship between this expansion capacity and model performance has remained unexplored. This paper presents a systematic investigation of GLU expansion ratios as a key metric for model pruning, using Llama-3.2 models (1B and 3B variants) as case studies. \r\n\r\nOur findings reveal that models with an expansion ratio of 140% consistently outperform others, achieving a balance between redundancy reduction, task-specific performance, and environmental sustainability. For instance, Llama-3.2-1B at 40% pruning and Llama-3.2-3B at 10% pruning surpassed their respective baselines in multiple benchmarks, including BoolQ, IFEval, and MUSR. \r\n\r\nBeyond performance, this study underscores significant environmental benefits. Specifically, pruning the 3B model to 140% expansion with just 10% pruning achieved a remarkable 50% reduction in CO2 emissions, showcasing the potential of pruning to enhance computational efficiency while maintaining robust performance. \r\n\r\nFuture research should explore extending these findings to other GLU-based architectures, such as Mistral, Qwen, and Microsoft Phi, to validate their broader applicability. This study provides a novel perspective on sustainable AI development, effectively bridging the goals of performance optimization and environmental responsibility.","url_abs":"https://doi.org/10.31219/osf.io/qgxea","url_pdf":"https://doi.org/10.31219/osf.io/qgxea","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exploring-glu-expansion-ratios-a-study-of","repo_url":"https://github.com/peremartra/Large-Language-Model-Notebooks-Course","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"network-pruning","task_name":"Network Pruning"}],"methods":[{"method_slug":"glu","method_name":"Gated Linear Unit"},{"method_slug":"pruning","method_name":"Pruning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}