{"url":"/method/llama","slug":"llama","name":"LLaMA","full_name":"LLaMA","full_name_withheld":false,"description_markdown":"**LLaMA** is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were subsequently proposed. The main difference with the original architecture are listed below.\r\n\r\n- RMSNorm normalizing function is used to improve the training stability, by normalizing the input of each transformer sub-layer, instead of normalizing the output.\r\n- The ReLU non-linearity is replaced by the SwiGLU activation function to improve performance.\r\n- Absolute positional embeddings are removed and instead rotary positional embeddings (RoPE) are added at each layer of the network.","description_state":"present","introduced_year":null,"introduced_by":{"title":"LLaMA: Open and Efficient Foundation Language Models","paper":"/paper/llama-open-and-efficient-foundation-language-1","first_author":"Hugo Touvron","n_authors":14,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/llama-open-and-efficient-foundation-language-1"},"source":{"url":"https://arxiv.org/abs/2302.13971v1","title":"LLaMA: Open and Efficient Foundation Language Models","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Language Models","url":"/methods/category/language-models","pwc_aliases":[]}],"n_papers_tagged":1062,"archive_num_papers":1062,"papers_newest_first":[{"paper":"/paper/making-language-model-a-hierarchical","title":"Making Language Model a Hierarchical Classifier and Generator","date":"2025-07-17","arxiv_id":"2507.12930","n_code_links":1,"syntology":null},{"paper":"/paper/simplifications-are-absolutists-how","title":"Simplifications are Absolutists: How Simplified Language Reduces Word Sense Awareness in LLM-Generated Definitions","date":"2025-07-16","arxiv_id":"2507.11981","n_code_links":1,"syntology":null},{"paper":"/paper/seq-vs-seq-an-open-suite-of-paired-encoders","title":"Seq vs Seq: An Open Suite of Paired Encoders and Decoders","date":"2025-07-15","arxiv_id":"2507.11412","n_code_links":1,"syntology":{"ran":0,"of":10,"unverified":10,"pointer_only":0}},{"paper":null,"title":"Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores","date":"2025-07-10","arxiv_id":"2507.08143","n_code_links":0,"syntology":null},{"paper":null,"title":"Evaluation of Habitat Robotics using Large Language Models","date":"2025-07-08","arxiv_id":"2507.06157","n_code_links":0,"syntology":null},{"paper":null,"title":"MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation","date":"2025-07-08","arxiv_id":"2507.05894","n_code_links":0,"syntology":null},{"paper":"/paper/any4-learned-4-bit-numeric-representation-for","title":"any4: Learned 4-bit Numeric Representation for LLMs","date":"2025-07-07","arxiv_id":"2507.04610","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":2}},{"paper":null,"title":"Model Inversion Attacks on Llama 3: Extracting PII from Large Language Models","date":"2025-07-06","arxiv_id":"2507.04478","n_code_links":0,"syntology":null},{"paper":null,"title":"Large Language Models Acing Chartered Accountancy","date":"2025-06-26","arxiv_id":"2506.21031","n_code_links":0,"syntology":null},{"paper":null,"title":"CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation","date":"2025-06-25","arxiv_id":"2506.20128","n_code_links":0,"syntology":null},{"paper":"/paper/octothinker-mid-training-incentivizes","title":"OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling","date":"2025-06-25","arxiv_id":"2506.20512","n_code_links":1,"syntology":{"ran":3,"of":17,"unverified":14,"pointer_only":0}},{"paper":null,"title":"Can LLMs Replace Humans During Code Chunking?","date":"2025-06-24","arxiv_id":"2506.19897","n_code_links":0,"syntology":null},{"paper":null,"title":"Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives","date":"2025-06-22","arxiv_id":"2506.18116","n_code_links":0,"syntology":null},{"paper":"/paper/pre-trained-llm-is-a-semantic-aware-and","title":"Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster","date":"2025-06-22","arxiv_id":"2506.18034","n_code_links":1,"syntology":null},{"paper":null,"title":"Shrinking the Generation-Verification Gap with Weak Verifiers","date":"2025-06-22","arxiv_id":"2506.18203","n_code_links":0,"syntology":null},{"paper":"/paper/a-minimalist-optimizer-design-for-llm","title":"A Minimalist Optimizer Design for LLM Pretraining","date":"2025-06-20","arxiv_id":"2506.16659","n_code_links":1,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":1}},{"paper":"/paper/all-is-not-lost-llm-recovery-without","title":"All is Not Lost: LLM Recovery without Checkpoints","date":"2025-06-18","arxiv_id":"2506.15461","n_code_links":1,"syntology":null},{"paper":null,"title":"I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution","date":"2025-06-18","arxiv_id":"2506.17323","n_code_links":0,"syntology":null},{"paper":null,"title":"PhantomHunter: Detecting Unseen Privately-Tuned LLM-Generated Text via Family-Aware Learning","date":"2025-06-18","arxiv_id":"2506.15683","n_code_links":0,"syntology":null},{"paper":"/paper/arctic-long-sequence-training-scalable-and","title":"Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences","date":"2025-06-16","arxiv_id":"2506.13996","n_code_links":2,"syntology":null},{"paper":"/paper/attribution-guided-pruning-for-compression","title":"Attribution-guided Pruning for Compression, Circuit Discovery, and Targeted Correction in LLMs","date":"2025-06-16","arxiv_id":"2506.13727","n_code_links":1,"syntology":null},{"paper":null,"title":"Delving Into the Psychology of Machines: Exploring the Structure of Self-Regulated Learning via LLM-Generated Survey Responses","date":"2025-06-16","arxiv_id":"2506.13384","n_code_links":0,"syntology":null},{"paper":"/paper/evaluating-large-language-models-for-phishing","title":"Evaluating Large Language Models for Phishing Detection, Self-Consistency, Faithfulness, and Explainability","date":"2025-06-16","arxiv_id":"2506.13746","n_code_links":1,"syntology":null},{"paper":null,"title":"Language Models Enable Data-Augmented Synthesis Planning for Inorganic Materials","date":"2025-06-14","arxiv_id":"2506.12557","n_code_links":0,"syntology":null},{"paper":"/paper/training-free-llm-merging-for-multi-task","title":"Training-free LLM Merging for Multi-task Learning","date":"2025-06-14","arxiv_id":"2506.12379","n_code_links":1,"syntology":null},{"paper":null,"title":"Beyond Homogeneous Attention: Memory-Efficient LLMs via Fourier-Approximated KV Cache","date":"2025-06-13","arxiv_id":"2506.11886","n_code_links":0,"syntology":null},{"paper":null,"title":"InfoFlood: Jailbreaking Large Language Models with Information Overload","date":"2025-06-13","arxiv_id":"2506.12274","n_code_links":0,"syntology":null},{"paper":null,"title":"LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs","date":"2025-06-12","arxiv_id":"2506.10527","n_code_links":0,"syntology":null},{"paper":null,"title":"Primender Sequence: A Novel Mathematical Construct for Testing Symbolic Inference and AI Reasoning","date":"2025-06-12","arxiv_id":"2506.10585","n_code_links":0,"syntology":null},{"paper":null,"title":"Contemporary AI foundation models increase biological weapons risk","date":"2025-06-12","arxiv_id":"2506.13798","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":158},{"task":"/task/language-modeling","name":"Language Modeling","papers":135},{"task":"/task/large-language-model","name":"Large Language Model","papers":115},{"task":"/task/question-answering","name":"Question Answering","papers":66},{"task":"/task/quantization","name":"Quantization","papers":62},{"task":null,"name":"GPU","papers":48},{"task":"/task/retrieval","name":"Retrieval","papers":46},{"task":"/task/rag","name":"RAG","papers":45},{"task":"/task/retrieval-augmented-generation","name":"Retrieval-augmented Generation","papers":39},{"task":"/task/text-generation","name":"Text Generation","papers":39},{"task":"/task/code-generation","name":"Code Generation","papers":36},{"task":"/task/benchmarking","name":"Benchmarking","papers":34},{"task":"/task/math","name":"Math","papers":34},{"task":"/task/prompt-engineering","name":"Prompt Engineering","papers":33},{"task":"/task/in-context-learning","name":"In-Context Learning","papers":32},{"task":"/task/decision-making","name":"Decision Making","papers":28},{"task":"/task/parameter-efficient-fine-tuning","name":"parameter-efficient fine-tuning","papers":28},{"task":"/task/instruction-following","name":"Instruction Following","papers":27},{"task":"/task/mmlu","name":"MMLU","papers":25},{"task":"/task/decoder","name":"Decoder","papers":24}],"tasks_shown":20,"n_tasks":450,"usage_by_year":[{"year":"2023","papers":30},{"year":"2024","papers":640},{"year":"2025","papers":392}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/llama"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}