Papers › Orca 2: Teaching Small Language Models How to Reason

Orca 2: Teaching Small Language Models How to Reason

18 Nov 2023arXiv:2311.11045archive 2025-07-28

Arindam Mitra, Luciano del Corro, Shweti Mahajan, Andres Codas, Clarisse Simoes, Sahaj Agarwal, Xuxi Chen, Anastasia Razdaibiedina, Erik Jones, Kriti Aggarwal, Hamid Palangi, Guoqing Zheng, Corby Rosset, Hamed Khanpour, Ahmed Awadallah

Orca 1 learns from rich signals, such as explanation traces, allowing it to outperform conventional instruction-tuned models on benchmarks like BigBench Hard and AGIEval. In Orca 2, we continue exploring how improved training signals can enhance smaller LMs' reasoning abilities. Research on training small LMs has often relied on imitation learning to replicate the output of more capable models. We contend that excessive emphasis on imitation may restrict the potential of smaller models. We seek to teach small LMs to employ different solution strategies for different tasks, potentially different from the one used by the larger model. For example, while larger models might provide a direct answer to a complex task, smaller models may not have the same capacity. In Orca 2, we teach the model various reasoning techniques (step-by-step, recall then generate, recall-reason-generate, direct answer, etc.). More crucially, we aim to help the model learn to determine the most effective solution strategy for each task. We evaluate Orca 2 using a comprehensive set of 15 diverse benchmarks (corresponding to approximately 100 tasks and over 36,000 unique prompts). Orca 2 significantly surpasses models of similar size and attains performance levels similar or better to those of models 5-10x larger, as assessed on complex tasks that test advanced reasoning abilities in zero-shot settings. make Orca 2 weights publicly available at aka.ms/orca-lm to support research on the development, evaluation, and alignment of smaller LMs

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Arithmetic ReasoningCommon Sense ReasoningCounterfactual ReasoningCrass AIImitation LearningMathematical ReasoningMulti-task Language UnderstandingQuestion AnsweringReading Comprehension

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Arithmetic Reasoning GSM8K Orca 2 13B Accuracy 59.14 #118 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Orca 2 13B Parameters (Billion) 13 #118 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Orca 2 7B Accuracy 47.23 #136 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Orca 2 7B Parameters (Billion) 7 #136 of 164 Archive leaderboard report
Crass AI BIG-bench Orca 2-13B Accuracy 86.86 #1 of 4 Archive leaderboard report
Crass AI BIG-bench Orca 2-7B Accuracy 84.31 #2 of 4 Archive leaderboard report
Multi-task Language Understanding BBH-nlp Orca 2-13B Average (%) 50.18 #14 of 15 Archive leaderboard report
Multi-task Language Understanding BBH-nlp Orca 2-7B Average (%) 45.93 #15 of 15 Archive leaderboard report
Question Answering AGI Eval Orca 2-13B Accuracy 49.93 #1 of 2 Archive leaderboard report
Question Answering AGI Eval Orca 2-7B Accuracy 45.1 #2 of 2 Archive leaderboard report
Question Answering DROP Test Orca 2-7B F1 60.26 #12 of 16 Archive leaderboard report
Question Answering DROP Test Orca 2-13B F1 57.97 #13 of 16 Archive leaderboard report
Reading Comprehension RACE Orca 2-13B Accuracy 82.87 #8 of 24 Archive leaderboard report
Reading Comprehension RACE Orca 2-7B Accuracy 80.79 #9 of 24 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

SET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections