Papers › The Claude 3 Model Family: Opus, Sonnet, Haiku

The Claude 3 Model Family: Opus, Sonnet, Haiku

4 Mar 2024Preprint 2024 3archive 2025-07-28

Anthropic

We introduce Claude 3, a new family of large multimodal models – Claude 3 Opus, our most capable offering, Claude 3 Sonnet, which provides a combination of skills and speed, and Claude 3 Haiku, our fastest and least expensive model. All new models have vision capabilities that enable them to process and analyze image data. The Claude 3 family demonstrates strong performance across benchmark evaluations and sets a new standard on measures of reasoning, math, and coding. Claude 3 Opus achieves state-of-the-art results on evaluations like GPQA [1], MMLU [2], MMMU [3] and many more. Claude 3 Haiku performs as well or better than Claude 2 [4] on most pure-text tasks, while Sonnet and Opus significantly outperform it. Additionally, these models exhibit improved fluency in non-English languages, making them more versatile for a global audience. In this report, we provide an in-depth analysis of our evaluations, focusing on core capabilities, safety, societal impacts, and the catastrophic risk assessments we committed to in our Responsible Scaling Policy.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

1 Image, 2*2 StitchingArithmetic ReasoningCode GenerationCommon Sense ReasoningHallucinationImage RetrievalLong-Context UnderstandingMMLUMathMulti-task Language UnderstandingQuestion Answeringmodel

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Arithmetic Reasoning GSM8K Claude 3 Opus (0-shot chain-of-thought) Accuracy 95 #8 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Claude 3 Sonnet (0-shot chain-of-thought) Accuracy 92.3 #15 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Claude 3 Haiku (0-shot chain-of-thought) Accuracy 88.9 #27 of 164 Archive leaderboard report
Code Generation MBPP Claude 3 Opus Accuracy 86.4 #11 of 99 Archive leaderboard report
Code Generation MBPP Claude 3 Haiku Accuracy 80.4 #22 of 99 Archive leaderboard report
Code Generation MBPP Claude 3 Sonnet Accuracy 79.4 #25 of 99 Archive leaderboard report
Common Sense Reasoning WinoGrande Claude 3 Opus (5-shot) Accuracy 88.5 #6 of 77 Archive leaderboard report
Common Sense Reasoning WinoGrande Claude 3 Sonnet (5-shot) Accuracy 75.1 #27 of 77 Archive leaderboard report
Common Sense Reasoning WinoGrande Claude 3 Haiku (5-shot) Accuracy 74.2 #29 of 77 Archive leaderboard report
Long-Context Understanding MMNeedle Claude 3 Opus 1 Image, 2*2 Stitching, Exact Accuracy 52.25 #6 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle Claude 3 Opus 1 Image, 4*4 Stitching, Exact Accuracy 12.3 #6 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle Claude 3 Opus 1 Image, 8*8 Stitching, Exact Accuracy 1.6 #6 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle Claude 3 Opus 10 Images, 1*1 Stitching, Exact Accuracy 66.93 #6 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle Claude 3 Opus 10 Images, 2*2 Stitching, Exact Accuracy 4.6 #6 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle Claude 3 Opus 10 Images, 4*4 Stitching, Exact Accuracy 0.4 #6 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle Claude 3 Opus 10 Images, 8*8 Stitching, Exact Accuracy 0 #6 of 12 Archive leaderboard report
Multi-task Language Understanding MML Claude 3 Sonnet (5-shot) Average (%) 79 #5 of 44 Archive leaderboard report
Multi-task Language Understanding MML Claude 3 Haiku (5-shot) Average (%) 75.2 #7 of 44 Archive leaderboard report
Question Answering PubMedQA Claude 3 Opus (5-shot) Accuracy 75.8 #15 of 30 Archive leaderboard report
Question Answering PubMedQA Claude 3 Opus (zero-shot) Accuracy 74.9 #18 of 30 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections