Papers › The Claude 3 Model Family: Opus, Sonnet, Haiku
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic
We introduce Claude 3, a new family of large multimodal models – Claude 3 Opus, our most capable offering, Claude 3 Sonnet, which provides a combination of skills and speed, and Claude 3 Haiku, our fastest and least expensive model. All new models have vision capabilities that enable them to process and analyze image data. The Claude 3 family demonstrates strong performance across benchmark evaluations and sets a new standard on measures of reasoning, math, and coding. Claude 3 Opus achieves state-of-the-art results on evaluations like GPQA [1], MMLU [2], MMMU [3] and many more. Claude 3 Haiku performs as well or better than Claude 2 [4] on most pure-text tasks, while Sonnet and Opus significantly outperform it. Additionally, these models exhibit improved fluency in non-English languages, making them more versatile for a global audience. In this report, we provide an in-depth analysis of our evaluations, focusing on core capabilities, safety, societal impacts, and the catastrophic risk assessments we committed to in our Responsible Scaling Policy.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Arithmetic Reasoning | GSM8K | Claude 3 Opus (0-shot chain-of-thought) | Accuracy | 95 | #8 of 164 | Archive leaderboard | report |
| Arithmetic Reasoning | GSM8K | Claude 3 Sonnet (0-shot chain-of-thought) | Accuracy | 92.3 | #15 of 164 | Archive leaderboard | report |
| Arithmetic Reasoning | GSM8K | Claude 3 Haiku (0-shot chain-of-thought) | Accuracy | 88.9 | #27 of 164 | Archive leaderboard | report |
| Code Generation | MBPP | Claude 3 Opus | Accuracy | 86.4 | #11 of 99 | Archive leaderboard | report |
| Code Generation | MBPP | Claude 3 Haiku | Accuracy | 80.4 | #22 of 99 | Archive leaderboard | report |
| Code Generation | MBPP | Claude 3 Sonnet | Accuracy | 79.4 | #25 of 99 | Archive leaderboard | report |
| Common Sense Reasoning | WinoGrande | Claude 3 Opus (5-shot) | Accuracy | 88.5 | #6 of 77 | Archive leaderboard | report |
| Common Sense Reasoning | WinoGrande | Claude 3 Sonnet (5-shot) | Accuracy | 75.1 | #27 of 77 | Archive leaderboard | report |
| Common Sense Reasoning | WinoGrande | Claude 3 Haiku (5-shot) | Accuracy | 74.2 | #29 of 77 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | Claude 3 Opus | 1 Image, 2*2 Stitching, Exact Accuracy | 52.25 | #6 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | Claude 3 Opus | 1 Image, 4*4 Stitching, Exact Accuracy | 12.3 | #6 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | Claude 3 Opus | 1 Image, 8*8 Stitching, Exact Accuracy | 1.6 | #6 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | Claude 3 Opus | 10 Images, 1*1 Stitching, Exact Accuracy | 66.93 | #6 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | Claude 3 Opus | 10 Images, 2*2 Stitching, Exact Accuracy | 4.6 | #6 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | Claude 3 Opus | 10 Images, 4*4 Stitching, Exact Accuracy | 0.4 | #6 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | Claude 3 Opus | 10 Images, 8*8 Stitching, Exact Accuracy | 0 | #6 of 12 | Archive leaderboard | report |
| Multi-task Language Understanding | MML | Claude 3 Sonnet (5-shot) | Average (%) | 79 | #5 of 44 | Archive leaderboard | report |
| Multi-task Language Understanding | MML | Claude 3 Haiku (5-shot) | Average (%) | 75.2 | #7 of 44 | Archive leaderboard | report |
| Question Answering | PubMedQA | Claude 3 Opus (5-shot) | Accuracy | 75.8 | #15 of 30 | Archive leaderboard | report |
| Question Answering | PubMedQA | Claude 3 Opus (zero-shot) | Accuracy | 74.9 | #18 of 30 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections