Papers › Model Card and Evaluations for Claude Models

Model Card and Evaluations for Claude Models

11 Jul 2023Technical Report 2023 7archive 2025-07-28

Anthropic

This report includes the model card [1] for Claude models, focusing on Claude 2, along with the results of a range of safety, alignment, and capabilities evaluations. We have been iterating on the training and evaluation of Claude-type models since our first work on Reinforcement Learning from Human Feedback (RLHF) [2]; the newest Claude 2 model represents a continuous evolution from those early and less capable ‘helpful and harmless’ language assistants. This report is not intended to be a scientific paper since most aspects of training and evaluating these models have been documented in our research papers. These include papers on preference modeling [3], reinforcement learning from human feedback for helpful and harmless models [2], red teaming language models [4], measuring representation of subjective global values in language models [5], honesty, (i.e., exploring language models’ ability to recognize what they know) [6], evaluating language models with language model-generated tests [7], moral self-correction [8], and Constitutional AI [9]. We also discussed Claude’s specific constitution in a recent blog post [10]. Our work using human evaluations to test model safety is most thoroughly documented in our paper “Red-Teaming Language Models to Reduce Harms” [4], while our recent work on automated safety evaluation is “Discovering Language Model Behaviors with Model-Written Evaluations” [7]. This report is also not comprehensive – we expect to release new findings as we continue our research and evaluations of frontier models. However, we hope it provides useful insight into Claude 2’s capabilities and limitations.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Arithmetic ReasoningBug fixingCode GenerationCommon Sense ReasoningLanguage ModelingLanguage ModellingMulti-task Language UnderstandingQuestion AnsweringRed TeamingSafety Alignmentmodelreinforcement-learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Arithmetic Reasoning GSM8K Claude 2 (0-shot chain-of-thought) Accuracy 88 #32 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Claude 1.3 (0-shot chain-of-thought) Accuracy 85.2 #44 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Claude Instant 1.1 (0-shot chain-of-thought) Accuracy 80.9 #68 of 164 Archive leaderboard report
Common Sense Reasoning ARC (Challenge) Claude 2 (few-shot, k=5) Accuracy 91 #5 of 54 Archive leaderboard report
Common Sense Reasoning ARC (Challenge) Claude 1.3 (few-shot, k=5) Accuracy 90 #6 of 54 Archive leaderboard report
Common Sense Reasoning ARC (Challenge) Claude Instant 1.1 (few-shot, k=5) Accuracy 85.7 #13 of 54 Archive leaderboard report
Question Answering QuALITY Claude 1.3 (5-shot) Accuracy 84.1 #1 of 4 Archive leaderboard report
Question Answering QuALITY Claude 2 (5-shot) Accuracy 83.2 #2 of 4 Archive leaderboard report
Question Answering QuALITY Claude Instant 1.1 (5-shot) Accuracy 80.5 #4 of 4 Archive leaderboard report
Question Answering TriviaQA Claude 2 (few-shot, k=5) EM 87.5 #1 of 56 Archive leaderboard report
Question Answering TriviaQA Claude 1.3 (few-shot, k=5) EM 86.7 #3 of 56 Archive leaderboard report
Question Answering TriviaQA Claude Instant 1.1 (few-shot, k=5) EM 78.9 #15 of 56 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections