| Multi-task Language Understanding |
BBH-nlp |
Qwen2.5-72B Average (%) 86.3 |
— |
— |
15 |
Compare |
| Common Sense Reasoning |
BIG-bench (Disambiguation QA) |
PaLM 2 (few-shot, k=3, Direct) Accuracy 78.8 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Common Sense Reasoning |
BIG-bench (Causal Judgment) |
PaLM 2 (few-shot, k=3, Direct) Accuracy 62.0 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Common Sense Reasoning |
BIG-bench (Date Understanding) |
PaLM 2 (few-shot, k=3, CoT) Accuracy 91.2 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Logical Reasoning |
BIG-bench (Formal Fallacies Syllogisms Negation) |
PaLM 2 (few-shot, k=3, Direct) Accuracy 64.8 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Logical Reasoning |
BIG-bench (Penguins In A Table) |
PaLM 2 (few-shot, k=3, CoT) Accuracy 84.9 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Logical Reasoning |
BIG-bench (Reasoning About Colored Objects) |
PaLM 2 (few-shot, k=3, CoT) Accuracy 91.2 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Logical Reasoning |
BIG-bench (Temporal Sequences) |
PaLM 2 (few-shot, k=3, CoT) Accuracy 100 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Multiple Choice Question Answering (MCQA) |
BIG-bench (Hyperbaton) |
Bloomberg GPT (few-shot, k=3) Accuracy 92 |
BloombergGPT: A Large Language Model for Finance |
yangletliu/finlora +1 |
9 |
Compare |
| Multiple Choice Question Answering (MCQA) |
BIG-bench (Movie Recommendation) |
PaLM 2 (few-shot, k=3, CoT) Accuracy 94.4 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Multiple Choice Question Answering (MCQA) |
BIG-bench (Navigate) |
PaLM 2 (few-shot, k=3, CoT) Accuracy 91.2 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Multiple Choice Question Answering (MCQA) |
BIG-bench (Ruin Names) |
PaLM 2 (few-shot, k=3, Direct) Accuracy 90 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
9 |
Compare |
| Common Sense Reasoning |
BIG-bench (Sports Understanding) |
PaLM 2(few-shot, k=3, CoT) Accuracy 98 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
8 |
Compare |
| Sarcasm Detection |
BIG-bench (SNARKS) |
PaLM 2(few-shot, k=3, CoT) Accuracy 84.8 |
PaLM 2 Technical Report |
eternityyw/tram-benchmark |
8 |
Compare |
| Multi-task Language Understanding |
BBH-alg |
code-davinci-002 175B (CoT) Average (%) 73.9 |
Evaluating Large Language Models Trained on Code |
THUDM/CodeGeeX +12 |
7 |
Compare |
| Word Sense Disambiguation |
BIG-bench (Anachronisms) |
Chinchilla-70B (few-shot, k=5) Accuracy 69.1 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
6 |
Compare |
| Common Sense Reasoning |
BIG-bench (Winowhy) |
PaLM-540B (few-shot, k=5) Accuracy 65.9 |
PaLM: Scaling Language Modeling with Pathways |
lucidrains/CoCa-pytorch +6 |
4 |
Compare |
| Crass AI |
BIG-bench |
Orca 2-13B Accuracy 86.86 |
Orca 2: Teaching Small Language Models How to Reason |
— |
4 |
Compare |
| Logical Reasoning |
BIG-bench (Logic Grid Puzzle) |
Chinchilla-70B (few-shot, k=5) Accuracy 44 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
4 |
Compare |
| Logical Reasoning |
BIG-bench (StrategyQA) |
PaLM-540B (few-shot, k=5) Accuracy 73.9 |
PaLM: Scaling Language Modeling with Pathways |
lucidrains/CoCa-pytorch +6 |
4 |
Compare |
| Multiple Choice Question Answering (MCQA) |
BIG-bench (Novel Concepts) |
PaLM-540B (few-shot, k=5) Accuracy 71.9 |
PaLM: Scaling Language Modeling with Pathways |
lucidrains/CoCa-pytorch +6 |
4 |
Compare |
| Auto Debugging |
Big-bench Lite |
PaLM 62B (few-shot, k=5) Exact string match 38.2 |
PaLM: Scaling Language Modeling with Pathways |
lucidrains/CoCa-pytorch +6 |
3 |
Compare |
| Common Sense Reasoning |
BIG-bench (Known Unknowns) |
PaLM-540B (few-shot, k=5) Accuracy 73.9 |
PaLM: Scaling Language Modeling with Pathways |
lucidrains/CoCa-pytorch +6 |
3 |
Compare |
| Language Modelling |
BIG-bench-lite |
GLM-130B (3-shot) Accuracy 15.11 |
GLM-130B: An Open Bilingual Pre-trained Model |
thudm/chatglm2-6b +8 |
3 |
Compare |
| Memorization |
BIG-bench (Hindu Knowledge) |
PaLM-540B (few-shot, k=5) Accuracy 95.4 |
PaLM: Scaling Language Modeling with Pathways |
lucidrains/CoCa-pytorch +6 |
3 |
Compare |
| Analogical Similarity |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 38.1 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Analytic Entailment |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 67.1 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Common Sense Reasoning |
BIG-bench (Logical Sequence) |
Chinchilla-70B (few-shot, k=5) Accuracy 64.1 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Crash Blossom |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 63.6 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
2 |
Compare |
| Dark Humor Detection |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 83.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
2 |
Compare |
| Discourse Marker Prediction |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 13.1 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Empirical Judgments |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 67.7 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| English Proverbs |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 82.4 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Entailed Polarity |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 94 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Epistemic Reasoning |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 60.6 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Evaluating Information Essentiality |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 17.6 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Fantasy Reasoning |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 69 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Figure Of Speech Detection |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 63.3 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| General Knowledge |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 94.3 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| GRE Reading Comprehension |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 53.1 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Human Organs Senses Multiple Choice |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 85.7 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Identify Odd Metapor |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 68.8 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Implicatures |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 75 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Implicit Relations |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 49.4 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Intent Recognition |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 92.8 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Irony Identification |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 73.0 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| LAMBADA |
BIG-bench |
Chinchilla-70B (zero-shot) Accuracy 77.4 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Logical Args |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 59.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
2 |
Compare |
| Logical Reasoning |
BIG-bench (Logical Fallacy Detection) |
Chinchilla-70B (few-shot, k=5) Accuracy 72.1 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Mathematical Induction |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 57.6 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
2 |
Compare |
| Metaphor Boolean |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 93.1 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Misconceptions |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 65.3 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Moral Permissibility |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 57.3 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Movie Dialog Same Or Different |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 54.5 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Nonsense Words Grammar |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 78 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Odd One Out |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 70.9 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Phrase Relatedness |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 94 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Physical Intuition |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 79 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Presuppositions As NLI |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 49.9 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Question Selection |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 52.6 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Riddle Sense |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 85.7 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Sentence Ambiguity |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 71.7 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Similarities Abstraction |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 87 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Timedial |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 68.8 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Understanding Fables |
BIG-bench |
Chinchilla-70B (few-shot, k=5) Accuracy 60.3 |
Training Compute-Optimal Large Language Models |
karpathy/llama2.c +1 |
2 |
Compare |
| Abstract Algebra |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 25.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Anatomy |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 56.3 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Astronomy |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 65.8 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Business Ethics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 70.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Clinical Knowledge |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 67.2 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| College Mathematics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 37.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| College Medicine |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 60.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Computer Security |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 65.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Econometrics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 43 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Elementary Mathematics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 33.6 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| FEVER (2-way) |
BIG-bench |
Gopher-280B (few-shot, k=10) Accuracy 77.5 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| FEVER (3-way) |
BIG-bench |
Gopher-280B (few-shot, k=15) Accuracy 77.5 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Formal Logic |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 35.7 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Global Facts |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 38.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School European History |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 72.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School Geography |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 76.8 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School Government and Politics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 83.9 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School Macroeconomics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 65.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School Mathematics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 23.7 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School Microeconomics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 66.4 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School Psychology |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 81.8 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School US History |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 78.9 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| High School World History |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 75.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Human Aging |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 66.4 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Human Sexuality |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 67.2 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| International Law |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 77.7 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Jurisprudence |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 71.3 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Logical Fallacies |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 72.4 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| BIG-bench Machine Learning |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 41.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Management |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 77.7 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Marketing |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 83.3 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Medical Genetics |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 69.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Miscellaneous |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 75.7 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Moral Disputes |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 66.8 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Moral Scenarios |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 40.2 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Natural Questions |
BIG-bench |
Gopher-280B (few-shot, k=64) Accuracy 28.2 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Nutrition |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 69.9 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
|
BIG-bench (Hyperbaton) |
CoT-T5 11B Accuracy 65.2 |
The CoT Collection: Improving Zero-shot and Few-shot... |
kaistai/cot-collection +1 |
1 |
Compare |
|
BIG-bench (Navigate) |
CoT-T5 11B Accuracy 60 |
The CoT Collection: Improving Zero-shot and Few-shot... |
kaistai/cot-collection +1 |
1 |
Compare |
|
BIG-bench (Ruin Names) |
CoT-T5 11B Accuracy 42.8 |
The CoT Collection: Improving Zero-shot and Few-shot... |
kaistai/cot-collection +1 |
1 |
Compare |
|
BIG-bench (SNARKS) |
CoT-T5 11B Accuracy 67.7 |
The CoT Collection: Improving Zero-shot and Few-shot... |
kaistai/cot-collection +1 |
1 |
Compare |
| Philosophy |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 68.8 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Prehistory |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 67.6 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Professional Accounting |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 44.3 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Professional Law |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 44.5 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Professional Medicine |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 64.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Professional Psychology |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 68.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Public Relations |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 71.8 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| RACE-h |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 71.6 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| RACE-m |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 75.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Security Studies |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 64.9 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Sociology |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 84.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| TriviaQA |
BIG-bench |
Gopher-280B (few-shot, k=64) Accuracy 57.1 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| US Foreign Policy |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 81.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| Virology |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 47.0 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |
| World Religions |
BIG-bench |
Gopher-280B (few-shot, k=5) Accuracy 84.2 |
Scaling Language Models: Methods, Analysis & Insights... |
allenai/dolma +2 |
1 |
Compare |