Datasets › BIG-bench

BIG-bench (Beyond the Imitation Game Benchmark)

Introduced by Aarohi Srivastava et al. in Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models9 Jun 2022 archive 2025-07-28

The Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark intended to probe large language models and extrapolate their future capabilities. Big-bench include more than 200 tasks.

Image source: https://arxiv.org/pdf/2206.04615.pdf

Benchmarks archive 2025-07-28

All 121 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Multi-task Language Understanding BBH-nlp Qwen2.5-72B Average (%) 86.3 — — 15 Compare
Common Sense Reasoning BIG-bench (Disambiguation QA) PaLM 2 (few-shot, k=3, Direct) Accuracy 78.8 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Common Sense Reasoning BIG-bench (Causal Judgment) PaLM 2 (few-shot, k=3, Direct) Accuracy 62.0 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Common Sense Reasoning BIG-bench (Date Understanding) PaLM 2 (few-shot, k=3, CoT) Accuracy 91.2 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Logical Reasoning BIG-bench (Formal Fallacies Syllogisms Negation) PaLM 2 (few-shot, k=3, Direct) Accuracy 64.8 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Logical Reasoning BIG-bench (Penguins In A Table) PaLM 2 (few-shot, k=3, CoT) Accuracy 84.9 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Logical Reasoning BIG-bench (Reasoning About Colored Objects) PaLM 2 (few-shot, k=3, CoT) Accuracy 91.2 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Logical Reasoning BIG-bench (Temporal Sequences) PaLM 2 (few-shot, k=3, CoT) Accuracy 100 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Multiple Choice Question Answering (MCQA) BIG-bench (Hyperbaton) Bloomberg GPT (few-shot, k=3) Accuracy 92 BloombergGPT: A Large Language Model for Finance yangletliu/finlora +1 9 Compare
Multiple Choice Question Answering (MCQA) BIG-bench (Movie Recommendation) PaLM 2 (few-shot, k=3, CoT) Accuracy 94.4 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Multiple Choice Question Answering (MCQA) BIG-bench (Navigate) PaLM 2 (few-shot, k=3, CoT) Accuracy 91.2 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Multiple Choice Question Answering (MCQA) BIG-bench (Ruin Names) PaLM 2 (few-shot, k=3, Direct) Accuracy 90 PaLM 2 Technical Report eternityyw/tram-benchmark 9 Compare
Common Sense Reasoning BIG-bench (Sports Understanding) PaLM 2(few-shot, k=3, CoT) Accuracy 98 PaLM 2 Technical Report eternityyw/tram-benchmark 8 Compare
Sarcasm Detection BIG-bench (SNARKS) PaLM 2(few-shot, k=3, CoT) Accuracy 84.8 PaLM 2 Technical Report eternityyw/tram-benchmark 8 Compare
Multi-task Language Understanding BBH-alg code-davinci-002 175B (CoT) Average (%) 73.9 Evaluating Large Language Models Trained on Code THUDM/CodeGeeX +12 7 Compare
Word Sense Disambiguation BIG-bench (Anachronisms) Chinchilla-70B (few-shot, k=5) Accuracy 69.1 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 6 Compare
Common Sense Reasoning BIG-bench (Winowhy) PaLM-540B (few-shot, k=5) Accuracy 65.9 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 4 Compare
Crass AI BIG-bench Orca 2-13B Accuracy 86.86 Orca 2: Teaching Small Language Models How to Reason — 4 Compare
Logical Reasoning BIG-bench (Logic Grid Puzzle) Chinchilla-70B (few-shot, k=5) Accuracy 44 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 4 Compare
Logical Reasoning BIG-bench (StrategyQA) PaLM-540B (few-shot, k=5) Accuracy 73.9 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 4 Compare
Multiple Choice Question Answering (MCQA) BIG-bench (Novel Concepts) PaLM-540B (few-shot, k=5) Accuracy 71.9 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 4 Compare
Auto Debugging Big-bench Lite PaLM 62B (few-shot, k=5) Exact string match 38.2 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 3 Compare
Common Sense Reasoning BIG-bench (Known Unknowns) PaLM-540B (few-shot, k=5) Accuracy 73.9 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 3 Compare
Language Modelling BIG-bench-lite GLM-130B (3-shot) Accuracy 15.11 GLM-130B: An Open Bilingual Pre-trained Model thudm/chatglm2-6b +8 3 Compare
Memorization BIG-bench (Hindu Knowledge) PaLM-540B (few-shot, k=5) Accuracy 95.4 PaLM: Scaling Language Modeling with Pathways lucidrains/CoCa-pytorch +6 3 Compare
Analogical Similarity BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 38.1 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Analytic Entailment BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 67.1 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Common Sense Reasoning BIG-bench (Logical Sequence) Chinchilla-70B (few-shot, k=5) Accuracy 64.1 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Crash Blossom BIG-bench Gopher-280B (few-shot, k=5) Accuracy 63.6 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 2 Compare
Dark Humor Detection BIG-bench Gopher-280B (few-shot, k=5) Accuracy 83.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 2 Compare
Discourse Marker Prediction BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 13.1 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Empirical Judgments BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 67.7 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
English Proverbs BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 82.4 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Entailed Polarity BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 94 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Epistemic Reasoning BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 60.6 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Evaluating Information Essentiality BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 17.6 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Fantasy Reasoning BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 69 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Figure Of Speech Detection BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 63.3 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
General Knowledge BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 94.3 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
GRE Reading Comprehension BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 53.1 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Human Organs Senses Multiple Choice BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 85.7 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Identify Odd Metapor BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 68.8 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Implicatures BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 75 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Implicit Relations BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 49.4 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Intent Recognition BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 92.8 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Irony Identification BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 73.0 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
LAMBADA BIG-bench Chinchilla-70B (zero-shot) Accuracy 77.4 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Logical Args BIG-bench Gopher-280B (few-shot, k=5) Accuracy 59.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 2 Compare
Logical Reasoning BIG-bench (Logical Fallacy Detection) Chinchilla-70B (few-shot, k=5) Accuracy 72.1 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Mathematical Induction BIG-bench Gopher-280B (few-shot, k=5) Accuracy 57.6 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 2 Compare
Metaphor Boolean BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 93.1 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Misconceptions BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 65.3 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Moral Permissibility BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 57.3 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Movie Dialog Same Or Different BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 54.5 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Nonsense Words Grammar BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 78 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Odd One Out BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 70.9 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Phrase Relatedness BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 94 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Physical Intuition BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 79 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Presuppositions As NLI BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 49.9 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Question Selection BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 52.6 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Riddle Sense BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 85.7 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Sentence Ambiguity BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 71.7 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Similarities Abstraction BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 87 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Timedial BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 68.8 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Understanding Fables BIG-bench Chinchilla-70B (few-shot, k=5) Accuracy 60.3 Training Compute-Optimal Large Language Models karpathy/llama2.c +1 2 Compare
Abstract Algebra BIG-bench Gopher-280B (few-shot, k=5) Accuracy 25.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Anatomy BIG-bench Gopher-280B (few-shot, k=5) Accuracy 56.3 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Astronomy BIG-bench Gopher-280B (few-shot, k=5) Accuracy 65.8 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Business Ethics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 70.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Clinical Knowledge BIG-bench Gopher-280B (few-shot, k=5) Accuracy 67.2 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
College Mathematics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 37.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
College Medicine BIG-bench Gopher-280B (few-shot, k=5) Accuracy 60.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Computer Security BIG-bench Gopher-280B (few-shot, k=5) Accuracy 65.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Econometrics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 43 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Elementary Mathematics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 33.6 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
FEVER (2-way) BIG-bench Gopher-280B (few-shot, k=10) Accuracy 77.5 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
FEVER (3-way) BIG-bench Gopher-280B (few-shot, k=15) Accuracy 77.5 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Formal Logic BIG-bench Gopher-280B (few-shot, k=5) Accuracy 35.7 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Global Facts BIG-bench Gopher-280B (few-shot, k=5) Accuracy 38.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School European History BIG-bench Gopher-280B (few-shot, k=5) Accuracy 72.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School Geography BIG-bench Gopher-280B (few-shot, k=5) Accuracy 76.8 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School Government and Politics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 83.9 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School Macroeconomics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 65.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School Mathematics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 23.7 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School Microeconomics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 66.4 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School Psychology BIG-bench Gopher-280B (few-shot, k=5) Accuracy 81.8 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School US History BIG-bench Gopher-280B (few-shot, k=5) Accuracy 78.9 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
High School World History BIG-bench Gopher-280B (few-shot, k=5) Accuracy 75.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Human Aging BIG-bench Gopher-280B (few-shot, k=5) Accuracy 66.4 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Human Sexuality BIG-bench Gopher-280B (few-shot, k=5) Accuracy 67.2 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
International Law BIG-bench Gopher-280B (few-shot, k=5) Accuracy 77.7 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Jurisprudence BIG-bench Gopher-280B (few-shot, k=5) Accuracy 71.3 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Logical Fallacies BIG-bench Gopher-280B (few-shot, k=5) Accuracy 72.4 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
BIG-bench Machine Learning BIG-bench Gopher-280B (few-shot, k=5) Accuracy 41.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Management BIG-bench Gopher-280B (few-shot, k=5) Accuracy 77.7 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Marketing BIG-bench Gopher-280B (few-shot, k=5) Accuracy 83.3 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Medical Genetics BIG-bench Gopher-280B (few-shot, k=5) Accuracy 69.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Miscellaneous BIG-bench Gopher-280B (few-shot, k=5) Accuracy 75.7 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Moral Disputes BIG-bench Gopher-280B (few-shot, k=5) Accuracy 66.8 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Moral Scenarios BIG-bench Gopher-280B (few-shot, k=5) Accuracy 40.2 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Natural Questions BIG-bench Gopher-280B (few-shot, k=64) Accuracy 28.2 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Nutrition BIG-bench Gopher-280B (few-shot, k=5) Accuracy 69.9 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
BIG-bench (Hyperbaton) CoT-T5 11B Accuracy 65.2 The CoT Collection: Improving Zero-shot and Few-shot... kaistai/cot-collection +1 1 Compare
BIG-bench (Navigate) CoT-T5 11B Accuracy 60 The CoT Collection: Improving Zero-shot and Few-shot... kaistai/cot-collection +1 1 Compare
BIG-bench (Ruin Names) CoT-T5 11B Accuracy 42.8 The CoT Collection: Improving Zero-shot and Few-shot... kaistai/cot-collection +1 1 Compare
BIG-bench (SNARKS) CoT-T5 11B Accuracy 67.7 The CoT Collection: Improving Zero-shot and Few-shot... kaistai/cot-collection +1 1 Compare
Philosophy BIG-bench Gopher-280B (few-shot, k=5) Accuracy 68.8 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Prehistory BIG-bench Gopher-280B (few-shot, k=5) Accuracy 67.6 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Professional Accounting BIG-bench Gopher-280B (few-shot, k=5) Accuracy 44.3 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Professional Law BIG-bench Gopher-280B (few-shot, k=5) Accuracy 44.5 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Professional Medicine BIG-bench Gopher-280B (few-shot, k=5) Accuracy 64.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Professional Psychology BIG-bench Gopher-280B (few-shot, k=5) Accuracy 68.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Public Relations BIG-bench Gopher-280B (few-shot, k=5) Accuracy 71.8 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
RACE-h BIG-bench Gopher-280B (few-shot, k=5) Accuracy 71.6 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
RACE-m BIG-bench Gopher-280B (few-shot, k=5) Accuracy 75.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Security Studies BIG-bench Gopher-280B (few-shot, k=5) Accuracy 64.9 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Sociology BIG-bench Gopher-280B (few-shot, k=5) Accuracy 84.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
TriviaQA BIG-bench Gopher-280B (few-shot, k=64) Accuracy 57.1 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
US Foreign Policy BIG-bench Gopher-280B (few-shot, k=5) Accuracy 81.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Virology BIG-bench Gopher-280B (few-shot, k=5) Accuracy 47.0 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
World Religions BIG-bench Gopher-280B (few-shot, k=5) Accuracy 84.2 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 349. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Orca 2: Teaching Small Language Models How to Reason 0 4 18 Nov 2023 not harvested
The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning 2 4 23 May 2023 not harvested
PaLM 2 Technical Report 1 28 17 May 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 63 30 Mar 2023 not harvested
Galactica: A Large Language Model for Science 1 4 16 Nov 2022 ran 0 of 2 samples (2 unverified)
Scaling Instruction-Finetuned Language Models 9 12 20 Oct 2022 ran 8 of 17 samples (9 unverified; 2 pointer-only for licence)
GLM-130B: An Open Bilingual Pre-trained Model 9 3 5 Oct 2022 ran 5 of 21 samples (16 unverified)
PaLM: Scaling Language Modeling with Pathways 7 12 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Training Compute-Optimal Large Language Models 2 60 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)
Scaling Language Models: Methods, Analysis & Insights from Training Gopher 3 113 8 Dec 2021 not harvested
Evaluating Large Language Models Trained on Code 13 2 7 Jul 2021 ran 6 of 39 samples (33 unverified; 2 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

taskLanguage ModellingCommon Sense ReasoningMultiple Choice Question Answering (MCQA)Logical ReasoningWord Sense DisambiguationSarcasm DetectionGeneral KnowledgeMulti-task Language UnderstandingIntent RecognitionBIG-bench Machine LearningRiddle SenseNatural QuestionsAnalogical SimilarityIdentify Odd MetaporOdd One OutCrash BlossomAuto DebuggingCrass AIDiscourse Marker PredictionEmpirical JudgmentsIrony IdentificationTimedialUnderstanding FablesDark Humor DetectionBusiness EthicsMoral DisputesMoral PermissibilityMoral ScenariosFEVER (2-way)FEVER (3-way)MisconceptionsSentence AmbiguityGlobal FactsMiscellaneousSimilarities AbstractionTriviaQAHigh School European HistoryHigh School US HistoryHigh School World HistoryInternational LawJurisprudenceLogical FallaciesManagementMarketingPhilosophyPrehistoryProfessional LawWorld ReligionsAnalytic EntailmentEntailed PolarityEpistemic ReasoningEvaluating Information EssentialityLogical ArgsMetaphor BooleanPhysical IntuitionPresuppositions As NLIAbstract AlgebraCollege MathematicsElementary MathematicsFormal LogicHigh School MathematicsMathematical InductionProfessional AccountingAnatomyClinical KnowledgeCollege MedicineHuman AgingHuman Organs Senses Multiple ChoiceMedical GeneticsNutritionProfessional MedicineVirologyEnglish ProverbsFantasy ReasoningFigure Of Speech DetectionGRE Reading ComprehensionImplicaturesImplicit RelationsLAMBADAMovie Dialog Same Or DifferentNonsense Words GrammarPhrase RelatednessQuestion SelectionRACE-hRACE-mAstronomyCollege BiologyCollege ChemistryCollege Computer ScienceCollege PhysicsComputer SecurityConceptual PhysicsElectrical EngineeringHigh School BiologyHigh School ChemistryHigh School Computer ScienceHigh School PhysicsHigh School StatisticsPhysics MCEconometricsHigh School GeographyHigh School Government and PoliticsHigh School MacroeconomicsHigh School MicroeconomicsHigh School PsychologyHuman SexualityProfessional PsychologyPublic RelationsSecurity StudiesSociologyUS Foreign PolicyMemorization

License archive 2025-07-28

Apache License 2.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • BBH-nlp
  • BBH-alg
  • Big-bench Hard
  • BIG-bench (Logical Sequence)
  • BIG-bench (Logical Fallacy Detection)
  • BIG-bench (Known Unknowns)
  • BIG-bench (Hindu Knowledge)
  • BIG-bench (Novel Concepts)
  • BIG-bench (StrategyQA)
  • BIG-bench (Winowhy)
  • BIG-bench (Logic Grid Puzzle)
  • BIG-bench (Anachronisms)
  • BIG-bench (Temporal Sequences)
  • BIG-bench (Sports Understanding)
  • BIG-bench (SNARKS)
  • BIG-bench (Ruin Names)
  • BIG-bench (Reasoning About Colored Objects)
  • BIG-bench (Penguins In A Table)
  • BIG-bench (Navigate)
  • BIG-bench (Movie Recommendation)
  • BIG-bench (Hyperbaton)
  • BIG-bench (Formal Fallacies Syllogisms Negation)
  • BIG-bench (Disambiguation QA)
  • BIG-bench (Date Understanding)
  • BIG-bench (Causal Judgment)
  • BIG-bench-lite
  • Big-bench Lite
  • BIG-bench

28 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections