Datasets › HellaSwag

HellaSwag

Introduced by Rowan Zellers et al. in HellaSwag: Can a Machine Really Finish Your Sentence? archive 2025-07-28

HellaSwag is a challenge dataset for evaluating commonsense NLI that is specially hard for state-of-the-art models, though its questions are trivial for humans (>95% accuracy).

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Sentence Completion HellaSwag CompassMTL 567M with Tailor Accuracy 96.1 Task Compass: Scaling Multi-task Pre-training with Task Prefix cooelf/compassmtl 89 Compare
parameter-efficient fine-tuning HellaSwag LLaMA2-7b Accuracy (% ) 76.68 GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient... On-Point-RND/GIFT_SW 3 Compare
Question Answering HellaSwag Shakti-LLM (2.5B) Accuracy 52.4 SHAKTI: A 2.5 Billion Parameter Small Language Model... — 1 Compare
Text Generation HellaSwag no rows — — 0 Compare
Text Generation HellaSwag (10-Shot) no rows — — 0 Compare
Text Generation HellaSwag TR no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 39 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 994. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments 0 1 15 Oct 2024 not harvested
GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs 1 1 27 Aug 2024 ran 3 of 5 samples (2 unverified; 3 pointer-only for licence)
Mixture-of-Subspaces in Low-Rank Adaptation 1 1 16 Jun 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts 2 3 22 Apr 2024 ran 6 of 11 samples (5 unverified)
DoRA: Weight-Decomposed Low-Rank Adaptation 5 1 14 Feb 2024 ran 8 of 15 samples (7 unverified; 14 pointer-only for licence)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks 2 2 5 Jan 2024 not harvested
LLM in a flash: Efficient Large Language Model Inference with Limited Memory 0 2 12 Dec 2023 not harvested
Mamba: Linear-Time Sequence Modeling with Selective State Spaces 35 2 1 Dec 2023 ran 18 of 62 samples (44 unverified; 28 pointer-only for licence)
The Falcon Series of Open Language Models 0 3 28 Nov 2023 not harvested
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning 2 3 10 Oct 2023 ran 3 of 3 samples (0 unverified)
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 4 18 Jul 2023 ran 31 of 52 samples (21 unverified; 16 pointer-only for licence)
Stay on topic with Classifier-Free Guidance 0 3 30 Jun 2023 not harvested
The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning 2 1 23 May 2023 not harvested
PaLM 2 Technical Report 1 3 17 May 2023 not harvested
LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions 1 6 27 Apr 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
GPT-4 Technical Report 11 2 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
Exploring the Benefits of Training Expert Language Models over Instruction Tuning 2 1 7 Feb 2023 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question Answering 1 2 29 Oct 2022 ran 1 of 7 samples (6 unverified)
Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models 0 1 28 Oct 2022 not harvested
DiscoSense: Commonsense Reasoning with Discourse Connectives 1 2 22 Oct 2022 not harvested
Task Compass: Scaling Multi-task Pre-training with Task Prefix 1 3 12 Oct 2022 not harvested
Guess the Instruction! Flipped Learning Makes Language Models Stronger Zero-Shot Learners 1 1 6 Oct 2022 not harvested
PaLM: Scaling Language Modeling with Pathways 7 3 5 Apr 2022 ran 30 of 37 samples (7 unverified)
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)
Efficient Language Modeling with Sparse all-MLP 0 5 14 Mar 2022 not harvested
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model 2 2 28 Jan 2022 not harvested
Scaling Language Models: Methods, Analysis & Insights from Training Gopher 3 1 8 Dec 2021 not harvested

The full list of 39 is in the JSON twin.

Dataset loaders archive 2025-07-28

8 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • HellaSwag
  • HellaSwag (10-Shot)
  • HellaSwag TR

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections