Datasets › SIQA

SIQA (Social Interaction QA)

Introduced by Maarten Sap et al. in SocialIQA: Commonsense Reasoning about Social Interactions22 Apr 2019 archive 2025-07-28

Social Interaction QA (SIQA) is a question-answering benchmark for testing social commonsense intelligence. Contrary to many prior benchmarks that focus on physical or taxonomic knowledge, Social IQa focuses on reasoning about people’s actions and their social implications. For example, given an action like "Jesse saw a concert" and a question like "Why did Jesse do this?", humans can easily infer that Jesse wanted "to see their favorite performer" or "to enjoy the music", and not "to see what's happening inside" or "to see if it works". The actions in Social IQa span a wide variety of social situations, and answer candidates contain both human-curated answers and adversarially-filtered machine-generated candidates. Social IQa contains over 37,000 QA pairs for evaluating models’ abilities to reason about the social implications of everyday events and situations.

Source: Social IQA Image Source: https://arxiv.org/pdf/1904.09728.pdf

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Question Answering SIQA Unicorn 11B (fine-tuned) Accuracy 83.2 UNICORN on RAINBOW: A Universal Commonsense Reasoning... allenai/rainbow 24 Compare

Papers archive 2025-07-28

12 shown of 12 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 120. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Mixture-of-Subspaces in Low-Rank Adaptation 1 1 16 Jun 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts 2 3 22 Apr 2024 ran 6 of 11 samples (5 unverified)
Textbooks Are All You Need II: phi-1.5 technical report 1 2 11 Sep 2023 not harvested
LLaMA: Open and Efficient Foundation Language Models 57 4 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question Answering 1 2 29 Oct 2022 ran 1 of 7 samples (6 unverified)
Task Compass: Scaling Multi-task Pre-training with Task Prefix 1 3 12 Oct 2022 not harvested
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)
Scaling Language Models: Methods, Analysis & Insights from Training Gopher 3 1 8 Dec 2021 not harvested
UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark 1 1 24 Mar 2021 ran 0 of 5 samples (5 unverified)
UnifiedQA: Crossing Format Boundaries With a Single QA System 2 1 2 May 2020 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
SocialIQA: Commonsense Reasoning about Social Interactions 1 4 22 Apr 2019 ran 0 of 10 samples (10 unverified)

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • SIQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections