Datasets › OCW

OCW (Only Connect Wall Dataset and creative problem solving tasks)

Introduced by Saeid Naeini et al. in Large Language Models are Fixated by Red Herrings: Exploring Creative Problem Solving and Einstellung Effect using the Only Connect Wall Dataset14 Dec 2023 archive 2025-07-28

The OCW dataset is for evaluating creative problem solving tasks by curating the problems and human performance results from the popular British quiz show Only Connect.

The OCW dataset contains 618 connecting wall puzzles and solutions in total from 15 seasons of the show. Each show episode has two walls.

The dataset has two tasks: Task 1 (Grouping), and Task 2 (Connections) are identical to the quiz-show’s human participant tasks.

Task 1 (Groupings) is evaluated via six metrics: number of solved walls, number of correct groups (max. four per wall), Adjusted Mutual Information (AMI), Adjusted Rand Index (ARI), Fowlkes Mallows Score (FMS), and Wasserstein Distance (WD), normalized to (0, 1) range, between predicted and ground-truth labels.

Task 2 (Connections) is evaluated with three metrics: exact string matching, ROUGE-1 F1, and BERTScore F1.

Baseline results with pre-trained language models and with few-shot In-context Learning (ICL) with LLMs such as GPT-4 are available here:

"Large Language Models are Fixated by Red Herrings: Exploring Creative Problem Solving and Einstellung Effect using the Only Connect Wall Dataset" Saeid Alavi Naeini, Raeid Saqur, Mozhgan Saeidi, John Giorgi, Babak Taati. 2023 https://neurips.cc/virtual/2023/poster/73547

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Only Connect Walls Dataset Task 1 (Grouping) OCW GPT-4 (5-shot) Wasserstein Distance (WD) 72.9 GPT-4 Technical Report openai/evals +10 22 Compare

Papers archive 2025-07-28

10 shown of 10 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 10. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Large Language Models are Fixated by Red Herrings: Exploring Creative Problem Solving and Einstellung Effect using the Only Connect Wall Dataset 1 1 19 Jun 2023 ran 0 of 1 samples (1 unverified)
GPT-4 Technical Report 11 10 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
Text Embeddings by Weakly-Supervised Contrastive Pre-training 1 2 7 Dec 2022 not harvested
MPNet: Masked and Permuted Pre-training for Language Understanding 7 1 20 Apr 2020 ran 5 of 8 samples (3 unverified; 6 pointer-only for licence)
Pre-Training of Deep Bidirectional Protein Sequence Representations with Structural Information 1 2 25 Nov 2019 not harvested
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter 37 1 2 Oct 2019 ran 19 of 27 samples (8 unverified)
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
Learning Word Vectors for 157 Languages 2 2 19 Feb 2018 not harvested
Deep contextualized word representations 46 1 15 Feb 2018 ran 23 of 58 samples (35 unverified; 25 pointer-only for licence)
GloVe: Global Vectors for Word Representation 4 1 1 Oct 2014 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • OCW

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections