{"url":"/dataset/only-connect-wall-ocw-dataset","name":"OCW","full_name":"Only Connect Wall Dataset and creative problem solving tasks","description_markdown":"The OCW dataset  is for evaluating creative problem solving tasks by curating the problems and human performance results from the popular British quiz show Only Connect. \r\n\r\nThe OCW dataset contains 618 connecting wall puzzles and solutions in total from 15 seasons of the show. Each show episode has two walls.\r\n\r\nThe dataset has two tasks: Task 1 (Grouping), and Task 2 (Connections) are identical to the quiz-show’s human participant tasks. \r\n\r\nTask 1 (Groupings) is evaluated via six metrics: number of solved walls, number of correct groups (max. four per wall), Adjusted Mutual Information (AMI), Adjusted Rand Index (ARI), Fowlkes Mallows Score (FMS), and Wasserstein Distance (WD), normalized to (0, 1) range, between predicted and ground-truth labels.\r\n\r\nTask 2 (Connections) is evaluated with three metrics: exact string matching, ROUGE-1 F1, and BERTScore F1.\r\n\r\nBaseline results with pre-trained language models and with few-shot In-context Learning (ICL) with LLMs such as GPT-4 are available here:\r\n\r\n\"Large Language Models are Fixated by Red Herrings: Exploring Creative Problem Solving and Einstellung Effect using the Only Connect Wall Dataset\"\r\nSaeid Alavi Naeini, Raeid Saqur, Mozhgan Saeidi, John Giorgi, Babak Taati.\r\n2023\r\nhttps://neurips.cc/virtual/2023/poster/73547","description_withheld":null,"homepage":"https://github.com/TaatiTeam/OCW","introduced_date":"2023-12-14","introduced_date_note":null,"introduced_by":{"paper":"/paper/large-language-models-are-fixated-by-red-1","title":"Large Language Models are Fixated by Red Herrings: Exploring Creative Problem Solving and Einstellung Effect using the Only Connect Wall Dataset","first_author":"Saeid Naeini","url":null},"license":{"name":"MIT License","url":"https://github.com/TaatiTeam/OCW/blob/master/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Natural Language Understanding","url":"/task/natural-language-understanding","datasets_with_task":"/datasets/task/natural-language-understanding"},{"name":"Only Connect Walls Dataset Task 1 (Grouping)","url":"/task/task-1-grouping","datasets_with_task":"/datasets/task/task-1-grouping"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["OCW"],"data_loaders":[],"num_papers_in_archive":10,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/task-1-grouping-on-ocw","task":"Only Connect Walls Dataset Task 1 (Grouping)","dataset_variant":"OCW","rows":22,"metrics":[" Wasserstein Distance (WD)","# Correct Groups","Fowlkes Mallows Score (FMS)","Adjusted Rand Index (ARI)","Adjusted Mutual Information (AMI)","# Solved Walls","Wasserstein Distance (WD)"],"first_row_in_archive_order":{"model":"GPT-4 (5-shot)","paper":"/paper/gpt-4-technical-report-1","metrics":{" Wasserstein Distance (WD)":"72.9","# Correct Groups":"269","# Solved Walls":"7","Adjusted Mutual Information (AMI)":"32.8 ","Adjusted Rand Index (ARI)":"29.1","Fowlkes Mallows Score (FMS)":"43.4"},"code_links":[{"title":"openai/evals","url":"https://github.com/openai/evals"},{"title":"shmsw25/factscore","url":"https://github.com/shmsw25/factscore"},{"title":"unispac/visual-adversarial-examples-jailbreak-large-language-models","url":"https://github.com/unispac/visual-adversarial-examples-jailbreak-large-language-models"},{"title":"gpt4life/alpagasus","url":"https://github.com/gpt4life/alpagasus"},{"title":"emrgnt-cmplxty/zero-shot-replication","url":"https://github.com/emrgnt-cmplxty/zero-shot-replication"},{"title":"ethz-privsec/superhuman-ai-consistency","url":"https://github.com/ethz-privsec/superhuman-ai-consistency"},{"title":"ethz-spylab/superhuman-ai-consistency","url":"https://github.com/ethz-spylab/superhuman-ai-consistency"},{"title":"eternityyw/tram-benchmark","url":"https://github.com/eternityyw/tram-benchmark"},{"title":"AUCOHL/RTL-Repo","url":"https://github.com/AUCOHL/RTL-Repo"},{"title":"zach-zhiling-zheng/reticular_chemist","url":"https://github.com/zach-zhiling-zheng/reticular_chemist"},{"title":"lflage/openfactscore","url":"https://github.com/lflage/openfactscore"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/large-language-models-are-fixated-by-red-1","title":"Large Language Models are Fixated by Red Herrings: Exploring Creative Problem Solving and Einstellung Effect using the Only Connect Wall Dataset","date":"2023-06-19","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/gpt-4-technical-report-1","title":"GPT-4 Technical Report","date":"2023-03-15","rows_on_this_dataset":10,"code_links":11,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":2,"samples_unverified":3,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/text-embeddings-by-weakly-supervised","title":"Text Embeddings by Weakly-Supervised Contrastive Pre-training","date":"2022-12-07","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/mpnet-masked-and-permuted-pre-training-for","title":"MPNet: Masked and Permuted Pre-training for Language Understanding","date":"2020-04-20","rows_on_this_dataset":1,"code_links":7,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":5,"samples_unverified":3,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/pre-training-of-deep-bidirectional-protein","title":"Pre-Training of Deep Bidirectional Protein Sequence Representations with Structural Information","date":"2019-11-25","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/distilbert-a-distilled-version-of-bert","title":"DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter","date":"2019-10-02","rows_on_this_dataset":1,"code_links":37,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":27,"samples_ran":19,"samples_unverified":8,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/roberta-a-robustly-optimized-bert-pretraining","title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","date":"2019-07-26","rows_on_this_dataset":1,"code_links":67,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":48,"samples_ran":22,"samples_unverified":26,"pointer_only_for_licence":23,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learning-word-vectors-for-157-languages","title":"Learning Word Vectors for 157 Languages","date":"2018-02-19","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/deep-contextualized-word-representations","title":"Deep contextualized word representations","date":"2018-02-15","rows_on_this_dataset":1,"code_links":46,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":58,"samples_ran":23,"samples_unverified":35,"pointer_only_for_licence":25,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/glove-global-vectors-for-word-representation","title":"GloVe: Global Vectors for Word Representation","date":"2014-10-01","rows_on_this_dataset":1,"code_links":4,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":6,"samples_harvested":147,"samples_ran":71,"samples_unverified":76,"pointer_only_for_licence":55,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}