{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/large-language-models-are-fixated-by-red-1","title":"Large Language Models are Fixated by Red Herrings: Exploring Creative Problem Solving and Einstellung Effect using the Only Connect Wall Dataset","arxiv_id":"2306.11167","date":"2023-06-19","proceeding":"NeurIPS 2023 12","authors":["Saeid Naeini","Raeid Saqur","Mozhgan Saeidi","John Giorgi","Babak Taati"],"abstract":"The quest for human imitative AI has been an enduring topic in AI research since its inception. The technical evolution and emerging capabilities of the latest cohort of large language models (LLMs) have reinvigorated the subject beyond academia to the cultural zeitgeist. While recent NLP evaluation benchmark tasks test some aspects of human-imitative behaviour (e.g., BIG-bench's 'human-like behavior' tasks), few, if not none, examine creative problem solving abilities. Creative problem solving in humans is a well-studied topic in cognitive neuroscience with standardized tests that predominantly use the ability to associate (heterogeneous) connections among clue words as a metric for creativity. Exposure to misleading stimuli - distractors dubbed red herrings - impede human performance in such tasks via the fixation effect and Einstellung paradigm. In cognitive neuroscience studies, such fixations are experimentally induced by pre-exposing participants to orthographically similar incorrect words to subsequent word-fragments or clues. The popular British quiz show Only Connect's Connecting Wall segment essentially mimics Mednick's Remote Associates Test (RAT) formulation with built-in, deliberate red herrings, which makes it an ideal proxy dataset to explore and study fixation effect and Einstellung paradigm from cognitive neuroscience in LLMs. In this paper we present the novel Only Connect Wall (OCW) dataset and report results from our evaluation of selected pre-trained language models and LLMs on creative problem solving tasks like grouping clue words by heterogeneous connections, and identifying correct open knowledge domain connections in respective groups. We synthetically generate two additional datasets: OCW-Randomized, OCW-WordNet to further analyze our red-herrings hypothesis in language models. The code and link to the dataset are available at https://github.com/TaatiTeam/OCW.","url_abs":"https://arxiv.org/abs/2306.11167v4","url_pdf":"https://arxiv.org/pdf/2306.11167v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"large-language-models-are-fixated-by-red-1","repo_url":"https://github.com/taatiteam/ocw","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"task-1-grouping","task_name":"Only Connect Walls Dataset Task 1 (Grouping)"},{"task_slug":"only-connect-walls-dataset-task-2-connections","task_name":"Only Connect Walls Dataset Task 2 (Connections)"}],"methods":[],"datasets_introduced":[{"slug":"only-connect-wall-ocw-dataset","name":"OCW","full_name":"Only Connect Wall Dataset and creative problem solving tasks"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/task-1-grouping-on-ocw","task":"Only Connect Walls Dataset Task 1 (Grouping)","dataset":"OCW","model":"Human Performance","rank_in_archive_order":19,"of":22,"metrics":{"# Correct Groups":"1405","# Solved Walls":"285"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2306.11167","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2306.11167"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/taatiteam/ocw","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"97bf761c6b996330","entry":"cosine_similarity","repo":"taatiteam/ocw","repo_kind":"official","path":"src/ocw/utils.py","file_url":"https://github.com/taatiteam/ocw/blob/HEAD/src/ocw/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"97bf761c6b996330"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}