{"url":"/task/misconceptions","name":"Misconceptions","slug":"misconceptions","description_markdown":"Measures whether a model can discern popular misconceptions from the truth.\r\n\r\nExample:\r\n\r\n```\r\n        input: The daddy longlegs spider is the most venomous spider in the world.\r\n        choice: T\r\n        choice: F\r\n        answer: F\r\n\r\n        input: Karl Benz is correctly credited with the invention of the first modern automobile.\r\n        choice: T\r\n        choice: F\r\n        answer: T\r\n```\r\n\r\nSource: [BIG-bench](https://github.com/google/BIG-bench/tree/main/bigbench/benchmark_tasks/misconceptions)","categories":[{"name":"Miscellaneous","url":"/area/miscellaneous"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":161,"papers_with_code":54,"benchmarks":1,"benchmark_tables_in_archive":1,"benchmark_tables_shown":1,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":1,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/misconceptions-on-big-bench","slug":"misconceptions-on-big-bench","dataset":"BIG-bench","dataset_url":"/dataset/big-bench","rows_in_archive":2,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"Chinchilla-70B (few-shot, k=5)","paper_title":"Training Compute-Optimal Large Language Models","paper_url":"/paper/training-compute-optimal-large-language","paper_date":"2022-03-29","arxiv_id":"2203.15556","code_links":[{"title":"karpathy/llama2.c","url":"https://github.com/karpathy/llama2.c"},{"title":"nkluge-correa/teenytinyllama","url":"https://github.com/nkluge-correa/teenytinyllama"}],"syntology":{"n":11,"n_ran":8,"n_unverified":3,"n_pointer_only":4}}}],"datasets":[{"url":"/dataset/big-bench","name":"BIG-bench","full_name":"Beyond the Imitation Game Benchmark","num_papers_in_archive":349}],"subtasks":[],"parent_tasks":[{"url":"/task/fact-checking","name":"Fact Checking"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":54,"tagged_in_all":161,"items":[{"url":"/paper/community-detection-in-networks-a-user-guide","title":"Community detection in networks: A user guide","date":"2016-07-30","arxiv_id":"1608.00163","repositories_listed":18,"syntology":null},{"url":"/paper/laplace-redux-effortless-bayesian-deep","title":"Laplace Redux -- Effortless Bayesian Deep Learning","date":"2021-06-28","arxiv_id":"2106.14806","repositories_listed":6,"syntology":{"n":5,"n_ran":1,"n_unverified":4,"n_pointer_only":1}},{"url":"/paper/factuality-enhanced-language-models-for-open","title":"Factuality Enhanced Language Models for Open-Ended Text Generation","date":"2022-06-09","arxiv_id":"2206.04624","repositories_listed":5,"syntology":{"n":11,"n_ran":2,"n_unverified":9,"n_pointer_only":4}},{"url":"/paper/parting-with-misconceptions-about-learning","title":"Parting with Misconceptions about Learning-based Vehicle Motion Planning","date":"2023-06-13","arxiv_id":"2306.07962","repositories_listed":3,"syntology":null},{"url":"/paper/scaling-language-models-methods-analysis-1","title":"Scaling Language Models: Methods, Analysis & Insights from Training Gopher","date":"2021-12-08","arxiv_id":"2112.11446","repositories_listed":3,"syntology":null},{"url":"/paper/truthfulqa-measuring-how-models-mimic-human","title":"TruthfulQA: Measuring How Models Mimic Human Falsehoods","date":"2021-09-08","arxiv_id":"2109.07958","repositories_listed":3,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/unveiling-contrastive-learning-s-capability","title":"Unveiling Contrastive Learning's Capability of Neighborhood Aggregation for Collaborative Filtering","date":"2025-04-14","arxiv_id":"2504.10113","repositories_listed":2,"syntology":null},{"url":"/paper/reliability-check-an-analysis-of-gpt-3-s","title":"Reliability Check: An Analysis of GPT-3's Response to Sensitive Topics and Prompt Wording","date":"2023-06-09","arxiv_id":"2306.06199","repositories_listed":2,"syntology":null},{"url":"/paper/training-compute-optimal-large-language","title":"Training Compute-Optimal Large Language Models","date":"2022-03-29","arxiv_id":"2203.15556","repositories_listed":2,"syntology":{"n":11,"n_ran":8,"n_unverified":3,"n_pointer_only":4}},{"url":"/paper/on-the-stability-of-fine-tuning-bert","title":"On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines","date":"2020-06-08","arxiv_id":"2006.04884","repositories_listed":2,"syntology":{"n":20,"n_ran":7,"n_unverified":13,"n_pointer_only":0}},{"url":"/paper/design-challenges-and-misconceptions-in","title":"Design Challenges and Misconceptions in Neural Sequence Labeling","date":"2018-06-12","arxiv_id":"1806.04470","repositories_listed":2,"syntology":null},{"url":"/paper/a-structured-unplugged-approach-for","title":"A Structured Unplugged Approach for Foundational AI Literacy in Primary Education","date":"2025-05-27","arxiv_id":"2505.21398","repositories_listed":1,"syntology":null},{"url":"/paper/when-ai-co-scientists-fail-spot-a-benchmark","title":"When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research","date":"2025-05-17","arxiv_id":"2505.11855","repositories_listed":1,"syntology":null},{"url":"/paper/harnessing-structured-knowledge-a-concept-map","title":"Harnessing Structured Knowledge: A Concept Map-Based Approach for High-Quality Multiple Choice Question Generation with Effective Distractors","date":"2025-05-02","arxiv_id":"2505.02850","repositories_listed":1,"syntology":null},{"url":"/paper/how-to-protect-yourself-from-5g-radiation","title":"How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation","date":"2025-03-12","arxiv_id":"2503.09598","repositories_listed":1,"syntology":null},{"url":"/paper/paths-and-ambient-spaces-in-neural-loss","title":"Paths and Ambient Spaces in Neural Loss Landscapes","date":"2025-03-05","arxiv_id":"2503.03382","repositories_listed":1,"syntology":null},{"url":"/paper/learning-to-correction-explainable-feedback","title":"Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor","date":"2024-12-08","arxiv_id":"2412.07801","repositories_listed":1,"syntology":null},{"url":"/paper/exploring-knowledge-tracing-in-tutor-student","title":"Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs","date":"2024-09-24","arxiv_id":"2409.16490","repositories_listed":1,"syntology":null},{"url":"/paper/enhancing-knowledge-tracing-with-concept-map","title":"Enhancing Knowledge Tracing with Concept Map and Response Disentanglement","date":"2024-08-23","arxiv_id":"2408.12996","repositories_listed":1,"syntology":null},{"url":"/paper/when-big-data-actually-are-low-rank-or","title":"When big data actually are low-rank, or entrywise approximation of certain function-generated matrices","date":"2024-07-03","arxiv_id":"2407.03250","repositories_listed":1,"syntology":null},{"url":"/paper/malalgoqa-a-pedagogical-approach-for","title":"MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education","date":"2024-07-01","arxiv_id":"2407.00938","repositories_listed":1,"syntology":null},{"url":"/paper/divert-distractor-generation-with-variational","title":"DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions","date":"2024-06-27","arxiv_id":"2406.19356","repositories_listed":1,"syntology":{"n":3,"n_ran":1,"n_unverified":2,"n_pointer_only":3}},{"url":"/paper/student-answer-forecasting-transformer-driven","title":"Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning","date":"2024-05-30","arxiv_id":"2405.20079","repositories_listed":1,"syntology":null},{"url":"/paper/a-closer-look-at-classification-evaluation","title":"A Closer Look at Classification Evaluation Metrics and a Critical Reflection of Common Evaluation Practice","date":"2024-04-25","arxiv_id":"2404.16958","repositories_listed":1,"syntology":null},{"url":"/paper/exploring-automated-distractor-generation-for","title":"Exploring Automated Distractor Generation for Math Multiple-choice Questions via Large Language Models","date":"2024-04-02","arxiv_id":"2404.02124","repositories_listed":1,"syntology":null},{"url":"/paper/the-pitfalls-of-next-token-prediction","title":"The pitfalls of next-token prediction","date":"2024-03-11","arxiv_id":"2403.06963","repositories_listed":1,"syntology":{"n":7,"n_ran":6,"n_unverified":1,"n_pointer_only":7}},{"url":"/paper/the-power-of-noise-toward-a-unified-multi","title":"Noise-powered Multi-modal Knowledge Graph Representation Framework","date":"2024-03-11","arxiv_id":"2403.06832","repositories_listed":1,"syntology":null},{"url":"/paper/watchat-explaining-perplexing-programs-by","title":"WatChat: Explaining perplexing programs by debugging mental models","date":"2024-03-08","arxiv_id":"2403.05334","repositories_listed":1,"syntology":null},{"url":"/paper/improving-the-validity-of-automatically","title":"Improving the Validity of Automatically Generated Feedback via Reinforcement Learning","date":"2024-03-02","arxiv_id":"2403.01304","repositories_listed":1,"syntology":{"n":13,"n_ran":9,"n_unverified":4,"n_pointer_only":13}},{"url":"/paper/clarify-improving-model-robustness-with","title":"Clarify: Improving Model Robustness With Natural Language Corrections","date":"2024-02-06","arxiv_id":"2402.03715","repositories_listed":1,"syntology":{"n":7,"n_ran":5,"n_unverified":2,"n_pointer_only":0}}],"syntology_records":9,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}