{"url":"/sota/semantic-parsing-on-wikitablequestions","task":{"name":"Semantic Parsing","url":"/task/semantic-parsing","note":null},"dataset":{"name":"WikiTableQuestions","url":"/dataset/wikitablequestions"},"category":"Natural Language Processing","categories":["Natural Language Processing"],"category_note":null,"description":"**Semantic Parsing** is the task of transducing natural language utterances into formal meaning representations. The target meaning representations can be defined according to a wide variety of formalisms. This include linguistically-motivated semantic representations that are designed to capture the meaning of any sentence such as λ-calculus or the abstract meaning representations. Alternatively, for more task-driven approaches to Semantic Parsing, it is common for meaning representations to represent executable programs such as SQL queries, robotic commands, smart phone instructions, and even general-purpose programming languages like Python and Java.\n\n\n<span class=\"description-source\">Source: [Tranx: A Transition-based Neural Abstract Syntax Parser for Semantic Parsing and Code Generation ](https://arxiv.org/abs/1810.02720)</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Accuracy (Test)","Accuracy (Dev)","Accuracy","Test Accuracy"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Accuracy (Test)":"higher","Accuracy (Dev)":"higher","Accuracy":"higher","Test Accuracy":"higher"}},"counts":{"rows":22,"rows_with_code":20,"rows_with_paper_page":22,"rows_dated":22,"rows_using_additional_data":3},"rows":[{"rank_in_archive_order":1,"model":"ARTEMIS-DA","metrics":{"Accuracy (Test)":"80.8"},"uses_additional_data":false,"paper_date":"2024-12-18","paper":"/paper/advanced-reasoning-and-transformation-engine","paper_url":"https://arxiv.org/abs/2412.14146v3","paper_title":"ARTEMIS-DA: An Advanced Reasoning and Transformation Engine for Multi-Step Insight Synthesis in Data Analytics","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":2,"model":"TabLaP","metrics":{"Accuracy (Dev)":"/","Accuracy (Test)":"76.6"},"uses_additional_data":false,"paper_date":"2024-10-10","paper":"/paper/accurate-and-regret-aware-numerical-problem","paper_url":"https://arxiv.org/abs/2410.12846v2","paper_title":"Accurate and Regret-aware Numerical Problem Solver for Tabular Question Answering","code":"https://github.com/yxw-11/tablap","n_code_links":1,"syntology":null},{"rank_in_archive_order":3,"model":"SynTQA (GPT)","metrics":{"Accuracy":"65.2","Accuracy (Test)":"74.4"},"uses_additional_data":true,"paper_date":"2024-09-25","paper":"/paper/syntqa-synergistic-table-based-question","paper_url":"https://arxiv.org/abs/2409.16682v2","paper_title":"SynTQA: Synergistic Table-based Question Answering via Mixture of Text-to-SQL and E2E TQA","code":"https://github.com/siyue-zhang/SynTableQA","n_code_links":1,"syntology":null},{"rank_in_archive_order":4,"model":"Mix SC","metrics":{"Accuracy (Dev)":"/","Accuracy (Test)":"73.6"},"uses_additional_data":false,"paper_date":"2023-12-27","paper":"/paper/rethinking-tabular-data-understanding-with","paper_url":"https://arxiv.org/abs/2312.16702v1","paper_title":"Rethinking Tabular Data Understanding with Large Language Models","code":"https://github.com/Leolty/tablellm","n_code_links":1,"syntology":{"n_ran":10,"n_unverified":1,"n_samples":11,"n_pointer_only_licence":0}},{"rank_in_archive_order":5,"model":"SynTQA (RF)","metrics":{"Accuracy (Dev)":"/","Accuracy (Test)":"71.6"},"uses_additional_data":true,"paper_date":"2024-09-25","paper":"/paper/syntqa-synergistic-table-based-question","paper_url":"https://arxiv.org/abs/2409.16682v2","paper_title":"SynTQA: Synergistic Table-based Question Answering via Mixture of Text-to-SQL and E2E TQA","code":"https://github.com/siyue-zhang/SynTableQA","n_code_links":1,"syntology":null},{"rank_in_archive_order":6,"model":"CABINET","metrics":{"Accuracy (Dev)":"/","Accuracy (Test)":"69.1"},"uses_additional_data":false,"paper_date":"2024-02-02","paper":"/paper/cabinet-content-relevance-based-noise","paper_url":"https://arxiv.org/abs/2402.01155v3","paper_title":"CABINET: Content Relevance based Noise Reduction for Table Question Answering","code":"https://github.com/sohanpatnaik106/cabinet_qa","n_code_links":1,"syntology":{"n_ran":3,"n_unverified":2,"n_samples":5,"n_pointer_only_licence":5}},{"rank_in_archive_order":7,"model":"NormTab+TabSQLify","metrics":{"Accuracy (Test)":"68.63"},"uses_additional_data":false,"paper_date":"2024-06-25","paper":"/paper/normtab-improving-symbolic-reasoning-in-llms","paper_url":"https://arxiv.org/abs/2406.17961v2","paper_title":"NormTab: Improving Symbolic Reasoning in LLMs Through Tabular Data Normalization","code":"https://github.com/mahadi-nahid/NormTab","n_code_links":1,"syntology":null},{"rank_in_archive_order":8,"model":"Chain-of-Table","metrics":{"Accuracy (Dev)":"/","Accuracy (Test)":"67.31"},"uses_additional_data":false,"paper_date":"2024-01-09","paper":"/paper/chain-of-table-evolving-tables-in-the","paper_url":"https://arxiv.org/abs/2401.04398v2","paper_title":"Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding","code":"https://github.com/cyqiq/multicot","n_code_links":2,"syntology":{"n_ran":6,"n_unverified":2,"n_samples":8,"n_pointer_only_licence":0}},{"rank_in_archive_order":9,"model":"Tab-PoT","metrics":{"Accuracy (Dev)":"/","Accuracy (Test)":"66.78"},"uses_additional_data":false,"paper_date":"2024-06-14","paper":"/paper/efficient-prompting-for-llm-based-generative","paper_url":"https://arxiv.org/abs/2406.10382v3","paper_title":"Efficient Prompting for LLM-based Generative Internet of Things","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":10,"model":"Dater","metrics":{"Accuracy (Dev)":"64.8","Accuracy (Test)":"65.9"},"uses_additional_data":false,"paper_date":"2023-01-31","paper":"/paper/large-language-models-are-versatile","paper_url":"https://arxiv.org/abs/2301.13808v3","paper_title":"Large Language Models are Versatile Decomposers: Decompose Evidence and Questions for Table-based Reasoning","code":"https://github.com/alibabaresearch/damo-convai","n_code_links":2,"syntology":null},{"rank_in_archive_order":11,"model":"LEVER","metrics":{"Accuracy (Dev)":"64.6","Accuracy (Test)":"65.8"},"uses_additional_data":false,"paper_date":"2023-02-16","paper":"/paper/lever-learning-to-verify-language-to-code","paper_url":"https://arxiv.org/abs/2302.08468v3","paper_title":"LEVER: Learning to Verify Language-to-Code Generation with Execution","code":"https://github.com/niansong1996/lever","n_code_links":1,"syntology":{"n_ran":18,"n_unverified":4,"n_samples":22,"n_pointer_only_licence":0}},{"rank_in_archive_order":12,"model":"TabSQLify (col+row)","metrics":{"Accuracy (Test)":"64.7"},"uses_additional_data":false,"paper_date":"2024-04-15","paper":"/paper/tabsqlify-enhancing-reasoning-capabilities-of","paper_url":"https://arxiv.org/abs/2404.10150v1","paper_title":"TabSQLify: Enhancing Reasoning Capabilities of LLMs Through Table Decomposition","code":"https://github.com/mahadi-nahid/tabsqlify","n_code_links":2,"syntology":{"n_ran":7,"n_unverified":8,"n_samples":15,"n_pointer_only_licence":15}},{"rank_in_archive_order":13,"model":"Binder","metrics":{"Accuracy (Dev)":"65.0","Accuracy (Test)":"64.6"},"uses_additional_data":false,"paper_date":"2022-10-06","paper":"/paper/binding-language-models-in-symbolic-languages","paper_url":"https://arxiv.org/abs/2210.02875v2","paper_title":"Binding Language Models in Symbolic Languages","code":"https://github.com/xlang-ai/binder","n_code_links":4,"syntology":{"n_ran":2,"n_unverified":1,"n_samples":3,"n_pointer_only_licence":0}},{"rank_in_archive_order":14,"model":"OmniTab-Large","metrics":{"Accuracy (Dev)":"62.5","Accuracy (Test)":"63.3"},"uses_additional_data":false,"paper_date":"2022-07-08","paper":"/paper/omnitab-pretraining-with-natural-and-1","paper_url":"https://arxiv.org/abs/2207.03637v1","paper_title":"OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering","code":"https://github.com/jzbjyb/omnitab","n_code_links":1,"syntology":{"n_ran":1,"n_unverified":1,"n_samples":2,"n_pointer_only_licence":2}},{"rank_in_archive_order":15,"model":"NormTab (Targeted) + SQL","metrics":{"Accuracy (Test)":"61.20"},"uses_additional_data":false,"paper_date":"2024-06-25","paper":"/paper/normtab-improving-symbolic-reasoning-in-llms","paper_url":"https://arxiv.org/abs/2406.17961v2","paper_title":"NormTab: Improving Symbolic Reasoning in LLMs Through Tabular Data Normalization","code":"https://github.com/mahadi-nahid/NormTab","n_code_links":1,"syntology":null},{"rank_in_archive_order":16,"model":"ReasTAP-Large","metrics":{"Accuracy (Dev)":"59.7","Accuracy (Test)":"58.7"},"uses_additional_data":false,"paper_date":"2022-10-22","paper":"/paper/reastap-injecting-table-reasoning-skills","paper_url":"https://arxiv.org/abs/2210.12374v1","paper_title":"ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning Examples","code":"https://github.com/yale-lily/reastap","n_code_links":1,"syntology":{"n_ran":8,"n_unverified":3,"n_samples":11,"n_pointer_only_licence":0}},{"rank_in_archive_order":17,"model":"TAPEX-Large","metrics":{"Accuracy (Dev)":"57.0","Accuracy (Test)":"57.5"},"uses_additional_data":false,"paper_date":"2021-07-16","paper":"/paper/tapex-table-pre-training-via-learning-a","paper_url":"https://arxiv.org/abs/2107.07653v3","paper_title":"TAPEX: Table Pre-training via Learning a Neural SQL Executor","code":"https://github.com/microsoft/Table-Pretraining","n_code_links":4,"syntology":null},{"rank_in_archive_order":18,"model":"MAPO + TABERTLarge (K = 3)","metrics":{"Accuracy (Dev)":"52.2","Accuracy (Test)":"51.8"},"uses_additional_data":false,"paper_date":"2020-05-17","paper":"/paper/tabert-pretraining-for-joint-understanding-of","paper_url":"https://arxiv.org/abs/2005.08314v1","paper_title":"TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data","code":"https://github.com/facebookresearch/tabert","n_code_links":1,"syntology":null},{"rank_in_archive_order":19,"model":"T5-3b(UnifiedSKG)","metrics":{"Accuracy (Dev)":"50.65","Accuracy (Test)":"49.29"},"uses_additional_data":false,"paper_date":"2022-01-16","paper":"/paper/unifiedskg-unifying-and-multi-tasking","paper_url":"https://arxiv.org/abs/2201.05966v3","paper_title":"UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models","code":"https://github.com/hkunlp/unifiedskg","n_code_links":1,"syntology":{"n_ran":2,"n_unverified":3,"n_samples":5,"n_pointer_only_licence":0}},{"rank_in_archive_order":20,"model":"TAPAS-Large (pre-trained on SQA)","metrics":{"Accuracy (Dev)":"/","Accuracy (Test)":"48.8"},"uses_additional_data":true,"paper_date":"2020-04-05","paper":"/paper/tapas-weakly-supervised-table-parsing-via-pre","paper_url":"https://arxiv.org/abs/2004.02349v2","paper_title":"TAPAS: Weakly Supervised Table Parsing via Pre-training","code":"https://github.com/huggingface/transformers","n_code_links":8,"syntology":{"n_ran":13,"n_unverified":1,"n_samples":14,"n_pointer_only_licence":0}},{"rank_in_archive_order":21,"model":"Structured Attention","metrics":{"Accuracy (Dev)":"43.7","Accuracy (Test)":"44.5"},"uses_additional_data":false,"paper_date":"2019-09-09","paper":"/paper/learning-semantic-parsers-from-denotations","paper_url":"https://arxiv.org/abs/1909.04165v1","paper_title":"Learning Semantic Parsers from Denotations with Latent Structured Alignments and Abstract Programs","code":"https://github.com/berlino/weaksp_em19","n_code_links":1,"syntology":null},{"rank_in_archive_order":22,"model":"SynTQA (Oracle)","metrics":{"Test Accuracy":"77.5"},"uses_additional_data":false,"paper_date":"2024-09-25","paper":"/paper/syntqa-synergistic-table-based-question","paper_url":"https://arxiv.org/abs/2409.16682v2","paper_title":"SynTQA: Synergistic Table-based Question Answering via Mixture of Text-to-SQL and E2E TQA","code":"https://github.com/siyue-zhang/SynTableQA","n_code_links":1,"syntology":null}],"since_archive":{"claim":"Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the archive rows.","extraction_file_present":true,"measurement":{"test_papers":883,"papers_with_output":881,"judged_true":108,"judged":110,"wilson95_lower":0.9361,"measured_on":"2026-09-24","frozen_commit":"0e3de0df94"},"measurement_note":"blind adjudication of accepted entries on a held-out split of archive papers, rules frozen before the test","coverage":{"sentence":"Syntology has checked 6,885 of the 9,623 papers on this site that are newer than the archive; results from the others appear after they are checked.","complete":false,"papers_newer_than_archive":9623,"papers_checked":6885,"papers_extracted_not_yet_verified":0,"boards_without_verdict":2,"papers_not_yet_extracted":2737},"order":"newest first by month (arXiv date, else the arXiv-id month), then arXiv id descending","columns":[],"entries":[]},"syntology":{"read_at":"2026-09-25T09:33:49+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":10,"rows_with_any_sample_ran":10,"distinct_papers_with_graph_line":10,"distinct_papers_with_any_sample_ran":10,"samples_over_distinct_papers":{"n_ran":70,"n_unverified":26,"n_samples":96,"n_pointer_only_licence":22,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":70,"n_unverified":26,"n_samples":96,"n_pointer_only_licence":22,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}