Datasets › Spider-Realistic

Spider-Realistic

Introduced by Xiang Deng et al. in Structure-Grounded Pretraining for Text-to-SQL24 Oct 2020 archive 2025-07-28

Spider dataset is used for evaluation in the paper "Structure-Grounded Pretraining for Text-to-SQL". The dataset is created based on the dev split of the Spider dataset (2020-06-07 version from https://yale-lily.github.io/spider). We manually modified the original questions to remove the explicit mention of column names while keeping the SQL queries unchanged to better evaluate the model's capability in aligning the NL utterance and the DB schema. For more details, please check our paper at https://arxiv.org/abs/2010.12773.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text-To-SQL spider XiYan-SQL Execution Accuracy (Test) 89.65 A Preview of XiYan-SQL: A Multi-Generator Ensemble... XGenerationLab/XiYan-SQL +4 20 Compare
Semantic Parsing spider RESDSQL-3B + NatSQL Accuracy 84.1 RESDSQL: Decoupling Schema Linking and Skeleton Parsing... ruckbreasoning/resdsql 10 Compare
Text-To-SQL SPIDER T5-3B+PICARD Exact Match Accuracy (in Dev) 75.5 PICARD: Parsing Incrementally for Constrained... servicenow/picard +2 4 Compare

Papers archive 2025-07-28

23 shown of 23 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 91. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
A Preview of XiYan-SQL: A Multi-Generator Ensemble Framework for Text-to-SQL 5 1 13 Nov 2024 not harvested
Learning Metadata-Agnostic Representations for Text-to-SQL In-Context Example Selection 0 1 17 Oct 2024 not harvested
MSc-SQL: Multi-Sample Critiquing Small Language Models For Text-To-SQL Translation 1 1 16 Oct 2024 ran 9 of 9 samples (0 unverified)
DataGpt-SQL-7B: An Open-Source Language Model for Text-to-SQL 0 2 24 Sep 2024 not harvested
PET-SQL: A Prompt-Enhanced Two-Round Refinement of Text-to-SQL with Cross-consistency 1 1 13 Mar 2024 ran 6 of 9 samples (3 unverified; 9 pointer-only for licence)
Knowledge-to-SQL: Enhancing SQL Generation with Data Expert LLM 1 1 18 Feb 2024 ran 2 of 3 samples (1 unverified)
Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation 1 1 29 Aug 2023 ran 1 of 3 samples (2 unverified)
C3: Zero-shot Text-to-SQL with ChatGPT 1 1 14 Jul 2023 ran 2 of 2 samples (0 unverified)
T5-SR: A Unified Seq-to-Seq Decoding Strategy for Semantic Parsing 1 1 14 Jun 2023 not harvested
Improving Generalization in Language Model-Based Text-to-SQL Semantic Parsing: Two Simple Semantic Boundary-Based Techniques 1 1 27 May 2023 not harvested
DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction 1 1 21 Apr 2023 ran 0 of 4 samples (4 unverified)
LEVER: Learning to Verify Language-to-Code Generation with Execution 1 2 16 Feb 2023 ran 6 of 22 samples (16 unverified)
RESDSQL: Decoupling Schema Linking and Skeleton Parsing for Text-to-SQL 1 2 12 Feb 2023 ran 0 of 1 samples (1 unverified)
Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing 1 2 18 Jan 2023 not harvested
RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model for Text-to-SQL 1 3 14 May 2022 ran 0 of 9 samples (9 unverified)
SADGA: Structure-Aware Dual Graph Aggregation Network for Text-to-SQL 1 2 1 Nov 2021 ran 5 of 7 samples (2 unverified; 7 pointer-only for licence)
PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models 3 4 10 Sep 2021 ran 4 of 7 samples (3 unverified)
Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training 3 2 18 Dec 2020 not harvested
GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing 1 1 29 Sep 2020 ran 6 of 7 samples (1 unverified; 7 pointer-only for licence)
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data 1 1 17 May 2020 not harvested
RYANSQL: Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases 1 1 7 Apr 2020 ran 1 of 2 samples (1 unverified)
RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers 4 1 10 Nov 2019 ran 4 of 7 samples (3 unverified)
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task 6 1 24 Sep 2018 ran 7 of 8 samples (1 unverified; 3 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • spider
  • Spider-Realistic

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections