Datasets › Spider-Realistic › Papers where code ran, page 1

Spider-Realistic

Papers archive 2025-07-28

papers with a benchmark row: 23 · with a code link: 21 · where Syntology ran a sample: 14 (10 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (14 of 23 with a benchmark row: 10 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument)

Syntology We ran code from the paper's repository; we did not run it on this dataset or check it against this dataset's benchmarks.

Page 1 of 1: papers 1 to 14 of the 14 papers with a benchmark row here where Syntology ran at least one harvested sample (10 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 91. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
MSc-SQL: Multi-Sample Critiquing Small Language Models For Text-To-SQL Translation 1 1 16 Oct 2024 official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
PET-SQL: A Prompt-Enhanced Two-Round Refinement of Text-to-SQL with Cross-consistency 1 1 13 Mar 2024 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (9 pointer-only for licence)
Knowledge-to-SQL: Enhancing SQL Generation with Data Expert LLM 1 1 18 Feb 2024 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified
Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation 1 1 29 Aug 2023 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
C3: Zero-shot Text-to-SQL with ChatGPT 1 1 14 Jul 2023 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction 1 1 21 Apr 2023 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
LEVER: Learning to Verify Language-to-Code Generation with Execution 1 2 16 Feb 2023 official (archive's flag): 18 ran · 18 ran (of which 0 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 1 violated, 17 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (1 pointer-only for licence)
RASAT: Integrating Relational Structures into Pretrained Seq2Seq Model for Text-to-SQL 1 3 14 May 2022 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (1 pointer-only for licence)
SADGA: Structure-Aware Dual Graph Aggregation Network for Text-to-SQL 1 2 1 Nov 2021 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (7 pointer-only for licence)
PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models 3 4 10 Sep 2021 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing 1 1 29 Sep 2020 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 1 unverified (7 pointer-only for licence)
RYANSQL: Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases 1 1 7 Apr 2020 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified
RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers 4 1 10 Nov 2019 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (4 pointer-only for licence)
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task 6 1 24 Sep 2018 official (archive's flag): 3 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 3 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence)