{"url":"/dataset/sede","name":"SEDE","full_name":"Stack Exchange Data Explorer","description_markdown":"**SEDE** is a dataset comprised of 12,023 complex and diverse SQL queries and their natural language titles and descriptions, written by real users of the Stack Exchange Data Explorer out of a natural interaction. These pairs contain a variety of real-world challenges which were rarely reflected so far in any other semantic parsing dataset. The goal of this dataset is to take a significant step towards evaluation of Text-to-SQL models in a real-world setting. Compared to other Text-to-SQL datasets, SEDE contains at least 10 times more SQL queries templates (queries after canonization and anonymization of values) than other datasets, and has the most diverse set of utterances and SQL queries (in terms of 3-grams) out of all single-domain datasets. SEDE introduces real-world challenges, such as under-specification, usage of parameters in queries, dates manipulation and more.","description_withheld":null,"homepage":"https://github.com/hirupert/sede","introduced_date":"2021-06-09","introduced_date_note":null,"introduced_by":{"paper":"/paper/text-to-sql-in-the-wild-a-naturally-occurring","title":"Text-to-SQL in the Wild: A Naturally-Occurring Dataset Based on Stack Exchange Data","first_author":"Moshe Hazoom","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Semantic Parsing","url":"/task/semantic-parsing","datasets_with_task":"/datasets/task/semantic-parsing"},{"name":"Text-To-SQL","url":"/task/text-to-sql","datasets_with_task":"/datasets/task/text-to-sql"}],"languages":[],"variants":["SEDE"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/hirupert/sede","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/sede","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":11,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/text-to-sql-on-sede","task":"Text-To-SQL","dataset_variant":"SEDE","rows":1,"metrics":["PCM-F1 (dev)","PCM-F1 (test)"],"first_row_in_archive_order":{"model":"T5-Large","paper":"/paper/text-to-sql-in-the-wild-a-naturally-occurring","metrics":{"PCM-F1 (dev)":"48.2","PCM-F1 (test)":"50.6"},"code_links":[{"title":"hirupert/sede","url":"https://github.com/hirupert/sede"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/text-to-sql-in-the-wild-a-naturally-occurring","title":"Text-to-SQL in the Wild: A Naturally-Occurring Dataset Based on Stack Exchange Data","date":"2021-06-09","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":0,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":5,"samples_ran":0,"samples_unverified":5,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}