{"url":"/dataset/safim","name":"SAFIM","full_name":"Syntax-Aware Fill-In-the-Middle","description_markdown":"Syntax-Aware Fill-in-the-Middle (SAFIM) is a benchmark for evaluating Large Language Models (LLMs) on the code Fill-in-the-Middle (FIM) task. SAFIM has three subtasks: Algorithmic Block Completion, Control-Flow Expression Completion, and API Function Call Completion. SAFIM is sourced from code submitted from April 2022 to January 2023 to minimize the impact of data contamination on evaluation results.\r\n\r\n- Authors: [Linyuan Gong](https://gonglinyuan.com), [Sida Wang](https://www.sidaw.xyz/), [Mostafa Elhoushi](https://www.linkedin.com/in/mostafaelhoushi), [Alvin Cheung](https://people.eecs.berkeley.edu/~akcheung/)\r\n- Paper: [https://arxiv.org/abs/2403.04814](https://arxiv.org/abs/2403.04814)\r\n- Huggingface Dataset: [https://huggingface.co/datasets/gonglinyuan/safim](https://huggingface.co/datasets/gonglinyuan/safim)\r\n- Leaderboard: [https://safimbenchmark.com](https://safimbenchmark.com)\r\n- Code & Submission Instructions: [https://github.com/gonglinyuan/safim](https://github.com/gonglinyuan/safim\r\n)\r\n\r\nThe SAFIM benchmark is partially derived from problem descriptions and code solutions from https://codeforces.com. According to the license of CodeForces, you may publish the texts of Codeforces problems in any open sources, but you must preserve a direct link to the site.","description_withheld":null,"homepage":"https://safimbenchmark.com","introduced_date":"2024-03-07","introduced_date_note":null,"introduced_by":{"paper":"/paper/evaluation-of-llms-on-syntax-aware-code-fill","title":"Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks","first_author":"Linyuan Gong","url":null},"license":{"name":"CC-BY-4.0","url":"https://creativecommons.org/licenses/by/4.0/legalcode.en"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Code Generation","url":"/task/code-generation","datasets_with_task":"/datasets/task/code-generation"},{"name":"Code Completion","url":"/task/code-completion","datasets_with_task":"/datasets/task/code-completion"},{"name":"Text-to-Code Generation","url":"/task/text-to-code-generation","datasets_with_task":"/datasets/task/text-to-code-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SAFIM"],"data_loaders":[],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/code-completion-on-safim","task":"Code Completion","dataset_variant":"SAFIM","rows":15,"metrics":["Average","Algorithmic","Control","API"],"first_row_in_archive_order":{"model":"deepseek-coder-33b-base","paper":"/paper/evaluation-of-llms-on-syntax-aware-code-fill","metrics":{"API":"75.16","Algorithmic":"60.78","Average":"69.01","Control":"71.10"},"code_links":[{"title":"gonglinyuan/safim","url":"https://github.com/gonglinyuan/safim"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/evaluation-of-llms-on-syntax-aware-code-fill","title":"Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks","date":"2024-03-07","rows_on_this_dataset":15,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}