Methods › Natural Language Processing › Language Models › CodeGen › Papers where code ran, page 1
CodeGen
Papers archive 2025-07-28
archive papers tagged: 25 · with a code link: 19 · where Syntology ran a sample: 10 (7 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (10 of 25 tagged: 7 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 1: papers 1 to 10 of the 10 tagged papers where Syntology ran at least one harvested sample (7 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
HumanEval on Latest GPT Models -- 2024 20 Feb 2024 · 1 repository · arXiv:2402.14852Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion 17 Oct 2023 · 1 repository · arXiv:2310.11248Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models 31 Aug 2023 · 1 repository · arXiv:2308.16458Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
How Effective Are Neural Networks for Fixing Security Vulnerabilities 29 May 2023 · 1 repository · arXiv:2305.18607Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Prompting with Pseudo-Code Instructions 19 May 2023 · 2 repositories · arXiv:2305.11790Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 13 harvested samples)
-
Learning Performance-Improving Code Edits 15 Feb 2023 · 2 repositories · arXiv:2302.07867Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 11 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
ReCode: Robustness Evaluation of Code Generation Models 20 Dec 2022 · 3 repositories · arXiv:2212.10264Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Benchmarking Large Language Models for Automated Verilog RTL Code Generation 13 Dec 2022 · 1 repository · arXiv:2212.11140Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation 17 Aug 2022 · 1 repository · arXiv:2208.08227Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis 25 Mar 2022 · 8 repositories · arXiv:2203.13474Syntology official (archive's flag): 3 ran · 10 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 1 pointer-only (licence)