Browse State-of-the-Art › HumanEval › Papers, page 3
HumanEval
Papers archive 2025-07-28
archive papers tagged: 264 · with a code link: 135 · where Syntology ran a sample: 76 (65 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (76 of 264 tagged: 65 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument)
Page 3 of 3: papers 201 to 264 of 264, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Context-Augmented Code Generation Using Programming Knowledge Graphs9 Oct 2024 0 repositories listed
-
AIME: AI System Optimization via Multiple LLM Evaluators4 Oct 2024 0 repositories listed
-
Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity24 Sep 2024 0 repositories listed
-
GRIN: GRadient-INformed MoE18 Sep 2024 0 repositories listed
-
RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation15 Sep 2024 0 repositories listed
-
CPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning Tasks13 Sep 2024 0 repositories listed
-
𝕌𝕊ℂ𝔻: Improving Code Generation of LLMs by Uncertainty-Aware Selective Contrastive Decoding9 Sep 2024 0 repositories listed
-
Arctic-SnowCoder: Demystifying High-Quality Data in Code Pretraining3 Sep 2024 0 repositories listed
-
CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution23 Aug 2024 0 repositories listed
-
DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation23 Aug 2024 0 repositories listed
-
AutoTest: Evolutionary Code Solution Selection with Test Cases22 Aug 2024 0 repositories listed
-
Concept Distillation from Strong to Weak Models via Hypotheses-to-Theories Prompting18 Aug 2024 0 repositories listed
-
Threshold Filtering Packing for Supervised Fine-Tuning: Training Related Samples within Packs18 Aug 2024 0 repositories listed
-
CodeMirage: Hallucinations in Code Generated by Large Language Models14 Aug 2024 0 repositories listed
-
CREST: Effectively Compacting a Datastore For Retrieval-Based Speculative Decoding8 Aug 2024 0 repositories listed
-
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models30 Jul 2024 0 repositories listed
-
Discrete Flow Matching22 Jul 2024 0 repositories listed
-
MaPPing Your Model: Assessing the Impact of Adversarial Attacks on LLM-based Programming Assistants12 Jul 2024 0 repositories listed
-
Brevity is the soul of wit: Pruning long files for code generation29 Jun 2024 0 repositories listed
-
Towards Large Language Model Aided Program Refinement26 Jun 2024 0 repositories listed
-
Qiskit HumanEval: An Evaluation Benchmark For Quantum Code Generative Models20 Jun 2024 0 repositories listed
-
Code-Optimise: Self-Generated Preference Data for Correctness and Efficiency18 Jun 2024 0 repositories listed
-
Reactor Mk.1 performances: MMLU, HumanEval and BBH test results15 Jun 2024 0 repositories listed
-
PLUM: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases11 Jun 2024 0 repositories listed
-
Validating LLM-Generated Programs with Metamorphic Prompt Testing11 Jun 2024 0 repositories listed
-
Does your data spark joy? Performance gains from domain upsampling at the end of training5 Jun 2024 0 repositories listed
-
Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code Generation30 May 2024 0 repositories listed
-
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths30 May 2024 0 repositories listed
-
Kotlin ML Pack: Technical Report29 May 2024 0 repositories listed
-
Qiskit Code Assistant: Training LLMs for generating Quantum Computing Code29 May 2024 0 repositories listed
-
On the Limitations of Embedding Based Methods for Measuring Functional Correctness for Code Generation26 Apr 2024 0 repositories listed
-
BASS: Batched Attention-optimized Speculative Sampling24 Apr 2024 0 repositories listed
-
NExT: Teaching Large Language Models to Reason about Code Execution23 Apr 2024 0 repositories listed
-
Low-Cost Language Models: Survey and Performance Evaluation on Python Code Generation17 Apr 2024 0 repositories listed
-
Exploring and Evaluating Hallucinations in LLM-Powered Code Generation1 Apr 2024 0 repositories listed
-
Reasoning Runtime Behavior of a Program with LLM: How Far Are We?25 Mar 2024 0 repositories listed
-
CodeShell Technical Report23 Mar 2024 0 repositories listed
-
23 Mar 2024 0 repositories listed
-
CodingTeachLLM: Empowering LLM's Coding Ability via AST Prior Knowledge13 Mar 2024 0 repositories listed
-
Software Vulnerability and Functionality Assessment using LLMs13 Mar 2024 0 repositories listed
-
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code12 Mar 2024 0 repositories listed
-
Test-Driven Development for Code Generation21 Feb 2024 0 repositories listed
-
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models13 Feb 2024 0 repositories listed
-
NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness29 Jan 2024 0 repositories listed
-
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs11 Jan 2024 0 repositories listed
-
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs8 Jan 2024 0 repositories listed
-
A Review of Repository Level Prompting for LLMs15 Dec 2023 0 repositories listed
-
Decoding Data Quality via Synthetic Corruptions: Embedding-guided Pruning of Code Data5 Dec 2023 0 repositories listed
-
Past as a Guide: Leveraging Retrospective Learning for Python Code Completion13 Nov 2023 0 repositories listed
-
Bridging Code Semantic and LLMs: Semantic Chain-of-Thought Prompting for Code Generation16 Oct 2023 0 repositories listed
-
CodeFuse-13B: A Pretrained Multi-lingual Code Large Language Model10 Oct 2023 0 repositories listed
-
The Program Testing Ability of Large Language Models for Code9 Oct 2023 0 repositories listed
-
LORD: Low Rank Decomposition Of Monolingual Code LLMs For One-Shot Compression25 Sep 2023 0 repositories listed
-
Can Programming Languages Boost Each Other via Instruction Tuning?31 Aug 2023 0 repositories listed
-
CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation17 Aug 2023 0 repositories listed
-
PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback27 Jul 2023 0 repositories listed
-
Textbooks Are All You Need20 Jun 2023 0 repositories listed
-
SelfEvolve: A Code Evolution Framework via Large Language Models5 Jun 2023 0 repositories listed
-
Structured Chain-of-Thought Prompting for Code Generation11 May 2023 0 repositories listed
-
Stochastic Code Generation14 Apr 2023 0 repositories listed
-
Large Language Models Meet NL2Code: A Survey19 Dec 2022 0 repositories listed
-
20 Nov 2022 0 repositories listed
-
Piloting Copilot, Codex, and StarCoder2: Hot Temperature, Cold Prompts, or Black Magic?26 Oct 2022 0 repositories listed
-
Interactive Code Generation via Test-Driven User-Intent Formalization11 Aug 2022 0 repositories listed