Browse State-of-the-Art › Code Generation › Papers, page 9
Code Generation
Papers archive 2025-07-28
archive papers tagged: 1,697 · with a code link: 745 · where Syntology ran a sample: 280 (238 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (280 of 1,697 tagged: 238 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 9 of 17: papers 801 to 900 of 1,697, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Rethinking the effects of data contamination in Code Intelligence3 Jun 2025 0 repositories listed
-
Flow2Code: Evaluating Large Language Models for Flowchart-based Code Generation Capability2 Jun 2025 0 repositories listed
-
ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code2 Jun 2025 0 repositories listed
-
SALAD: Systematic Assessment of Machine Unlearing on LLM-Aided Hardware Design2 Jun 2025 0 repositories listed
-
Legal Compliance Evaluation of Smart Contracts Generated By Large Language Models1 Jun 2025 0 repositories listed
-
CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval31 May 2025 0 repositories listed
-
A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming30 May 2025 0 repositories listed
-
Cascading Adversarial Bias from Injection to Distillation in Language Models30 May 2025 0 repositories listed
-
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX30 May 2025 0 repositories listed
-
HardTests: Synthesizing High-Quality Test Cases for LLM Coding30 May 2025 0 repositories listed
-
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation30 May 2025 0 repositories listed
-
Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards30 May 2025 0 repositories listed
-
Enhancing LLM-Based Code Generation with Complexity Metrics: A Feedback-Driven Approach29 May 2025 0 repositories listed
-
Infinite-Instruct: Synthesizing Scaling Code instruction Data with Bidirectional Synthesis and Static Verification29 May 2025 0 repositories listed
-
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving29 May 2025 0 repositories listed
-
DeepRTL2: A Versatile Model for RTL-Related Tasks28 May 2025 0 repositories listed
-
HiLDe: Intentional Code Generation via Human-in-the-Loop Decoding28 May 2025 0 repositories listed
-
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks27 May 2025 0 repositories listed
-
Rendering-Aware Reinforcement Learning for Vector Graphics Generation27 May 2025 0 repositories listed
-
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation26 May 2025 0 repositories listed
-
Large Language Models for IT Automation Tasks: Are We There Yet?26 May 2025 0 repositories listed
-
Large Language Models in Code Co-generation for Safe Autonomous Vehicles26 May 2025 0 repositories listed
-
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment25 May 2025 0 repositories listed
-
Architectures of Error: A Philosophical Inquiry into AI and Human Code Generation25 May 2025 0 repositories listed
-
Autocomp: LLM-Driven Code Optimization for Tensor Accelerators24 May 2025 0 repositories listed
-
24 May 2025 0 repositories listed Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?24 May 2025 0 repositories listed
-
HD-PiSSA: High-Rank Distributed Orthogonal Adaptation24 May 2025 0 repositories listed
-
PromptWise: Online Learning for Cost-Aware Prompt Assignment in Generative Models24 May 2025 0 repositories listed
-
Evaluating the Energy-Efficiency of the Code Generated by LLMs23 May 2025 0 repositories listed
-
PPT: A Process-based Preference Learning Framework for Self Improving Table Question Answering Models23 May 2025 0 repositories listed
-
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks22 May 2025 0 repositories listed
-
MAPS: A Multilingual Benchmark for Global Agent Performance and Security21 May 2025 0 repositories listed
-
SIMCOPILOT: Evaluating Large Language Models for Copilot-Style Code Generation21 May 2025 0 repositories listed
-
Towards a Science of Causal Interpretability in Deep Learning for Software Engineering21 May 2025 0 repositories listed
-
Cheaper, Better, Faster, Stronger: Robust Text-to-SQL without Chain-of-Thought or Fine-Tuning20 May 2025 0 repositories listed
-
From Reasoning to Code: GRPO Optimization for Underrepresented Languages20 May 2025 0 repositories listed
-
Knowledge Graph Based Repository-Level Code Generation20 May 2025 0 repositories listed
-
Self-Evolving Curriculum for LLM Reasoning20 May 2025 0 repositories listed
-
Text Generation Beyond Discrete Token Sampling20 May 2025 0 repositories listed
-
AutoGEEval: A Multimodal and Automated Framework for Geospatial Code Generation on GEE with Large Language Models19 May 2025 0 repositories listed
-
19 May 2025 0 repositories listed
-
On-Policy Optimization with Group Equivalent Preference for Multi-Programming Language Understanding19 May 2025 0 repositories listed
-
Selective Code Generation for Functional Guarantees19 May 2025 0 repositories listed
-
Understanding Complexity in VideoQA via Visual Program Generation19 May 2025 0 repositories listed
-
EVALOOP: Assessing LLM Robustness in Programming from a Self-consistency Perspective18 May 2025 0 repositories listed
-
SOCIA: An End-to-End Agentic Framework for Automated Cyber-Physical-Social Simulator Generation17 May 2025 0 repositories listed
-
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations16 May 2025 0 repositories listed
-
Code-Driven Planning in Grid Worlds with Large Language Models15 May 2025 0 repositories listed
-
CRPE: Expanding The Reasoning Capability of Large Language Model for Code Generation15 May 2025 0 repositories listed
-
Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models15 May 2025 0 repositories listed
-
Agent-as-a-Service based on Agent Network13 May 2025 0 repositories listed
-
CAD-Coder:Text-Guided CAD Files Code Generation13 May 2025 0 repositories listed
-
Evaluating LLM Metrics Through Real-World Capabilities13 May 2025 0 repositories listed
-
Generalizing Large Language Model Usability Across Resource-Constrained13 May 2025 0 repositories listed
-
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation13 May 2025 0 repositories listed
-
One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models12 May 2025 0 repositories listed
-
Code Retrieval for MILP Instance Generation11 May 2025 0 repositories listed
-
RTL++: Graph-enhanced LLM for RTL Code Generation11 May 2025 0 repositories listed
-
CodeMixBench: Evaluating Large Language Models on Code Generation with Code-Mixed Prompts8 May 2025 0 repositories listed
-
A Proposal for Evaluating the Operational Risk for ChatBots based on Large Language Models7 May 2025 0 repositories listed
-
LLM Code Customization with Visual Results: A Benchmark on TikZ7 May 2025 0 repositories listed
-
YABLoCo: Yet Another Benchmark for Long Context Code Generation7 May 2025 0 repositories listed
-
Capability-Driven Skill Generation with LLMs: A RAG-Based Approach for Reusing Existing Libraries and Interfaces6 May 2025 0 repositories listed
-
MARCO: Multi-Agent Code Optimization with Real-Time Knowledge Integration for High-Performance Computing6 May 2025 0 repositories listed
-
Scratch Copilot: Supporting Youth Creative Coding with AI6 May 2025 0 repositories listed
-
STORY2GAME: Generating (Almost) Everything in an Interactive Fiction Game6 May 2025 0 repositories listed
-
AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks5 May 2025 0 repositories listed
-
QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic Approach4 May 2025 0 repositories listed
-
A Rusty Link in the AI Supply Chain: Detecting Evil Configurations in Model Repositories2 May 2025 0 repositories listed
-
CHORUS: Zero-shot Hierarchical Retrieval and Orchestration for Generating Linear Programming Code2 May 2025 0 repositories listed
-
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding2 May 2025 0 repositories listed
-
Assessing LLM code generation quality through path planning tasks30 Apr 2025 0 repositories listed
-
ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement29 Apr 2025 0 repositories listed
-
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation29 Apr 2025 0 repositories listed
-
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges29 Apr 2025 0 repositories listed
-
SecRepoBench: Benchmarking LLMs for Secure Code Generation in Real-World Repositories29 Apr 2025 0 repositories listed
-
Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs29 Apr 2025 0 repositories listed
-
The Hidden Risks of LLM-Generated Web Application Code: A Security-Centric Evaluation of Code Generation Capabilities in Large Language Models29 Apr 2025 0 repositories listed
-
An Automated Reinforcement Learning Reward Design Framework with Large Language Model for Cooperative Platoon Coordination28 Apr 2025 0 repositories listed
-
Evaluating Grounded Reasoning by Code-Assisted Large Language Models for Mathematics24 Apr 2025 0 repositories listed
-
High-Fidelity And Complex Test Data Generation For Real-World SQL Code Generation Services24 Apr 2025 0 repositories listed
-
ClarifyCoder: Clarification-Aware Fine-Tuning for Programmatic Problem Solving23 Apr 2025 0 repositories listed
-
EduBot -- Can LLMs Solve Personalized Learning and Programming Assignments?23 Apr 2025 0 repositories listed
-
A Large-scale Class-level Benchmark Dataset for Code Generation with LLMs22 Apr 2025 0 repositories listed
-
Insights from Verification: Training a Verilog Generation LLM with Reinforcement Learning with Testbench Feedback22 Apr 2025 0 repositories listed
-
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation22 Apr 2025 0 repositories listed
-
Empowering AI to Generate Better AI Code: Guided Generation of Deep Learning Projects with LLMs21 Apr 2025 0 repositories listed
-
Evaluating Code Generation of LLMs in Advanced Computer Science Problems21 Apr 2025 0 repositories listed
-
Improving RL Exploration for LLM Reasoning through Retrospective Replay19 Apr 2025 0 repositories listed
-
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation18 Apr 2025 0 repositories listed
-
Do Prompt Patterns Affect Code Quality? A First Empirical Assessment of ChatGPT-Generated Code18 Apr 2025 0 repositories listed
-
Towards End-to-End Network Intent Management with Large Language Models18 Apr 2025 0 repositories listed
-
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation17 Apr 2025 0 repositories listed
-
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo17 Apr 2025 0 repositories listed
-
Rethinking the Generation of High-Quality CoT Data from the Perspective of LLM-Adaptive Question Difficulty Grading16 Apr 2025 0 repositories listed
-
Themisto: Jupyter-Based Runtime Benchmark16 Apr 2025 0 repositories listed
-
The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance14 Apr 2025 0 repositories listed
-
Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation13 Apr 2025 0 repositories listed
-
Iterative Self-Training for Code Generation via Reinforced Re-Ranking13 Apr 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.