Papers › Effective Test Generation Using Pre-trained Large Language Models and Mutation Testing

Effective Test Generation Using Pre-trained Large Language Models and Mutation Testing

31 Aug 2023arXiv:2308.16557links table onlyarchive 2025-07-28

Arghavan Moradi Dakhel, Amin Nikanjam, Vahid Majdinasab, Foutse Khomh, Michel C. Desmarais

The archive published only this paper's code-link row. Authors, date and abstract are from arXiv's metadata (CC0), read from the Kaggle arXiv metadata snapshot of 2026-09-12 where its title matched the archive's; the title is the archive's.

One of the critical phases in software development is software testing. Testing helps with identifying potential bugs and reducing maintenance costs. The goal of automated test generation tools is to ease the development of tests by suggesting efficient bug-revealing tests. Recently, researchers have leveraged Large Language Models (LLMs) of code to generate unit tests. While the code coverage of generated tests was usually assessed, the literature has acknowledged that the coverage is weakly correlated with the efficiency of tests in bug detection. To improve over this limitation, in this paper, we introduce MuTAP for improving the effectiveness of test cases generated by LLMs in terms of revealing bugs by leveraging mutation testing. Our goal is achieved by augmenting prompts with surviving mutants, as those mutants highlight the limitations of test cases in detecting bugs. MuTAP is capable of generating effective test cases in the absence of natural language descriptions of the Program Under Test (PUTs). We employ different LLMs within MuTAP and evaluate their performance on different benchmarks. Our results show that our proposed method is able to detect up to 28% more faulty human-written code snippets. Among these, 17% remained undetected by both the current state-of-the-art fully automated test generation tool (i.e., Pynguin) and zero-shot/few-shot learning approaches on LLMs. Furthermore, MuTAP achieves a Mutation Score (MS) of 93.57% on synthetic buggy code, outperforming all other approaches in our evaluation. Our findings suggest that although LLMs can serve as a useful tool to generate test cases, they require specific post-processing steps to enhance the effectiveness of the generated test cases which may suffer from syntactic or functional errors and may be ineffective in detecting certain types of bugs and testing corner cases PUTs.

PaperPDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2308.16557")

Code

Syntology Ran 14 of 16 code samples harvested from 1 repository linked to this paper; 2 have no recorded run. Of those that ran: 14 ran with no contract checked.

By repository: official repository: 16 samples from 1 repository, 14 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

expertisemodel/mutap officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

16 samples harvested; 14 ran; 0 honoured the contract we drafted; 2 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

14ran
2unverified

Licence: 16 of the 16 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from expertisemodel/mutap. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

Mut_all_in_one expertisemodel/mutap/MuTAP/Merge_all_mut.py official repository ran no licence file found · pointer only · 3d640759c588109a · report
Write_into_pickle expertisemodel/mutap/Read_json_data.py official repository ran no licence file found · pointer only · e755516d4ffe111d · report
check_test_oracle_sematic expertisemodel/mutap/MuTAP/Sematic_err_correction.py official repository ran no licence file found · pointer only · a5bc80a559b5038c · report
directory_crawler expertisemodel/mutap/MuTAP/greedy_test_generator.py official repository ran fingerprinted no licence file found · pointer only · ce15f08dfed3da2d · report
fewshotMutantPromptGenerator expertisemodel/mutap/MuTAP/augmented_prompt.py official repository ran fingerprinted no licence file found · pointer only · 54591bb7dd8a8e60 · report
fewshotPromptGenerator expertisemodel/mutap/MuTAP/generate_test_oracle.py official repository ran fingerprinted no licence file found · pointer only · f43c94b84ab25a2c · report
initalPromptGenerator expertisemodel/mutap/MuTAP/generate_test_oracle.py official repository ran fingerprinted no licence file found · pointer only · 2a493a7d68af773f · report
load_model expertisemodel/mutap/llama_util/model_utils.py official repository ran no licence file found · pointer only · bfc6d0ec102ca108 · report
merge_asserts expertisemodel/mutap/MuTAP/Merge_all_mut.py official repository ran no licence file found · pointer only · 0bd08159faf7b105 · report
mutatePromptGenerator expertisemodel/mutap/MuTAP/augmented_prompt.py official repository ran fingerprinted no licence file found · pointer only · 10e879155a44ead2 · report
paths_list expertisemodel/mutap/MuTAP/Mutation_Score.py official repository ran no licence file found · pointer only · 4e2d817f03e60189 · report
read_csv expertisemodel/mutap/MuTAP/augmented_prompt.py official repository ran no licence file found · pointer only · 77d3223b6589a4e6 · report
separate_asserts expertisemodel/mutap/MuTAP/greedy_test_generator.py official repository ran no licence file found · pointer only · f7c84817041483b2 · report
synxPromptGenerator expertisemodel/mutap/MuTAP/generate_test_oracle.py official repository ran fingerprinted no licence file found · pointer only · ba4811601ac7b530 · report
apply_semantix_fix expertisemodel/mutap/MuTAP/Sematic_err_correction.py official repository unverified no licence file found · pointer only · c129303ba93e3e99 · report
run_asserts_on_scripts expertisemodel/mutap/MuTAP/greedy_test_generator.py official repository unverified no licence file found · pointer only · 7574ac527059c5f2 · report

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections