Browse State-of-the-Art › software testing
software testing
39 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 39 papers with code (135 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 Mar 2024 3 repositories listed Syntology ran 12 of 17 samples · 5 unverifiedCompared to MuTAP, a mutation- and LLM-based test generator, CoverUp achieves an overall line+branch coverage of 89% (vs.
-
28 Jul 2018 3 repositories listed Syntology ran 1 of 12 samples · 11 unverifiedWe then discuss the application of CGF to the following goals: finding numerical errors in trained neural networks, generating disagreements between neural networks and quantized versions of those networks, and…
-
17 Jun 2025 2 repositories listedThis paper reviews recent research on AI augmentation in software test automation, from no automation to full automation.
-
13 Feb 2022 2 repositories listedThis paper investigates the parameter space of machine learning (ML) algorithms in aggravating or mitigating fairness bugs.
-
5 Jun 2020 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We propose D-RISE, a method for generating visual explanations for the predictions of object detectors.
-
26 Apr 2025 1 repository listedIn-context learning (ICL) has emerged as a powerful capability of large language models (LLMs), enabling them to perform new tasks based on a few provided examples without explicit fine-tuning.
-
30 Mar 2025 1 repository listedGUI agents, powered by large foundation models, can interact with digital interfaces, enabling various applications in web automation, mobile navigation, and software testing.
-
3 Feb 2025 1 repository listedData augmentation has become a standard practice in software engineering to address limited or imbalanced data sets, particularly in specialized domains like test classification and bug detection where data can be…
-
15 Oct 2024 1 repository listedHowever, it is important to ensure that identified failures are spread throughout the entire failure-inducing area of a search domain and not clustered in a sub-region.
-
21 Aug 2024 1 repository listedAutomated unit test generators, particularly search-based software testing tools like EvoSuite, are capable of generating tests with high coverage.
-
18 Jun 2024 1 repository listed Syntology ran 5 of 6 samples · 1 unverifiedWe find that LLMs generally perform surprisingly well at generating relevant test cases, with Code Agents designed for code repair exceeding the performance of systems designed specifically for test generation.
-
16 Apr 2024 1 repository listedTo address this problem, we propose TrickCatcher, an LLM-powered approach to generating test cases for uncovering bugs in plausible programs.
-
22 Mar 2024 1 repository listedBehavior-driven development (BDD) is an Agile testing methodology fostering collaboration among developers, QA analysts, and stakeholders.
-
20 Mar 2024 1 repository listedWe study two types of settings: one where there is iid noise in the observation, and a more challenging setting where there is also the presence of exogenous noise, which is non-iid noise that is temporally correlated,…
-
20 Feb 2024 1 repository listedSubsequently, QuanTest formulates the problem of generating test inputs that maximize the quantum entanglement adequacy and capture incorrect behaviors of the QNN system as a joint optimization problem and solves it in…
-
1 Feb 2024 1 repository listedTo our best knowledge, this is the first work that covers the intersection of three areas, including LLMs, fuzzing test, and fuzzing test generated based on LLMs.
-
31 Jan 2024 1 repository listedThe results show that LLMs can successfully generate realistic test data generators in a wide range of domains at all three levels of integrability.
-
13 Jan 2024 1 repository listedAutomated test case generation has proven to be useful to reduce the usually high expenses of software testing.
-
23 Dec 2023 1 repository listedSynthetic data has numerous applications, including but not limited to software testing at scale, privacy-preserving data sharing to enable smoother collaboration between stakeholders, and data augmentation for…
-
4 Oct 2023 1 repository listedTo this end, we propose an automated approach which exploits both structural and semantic properties of source code methods and test cases to recommend the most relevant and useful unit tests to the developers.
-
11 Apr 2023 1 repository listedOur experimental study shows that (1) lexical, syntactic and structural properties of source code are encoded in the lower, intermediate, and higher layers, respectively, while the semantic property spans across the…
-
6 Mar 2023 1 repository listedTo understand and debug convolutional neural networks (CNNs) we propose techniques for testing the channels of CNNs.
-
2 Mar 2023 1 repository listedWe introduce Reasoning-Based Software Testing (RBST), a new way of thinking at the testing problem as a causal reasoning task.
-
14 Feb 2023 1 repository listedOngoing progress in computational intelligence (CI) has led to an increased desire to apply CI techniques for the purpose of improving software engineering processes, particularly software testing.
-
3 Feb 2023 1 repository listedThis paper presents SEER, a learning-based approach that in the absence of test assertions or other types of oracle, can determine whether a unit test passes or fails on a given method under test (MUT).
-
18 Jan 2023 1 repository listedHowever, typically, the variables of a dataset depend on one another, and these dependencies are not considered in data generation leading to the creation of implausible records.
-
25 Aug 2022 1 repository listedIn this paper, we empirically investigate the applications of carefully selected DRL algorithms on two important software testing tasks: test case prioritization in the context of Continuous Integration (CI) and game…
-
25 Jul 2022 1 repository listedThe execution of the feasible tests revealed that there is a large amount of deviations for the scores and classes.
-
8 Jul 2022 1 repository listedSPLADE efficiency can be controlled via a regularization factor, but solely controlling this regularization has been shown to not be efficient enough.
-
1 May 2022 1 repository listedWe hope to open a discussion on the best methodologies to handle social bias testing in language models.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections