Academic Journal

Adaptive Multi‐Metric Test Case Selection for Deep Neural Networks Based on Genetic Algorithm.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Adaptive Multi‐Metric Test Case Selection for Deep Neural Networks Based on Genetic Algorithm.
Συγγραφείς: Wang, Weiwei1 (AUTHOR) wangweiwei@bipt.edu.cn, Chen, Qingshuai1 (AUTHOR), Zhao, Zimo1 (AUTHOR), Liu, Xuejun1 (AUTHOR), Zhao, Ruilian2 (AUTHOR)
Πηγή: Journal of Software: Evolution & Process. Jul2026, Vol. 38 Issue 7, p1-21. 21p.
Θεματικοί όροι: *Genetic algorithms, *Computer software testing, *Fault diagnosis, *Artificial neural networks, *Uncertainty (Information theory), *Multi-objective optimization
Περίληψη: As deep neural networks (DNNs) are increasingly deployed in safety‐ and mission‐critical domains, their latent defects and security risks have become a growing concern. Test case selection (TCS) aims to identify and label those test cases most likely to be misclassified within a limited annotation budget, thereby maximizing the exposure of real faults at minimal cost and providing high‐value data for subsequent localization, repair, and regression testing. However, current TCS methods predominantly utilize single‐dimensional metrics, such as neuron coverage or mutant‐killing rate, limiting their ability to comprehensively capture diverse fault patterns in complex models. This paper proposes a novel DNN TCS framework based on adaptive multi‐metric optimization, which takes both prediction uncertainty and divergence into account and adjust the weights of them to qualify the effectiveness of test cases by genetic algorithm. Concretely, firstly, multiple slightly mutated models are constructed to capture the behavioral discrepancies of each test case across different models. It then establishes two complementary effectiveness metrics‐prediction uncertainty and prediction divergence‐to quantify a test case's local decision boundary sensitivity and global behavioral diversity, respectively. Finally, a genetic algorithm adaptively optimizes the weights of these metrics, enabling a precise assessment of each test case and the identification of the subset with the highest fault‐revealing potential. To verify the effectiveness of our approach, experiments and evaluations are conducted on DNNs across both computer vision and natural language processing domains. And the experimental results show that our approach surpasses the best baseline by improving the fault detection rate by an average of 3.32% on the Top‐5% candidate subset and 9.54% on the Top‐10% subset. [ABSTRACT FROM AUTHOR]
Βάση Δεδομένων: Academic Search Index