Academic Journal

SemuFuzz: An efficient fuzzing approach for deep learning libraries based on semantic representation.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: SemuFuzz: An efficient fuzzing approach for deep learning libraries based on semantic representation.
Συγγραφείς: Chen, Zhuoyi1,2 (AUTHOR) 2212408053@stmail.ujs.edu.cn, Chen, Jinfu1,2 (AUTHOR) jinfuchen@ujs.edu.cn, Cai, Saihua1,2 (AUTHOR) caisaih@ujs.edu.cn, Chen, Jingyi1,2 (AUTHOR) jychen@stmail.ujs.edu.cn, Gu, Wenjie3 (AUTHOR) wenjie.gu@connect.polyu.hk
Πηγή: Journal of Systems & Software. Nov2026, Vol. 241, pN.PAG-N.PAG. 1p.
Θεματικοί όροι: *Software libraries (Computer programming), *Defect tracking (Computer software development), Computer software security, Encoding, Computer software testing
Περίληψη: Deep learning libraries form the foundational architecture of modern artificial intelligence systems, supporting applications in fields such as computer vision, natural language processing, and software engineering. However, their massive APIs and high-dimensional input spaces make them vulnerable to serious vulnerabilities such as buffer overflows and illegal memory accesses, jeopardizing system security and reliability. While fuzzing methods using large language models have shown promise in API-level testing in recent years, they typically rely on high-dimensional LLM-generated embeddings for boundary-use case clustering, which leads to inefficiency and limited controllability. To address these limitations, we present SemuFuzz, a systematic refinement of the DFuzz framework. SemuFuzz retains the overall architecture of DFuzz while improving the handling of semantic constraints by replacing LLM-generated embedding vectors with lightweight semantic embeddings and clustering analysis, enabling more efficient and controllable grouping of semantically similar test cases. This design makes it possible to systematically, scalably, and easily extend vulnerability discovery within deep learning libraries. Empirical evaluations on TensorFlow and PyTorch demonstrate that SemuFuzz achieves higher API coverage and stronger vulnerability-triggering capabilities than state-of-the-art tools. It discovered a total of 36 different bugs, including 10 previously unknown vulnerabilities, two of which have been officially fixed by the maintainers. This verifies the effectiveness and practical value of our method. • Proposes SemuFuzz for semantic-aware API edge-case generation. • Enables reusable test inputs across APIs using semantic constraints. • Found 36 crashes; 2 confirmed and patched by developers. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Systems & Software is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Business Source Index
Περιγραφή
ISSN:01641212
DOI:10.1016/j.jss.2026.112999