Academic Journal

Issue classification with LLMs: An empirical study of the NASA flight software systems.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Issue classification with LLMs: An empirical study of the NASA flight software systems.
Συγγραφείς: Colavito, Giuseppe1 (AUTHOR) giuseppe.colavito@uniba.it, Lanubile, Filippo1 (AUTHOR) filippo.lanubile@uniba.it, Novielli, Nicole1 (AUTHOR) nicole.novielli@uniba.it, Arreza, Christopher2 (AUTHOR) chris.arreza@nasa.gov, Shi, Ying2 (AUTHOR) ying.shi@nasa.gov
Πηγή: Journal of Systems & Software. Jul2026, Vol. 237, pN.PAG-N.PAG. 1p.
Θεματικοί όροι: *Defect tracking (Computer software development), *Generative artificial intelligence, *Systems software, Language models, Machine learning
Εταιρία/Οντότητα: United States. National Aeronautics & Space Administration
Περίληψη: • This study compares different approaches for automatically classifying software issue reports into "bug" vs "non-bug" categories, evaluating both traditional encoder-only models (like BERT) and newer generative LLMs on NASA's Core Flight System (cFS) and F´ datasets. • Fine-tuned encoder-only models (specifically SetFit) consistently outperformed zero-shot generative LLMs, achieving better F1 scores and bug recall. Notably, SetFit required only 40 labeled examples per class to outperform zero-shot LLMs, demonstrating that effective classification can be achieved with small, well-curated training sets rather than large labeled datasets. However, generative LLMs demonstrated comparable zero-shot performance, thus representing a viable solution in the absence of training data. • Model performance varied considerably between the two NASA datasets. Models performed better on the F´ dataset (from a single project with immediate, author-driven labeling) compared to cFS (multiple projects with delayed, maintainer-driven labeling), highlighting how dataset characteristics and labeling practices impact classification accuracy. • Substantially different computational costs are associated with the use of encoder-only models vs. generative LLMs. BERT-like models achieve inference in seconds on a single GPU, while generative LLMs require 25+ minutes using multiple GPUs, highlighting the trade-off between classification performance and computational resources for LLM deployment in production environments. NASA collects vast amounts of problem data for space projects, which includes not only descriptions of defects but also enhancements and other issue reports. The growing complexity of Flight Software has led to an increase in the volume of problem reports, presenting both opportunities and challenges in data analysis. This paper explores AI-based solutions for classifying software issue reports in NASA's spacecraft control systems. In particular, we aim to develop an accurate classifier for identifying bug tickets, building on previous research in automated issue labeling. We conduct a benchmark study for comparing various language models and provide insights on their performance and deployment costs, with the goal of improving issue classification for NASA's growing software complexity. Based on our empirical results, we provide empirically-driven guidelines on how to address the tradeoff between the need for manual labeling of training data and the computational costs associated with the on-premise deployment of LLMs that could be used in a zero-shot setting. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Systems & Software is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Business Source Index