Academic Journal

基于大语言模型的 HTTP/HTTPS 网络资产设备类型识别方法.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: 基于大语言模型的 HTTP/HTTPS 网络资产设备类型识别方法. (Chinese)
Alternate Title: Device type identification of HTTP / HTTPS network asset based on large language models. (English)
Συγγραφείς: 陈倩怡, 苏马婧, 陈紫璇, 张永奇, 马 琰
Πηγή: Cyber Security & Data Governance; 2026, Vol. 45 Issue 5, p1-10, 10p
Θεματικοί όροι: HTTP (Computer network protocol), Language models, Internet protocols
Abstract (English): To address the limited generalization ability of traditional network asset identification methods based on static fingerprint rules and discriminative models in complex and open environments, this paper proposes an HTTP/ HTTPS network asset device type identification method based on instruction fine-tuning of a large language model. A multi-source data collection scheme with multi-platform label aggregation is designed to construct the original network asset dataset. A data preprocessing strategy that prioritizes key feature retention is applied to reduce redundant noise in model inputs. Multiple heterogeneous features, including HTTP/ HTTPS response bodies, response headers, SSL certificates, ports, and protocols, are further integrated to construct a unified serialized representation. Based on this representation, the LoRA technique is employed to perform parameter-efficient fine-tuning on the LLaMA-3-8B-Instruct model, enabling the model to learn the semantic associations between network asset characteristics and device types. Experimental results on a test dataset containing 380 000 real-world network assets demonstrate that the proposed method maintains stable performance under highly imbalanced samples and long-tail device scenarios, achieving a Weighted F1-score of 0. 959 1, which significantly outperforms the unfine-tuned base model. In addition, the model inference throughput is improved by 62. 81% . These results verify the effectiveness and practicality of the proposed method for large-scale automated network asset device identification. [ABSTRACT FROM AUTHOR]
Abstract (Chinese): 针对传统基于静态指纹规则和判别式模型在复杂开放环境下泛化能力不足的现状, 提出了一种基于大 语言模型指令微调的 HTTP/HTTPS 网络资产设备类型识别方法。 通过多源采集与多平台标签聚合的数据收集方 案构造原始网络资产数据集, 对该数据集进行关键特征优先保留的数据预处理, 有效降低冗余噪声对模型输入 的影响, 然后通过融合 HTTP/HTTPS 响应体、 响应头、 SSL 证书、 端口及协议等多源异构特征, 构建统一的序 列化表示; 在此基础上, 利用 LoRA 技术对 LLaMA-3-8B-Instruct 模型进行参数高效微调, 引导模型学习网络资 产特征与设备类型之间的语义关联关系。 实验结果表明, 在包含 38 万条真实网络资产的测试集中, 该方法在样 本高度不均衡和长尾设备场景下仍能保持稳定性能, Weighted F1-score 达到 0. 959 1, 相比未微调模型效果显著 提升。 同时, 模型推理吞吐量提高 62. 81%, 验证了所提方法在大规模网络资产自动识别任务中的有效性与实 用性。 [ABSTRACT FROM AUTHOR]
Copyright of Cyber Security & Data Governance is the property of Editorial Office of Information Technology & Network Security and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Βάση Δεδομένων: Complementary Index
Περιγραφή
ISSN:20971788
DOI:10.19358/j.issn.2097-1788.2026.05.001