Academic Journal

Research on multimodal dense small object detection algorithms guided by cross-modal information.

Bibliographic Details
Title: Research on multimodal dense small object detection algorithms guided by cross-modal information.
Authors: Hu J; A. B. Freeman School of Business, Tulane University, New Orleans, United States of America., Ma Y; Ira A. Fulton School of Engineering, Arizona State University, Tempe, United States of America., Xing Y; School of Engineering and Applied Science, University of Pennsylvania, Philadelphia, United States of America., Zi Y; School of Computer Science, Georgia Institute of Technology, Atlanta, United States of America., Deng Y; School of Computer Science, Georgia Institute of Technology, Atlanta, United States of America., Wang M; College of Graduate and Professional Studies, Trine University, Phoenix, United States of America., Liu H; College of Engineering, Northeastern University, Boston, United States of America., Du J; MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University, Shanghai, China.
Source: PloS one [PLoS One] 2026 Sep 08; Vol. 21 (9), pp. e0348533. Date of Electronic Publication: 2026 Sep 08 (Print Publication: 2026).
Publication Type: Journal Article
Language: English
Journal Info: Publisher: Public Library of Science Country of Publication: United States NLM ID: 101285081 Publication Model: eCollection Cited Medium: Internet ISSN: 1932-6203 (Electronic) Linking ISSN: 19326203 NLM ISO Abbreviation: PLoS One Subsets: MEDLINE
Imprint Name(s): Original Publication: San Francisco, CA : Public Library of Science
MeSH Terms: Image Processing, Computer-Assisted*/methods , Pattern Recognition, Automated*/methods , Detection Algorithms*, Algorithms ; Humans
Abstract: To address the challenges of modality heterogeneity, scale inconsistency, and background interference in dense small object detection under multimodal conditions, this paper proposes a novel detection framework based on cross-modality guidance and hierarchical scale refinement. Built upon the RT-DETR backbone, the framework integrates a Cross-Modality Guided Dynamic Fusion (CMG-DF) module, which performs semantic-level recalibration between infrared and visible features via a learnable modality attention mechanism, and a Hierarchical Scale Refinement Network (HSRN), which enhances semantic consistency and boundary continuity across scales through bidirectional residual flow and graph-based relational modeling. To validate the effectiveness of the proposed method, extensive comparison and ablation studies are conducted on two public multimodal benchmarks, SMOD and LLVIP. Experimental results show that the proposed method achieves 92.7% mAP@50 and 69.5% mAP@50:95 on SMOD, as well as 77.3% mAP@50 and 43.1% mAP@50:95 on LLVIP, consistently outperforming existing state-of-the-art multimodal detection algorithms. Qualitative visualizations further confirm the robustness and enhancement capability of the method for small objects under low illumination, occlusion, and complex background conditions, highlighting its strong structural generalization and practical deployment potential.
(Copyright: © 2026 Hu et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.)
Competing Interests: The authors have declared that no competing interests exist.
References: Neural Netw. 2026 Apr;196:108450. (PMID: 41406644)
IEEE Trans Neural Netw Learn Syst. 2025 Mar;36(3):4145-4159. (PMID: 34437075)
IEEE Trans Cybern. 2016 Aug 04;47(11):3980-3990. (PMID: 28708566)
PLoS One. 2021 Oct 29;16(10):e0259283. (PMID: 34714878)
IEEE Trans Image Process. 2026;35:6517-6529. (PMID: 42301838)
Entry Date(s): Date Created: 20260908 Date Completed: 20260908 Latest Revision: 20260910
Update Code: 20260910
PubMed Central ID: PMC13552963
DOI: 10.1371/journal.pone.0348533
PMID: 42709862
Database: MEDLINE
Description
ISSN:1932-6203
DOI:10.1371/journal.pone.0348533