C3Net: A cross-modal collaborative calibration of features for object detection using frames and events.

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: C3Net: A cross-modal collaborative calibration of features for object detection using frames and events.
Συγγραφείς: Chen Y; School of Computer Science and Technology, Guangdong University of Technology, Guangzhou, China., Zhong J; School of Computer Science and Technology, Guangdong University of Technology, Guangzhou, China., Guo Y; School of Electronic Information, Wuhan University, Wuhan, China., Xie Z; School of Computer Science and Technology, Guangdong University of Technology, Guangzhou, China., Xiao J; School of Electronic Information, Wuhan University, Wuhan, China. Electronic address: xiaojs@whu.edu.cn., Chen P; School of Computer Science and Technology, Guangdong University of Technology, Guangzhou, China. Electronic address: phchen@gdut.edu.cn.
Πηγή: Neural networks : the official journal of the International Neural Network Society [Neural Netw] 2026 Jul; Vol. 199, pp. 108651. Date of Electronic Publication: 2026 Feb 02.
Τύπος έκδοσης: Journal Article
Γλώσσα: English
Στοιχεία περιοδικού: Publisher: Pergamon Press Country of Publication: United States NLM ID: 8805018 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1879-2782 (Electronic) Linking ISSN: 08936080 NLM ISO Abbreviation: Neural Netw Subsets: MEDLINE
Imprint Name(s): Original Publication: New York : Pergamon Press, [c1988-
Ιατρικοί όροι (MeSH): Image Processing, Computer-Assisted* , Detection Algorithms*, Calibration
Περίληψη: Object detection by fusing RGB frames and event streams is challenging due to their inherent heterogeneity and significant statistical disparities, which often lead to suboptimal fusion in existing methods. To address this, we introduce C3Net, a novel framework built upon a paradigm shift from direct feature merging to Collaborative Calibration. First, we propose an Adaptive Balancing Time Surface (ABTS) to generate motion-robust event representations by mitigating spatial inconsistencies caused by varying object velocities. Second, the core Cross-Modal Feature Collaborative Calibration Module (CM-FCCM) performs mutual calibration of RGB and event features across channel and spatial dimensions, reducing modality discrepancies before fusion; the calibrated features are then fed back to the respective backbones for enriched feature learning. Finally, an Adaptive Channel Fusion Module (ACFM) dynamically integrates the modalities based on channel-wise confidence. Extensive experiments on PKU-DAVIS-SOD, DSEC-MOD, and PKU-DDD17-CAR datasets demonstrate that C3Net achieves state-of-the-art performance, showcasing its superior ability to leverage the complementary strengths of frames and events.
(Copyright © 2026 Elsevier Ltd. All rights reserved.)
Competing Interests: Declaration of competing interest The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Yunhua Chen reports financial support was provided by Department of Science and Technology of Guangdong Province. Pinghua Chen reports financial support was provided by Department of Science and Technology of Guangdong Province. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Contributed Indexing: Keywords: Event camera; Multimodal fusion; Object detection
Entry Date(s): Date Created: 20260212 Date Completed: 20260827 Latest Revision: 20260827
Update Code: 20260827
DOI: 10.1016/j.neunet.2026.108651
PMID: 41679046
Βάση Δεδομένων: MEDLINE
Περιγραφή
ISSN:1879-2782
DOI:10.1016/j.neunet.2026.108651