Abstract:
To enhance the accuracy and robustness of infrared and visible light object detection, this paper proposes a multi-modal detection network that integrates deep learning with multi-source information fusion, leveraging complementary features and decision-level cues. The network first introduces lightweight separable convolution units designed according to the criticality of feature channels, which are then concatenated and stacked to construct a backbone for feature extraction. On top of this backbone, a parallel multi-branch architecture is adopted to extract object features from both infrared and visible spectra. A feature fusion branch, guided by a stepwise fusion strategy, is further incorporated to achieve multimodal feature complementarity. Subsequently, a multi-scale adaptive fusion detection structure is designed to integrate features from the infrared, visible, and fusion branches at different scales, enabling separate predictions of object categories and locations. At the decision-level fusion stage, an improved soft non-maximum suppression strategy—incorporating predicted bounding box confidence scores, overlap area, and center-point distance—is employed to select the final object bounding boxes. Experimental results on standard benchmarks demonstrate that each component of the proposed method contributes positively to detection performance, effectively realizing multimodal information complementarity while mitigating inter-object feature interference. Compared with state-of-the-art alternatives, the proposed method exhibits marked advantages in robustness and generalization, and proves capable of reliably performing object detection in complex scenarios.