基于特征与决策互补的红外及可见光融合目标检测

Infrared and Visible Light Fusion Object Detection Based on Feature and Decision Complementarity

  • 摘要: 为提升红外和可见光目标检测方法的准确性和鲁棒性,将深度学习技术和多源信息融合技术相结合,提出了一种基于特征与决策信息互补的多模态目标检测网络。网络首先基于特征通道关键度来设计轻量级分离卷积单元,并通过串联堆叠的方式构建特征提取骨干结构;其次,基于骨干结构以并列多分支架构分别提取目标在红外和可见光波段下的特征,并结合分步融合策略设计特征融合分支,从特征层面实现目标多模态信息互补;然后,设计多尺度自适应融合检测结构,将红外、可见光以及融合分支特征进行多尺度融合后分别预测目标类别及位置;最后,通过结合预测框置信度、重叠面积、中心点距离等信息,利用改进的非极大值软抑制策略,从决策级融合层面筛选出最终的目标框。通过在标准数据集上的实验表明:所提方法各个结构对目标检测性能都有一定提升效果,能够有效实现多模态信息互补,同时避免不同目标间的特征相互干扰。相较于同类型方法,该方法在模型鲁棒性和泛化能力等方面均展现出显著的优势,能够在复杂场景下有效地完成目标检测任务。

     

    Abstract: To enhance the accuracy and robustness of infrared and visible light object detection, this paper proposes a multi-modal detection network that integrates deep learning with multi-source information fusion, leveraging complementary features and decision-level cues. The network first introduces lightweight separable convolution units designed according to the criticality of feature channels, which are then concatenated and stacked to construct a backbone for feature extraction. On top of this backbone, a parallel multi-branch architecture is adopted to extract object features from both infrared and visible spectra. A feature fusion branch, guided by a stepwise fusion strategy, is further incorporated to achieve multimodal feature complementarity. Subsequently, a multi-scale adaptive fusion detection structure is designed to integrate features from the infrared, visible, and fusion branches at different scales, enabling separate predictions of object categories and locations. At the decision-level fusion stage, an improved soft non-maximum suppression strategy—incorporating predicted bounding box confidence scores, overlap area, and center-point distance—is employed to select the final object bounding boxes. Experimental results on standard benchmarks demonstrate that each component of the proposed method contributes positively to detection performance, effectively realizing multimodal information complementarity while mitigating inter-object feature interference. Compared with state-of-the-art alternatives, the proposed method exhibits marked advantages in robustness and generalization, and proves capable of reliably performing object detection in complex scenarios.

     

/

返回文章
返回