基于通道对齐与自注意融合的可见光-红外遥感图像目标检测算法

Target Detection for Visible-Infrared Remote Sensing Images Using Channel Alignment and Self-Attentive Fusion

  • 摘要: 针对当前无人机多模态遥感目标检测中存在的跨模态交互效率低下及融合不充分问题,本文提出一种新型目标检测算法PCMFNet。首先,设计渐进式跨模态一致性交互模块(CMCE),通过计算可见光与红外模态的通道一致性图实现特征对齐并捕获互补信息,同时利用残差连接保留模态内有效特征,从而高效促进跨模态交互。其次,引入基于融合结果引导的轻量化跨模态自注意力融合模块( CSAF),以深层红外通道特征作为值向量建模全局上下文信息,实现跨模态特征的有效融合并增强表征能力。实验结果表明,该方法在两个可见光-红外遥感目标检测数据集DroneVehicle和VTUAV-det上均取得明显性能提升,mAP0.5较基线模型分别提高了3.2%和0.7%,有效地提升了无人机在复杂场景下的多模态目标检测精度。

     

    Abstract: To address the challenges of inefficient cross-modal interaction and insufficient fusion in current multimodal remote sensing object detection for unmanned aerial vehicles (UAVs), this paper proposes a novel object detection algorithm named PCMFNet. First, a Progressive Cross-modal Consistency Enhancement module (CMCE) is designed to achieve feature alignment and capture complementary information by computing channel-wise consistency maps between visible and infrared modalities. Simultaneously, residual connections are utilized to preserve intra-modal discriminative features, thereby efficiently facilitating crossmodal interaction. Second, a lightweight Cross-modal Self-Attention Fusion module (CSAF) guided by fusion outcomes is introduced. This module employs deep infrared channel features as value vectors to model global contextual information, enabling effective cross-modal feature fusion and enhancing representation capacity. Experimental results demonstrate that the proposed method achieves significant performance improvements on two visible-infrared remote sensing object detection datasets, DroneVehicle and VTUAV-det. Specifically, it elevates mAP0.5 by 3.2% and 0.7%, respectively, over the baseline model, significantly enhancing UAVbased multimodal object detection accuracy in complex scenarios.

     

/

返回文章
返回