CNN联合直方图Transformer的红外与可见光图像融合

Fusion of Infrared and Visible Images Using CNN Joint Histogram Transformer

  • 摘要: 针对现有红外与可见光融合算法不能对跨模态特征进行充分融合,导致高亮区域暗淡、背景信息丢失的问题。本文提出一种CNN(Convolutional Neural Network,CNN)联合直方图Transformer的网络(HSFusion)来实现红外和可见光图像融合。首先设计一个直方图Transformer-CNN结构,利用它们各自的归纳偏置来更好地对跨模态特征进行建模,实现全局特征和局部特征的联合学习。然后网络采用类U-Net架构,获取多个尺度的长程和短程依赖关系,以集成源信息。经过在MSRS、TNO和RoadScene数据集的实验结果表明,本文算法与当前其他先进算法相比,实现了更好的融合效果,同时证明可以提高下游多模态目标检测和多模态语义分割任务的精度。

     

    Abstract: In response to the problem that existing infrared and visible light fusion algorithms cannot fully fuse cross modal features, resulting in dim bright areas and loss of background information. This article proposes a Convolutional Neural Network (CNN) combined with Histogram Transformer network (HSFusion) to achieve infrared and visible image fusion. Firstly, design a histogram Transformer-CNN structure that utilizes their respective inductive biases to better model cross modal features and achieve joint learning of global and local features. Then the network adopts a U-Net-like architecture to obtain long-range and short-range dependencies at multiple scales, in order to integrate source information. The experimental results on the MSRS, TNO, and RoadScene datasets show that our algorithm achieves better fusion performance compared to other advanced algorithms, and also proves to improve the accuracy of downstream multimodal object detection and multimodal semantic segmentation tasks.

     

/

返回文章
返回