Abstract:
In response to the problem that existing infrared and visible light fusion algorithms cannot fully fuse cross modal features, resulting in dim bright areas and loss of background information. This article proposes a Convolutional Neural Network (CNN) combined with Histogram Transformer network (HSFusion) to achieve infrared and visible image fusion. Firstly, design a histogram Transformer-CNN structure that utilizes their respective inductive biases to better model cross modal features and achieve joint learning of global and local features. Then the network adopts a U-Net-like architecture to obtain long-range and short-range dependencies at multiple scales, in order to integrate source information. The experimental results on the MSRS, TNO, and RoadScene datasets show that our algorithm achieves better fusion performance compared to other advanced algorithms, and also proves to improve the accuracy of downstream multimodal object detection and multimodal semantic segmentation tasks.