基于语义引导Cross-Mamba的红外与可见光图像融合

A Semantic-Guided Cross-Mamba Network for Infrared and Visible Image Fusion

  • 摘要: 针对现有红外与可见光图像融合方法多侧重低层纹理与对比度增强,而对跨模态高层语义一致性及目标检测等下游视觉任务适应性关注不足的问题,提出一种语义引导Mamba融合网络(SemanticGuided Mamba Fusion,SGMFusion)。该网络采用双路ResNet-18编码器提取多尺度特征,在浅层通过局部空间注意融合模块保留纹理细节,在深层通过Cross-Mamba融合模块建模跨模态长距离依赖,形成深浅层异构融合范式。融合特征经多尺度语义聚合后,通过语义注入模块以条件归一化方式逐级注入重建路径,增强融合结果的语义一致性。同时引入基底图–残差输出范式以提升融合稳定性。训练阶段从强度、梯度、结构相似性与统计相关性等方面联合约束融合结果,从像素保真、结构清晰度、边缘方向与统计相关性方面联合约束融合质量。在MSRS、TNO和LLVIP数据集上与7种先进方法对比实验结果表明,SGMFusion在多项融合指标上表现较优,并在YOLOv8目标检测任务中取得最优mAP,验证了方法的有效性与检测友好性。

     

    Abstract: Infrared and visible image fusion aims to integrate thermal target information and visible texture details into a single image. Existing methods usually emphasize low-level texture preservation and contrast enhancement, while the consistency of high-level cross-modal semantics and the requirements of downstream tasks such as object detection are often less considered. In this paper, we propose SGMFusion, a semanticguided Mamba fusion network. The network uses dual ResNet-18 encoders to extract multi-scale features. For shallow features, a local spatial attention fusion module is adopted to preserve texture and edge details. For deep features, a Cross-Mamba Fusion module is designed to capture long-range cross-modal dependencies. The fused features are further processed by multi-scale semantic aggregation, and semantic cues are injected into the reconstruction path through a conditional-normalization-based semantic injection module. This design helps improve semantic consistency during image reconstruction. Meanwhile, a base-residual output strategy is introduced to improve fusion stability. During training, the fusion results are jointly constrained in terms of intensity, gradient information, structural similarity, and statistical correlation, thereby improving pixel fidelity, structural clarity, edge consistency, and statistical relevance. Comparative experiments against 7 state-of-theart methods on the MSRS, TNO, and LLVIP datasets show that SGMFusion achieves competitive performance on several fusion metrics. In the YOLOv8 object detection experiment, it also obtains the highest mAP among the compared methods, indicating that the proposed method is suitable for detection-oriented fusion tasks.

     

/

返回文章
返回