DK-CLIP:基于红外-可见光特征融合与语义引导的跨模态行人重识别方法

DK-CLIP: A Cross-Modal Person Re-Identification Method Based on Infrared-Visible Feature Fusion and Semantic Guidance

  • 摘要: 红外成像技术在夜间监控、恶劣天气及隐蔽侦察等场景中具有不可替代的优势,与可见光图像融合可显著提升全天候行人重识别系统的鲁棒性。然而,可见光与红外图像间存在显著的光谱与语义差异,现有跨模态识别方法在特征对齐与融合方面仍面临挑战。为此,本文提出一种面向红外-可见光融合的对比语言-图像预训练网络DK-CLIP,通过语义引导的特征增强与动态融合机制,实现跨模态行人重识别。首先,设计文本描述符为红外与可见光图像生成高层语义描述,增强特征表达的模态不变性;其次,构建基于KAN的Transformer编码器KANformer,提升对红外图像弱纹理特征的提取能力,并引入对比损失促进模态间对齐;最后,提出动态卷积核融合模块,自适应融合双模态特征,有效抑制模态差异。在SYSU-MM01数据集上的实验表明,本文方法的平均精度mAP与首位命中率Rank-1分别达到70.69%和75.01%;在RegDB数据集上同样表现优异。该方法为红外与可见光协同感知提供有效技术路径,在智能安防与全天候监控系统中具有良好应用前景。

     

    Abstract: Infrared imaging technology offers irreplaceable advantages in night surveillance, adverse weather conditions, and covert reconnaissance. Its fusion with visible light images can significantly enhance the robustness of all-weather person re-identification systems. However, significant spectral and semantic discrepancies exist between visible and infrared images, posing challenges for existing cross-modal recognition methods in feature alignment and fusion. To address these issues, this paper proposes DK-CLIP, a contrastive language-image pre-training network designed for infrared-visible fusion, which achieves crossmodal person re-identification through semantically-guided feature enhancement and a dynamic fusion mechanism. First, a text descriptor is designed to generate high-level semantic descriptions for both infrared and visible images, thereby enhancing the modality invariance of feature representation. Second, a Transformer encoder based on the KAN, termed KANformer, is constructed to improve the extraction of weaktexture features in infrared images, while a contrastive loss is introduced to promote inter-modal alignment. Finally, a dynamic convolutional kernel fusion module is proposed to adaptively fuse dual-modal features and effectively suppress moda l discrepancies. Experiments on the SYSU-MM01 dataset demonstrate that the proposed method achieves a mean average precision (mAP) of 70.69% and a Rank-1 accuracy of 75.01%. It also exhibits excellent performance on the RegDB dataset. The proposed method provides an effective technical pathway for infrared-visible cooperative perception and holds promising application prospects in intelligent security and all-weather surveillance systems.

     

/

返回文章
返回