Abstract:
Infrared imaging technology offers irreplaceable advantages in night surveillance, adverse weather conditions, and covert reconnaissance. Its fusion with visible light images can significantly enhance the robustness of all-weather person re-identification systems. However, significant spectral and semantic discrepancies exist between visible and infrared images, posing challenges for existing cross-modal recognition methods in feature alignment and fusion. To address these issues, this paper proposes DK-CLIP, a contrastive language-image pre-training network designed for infrared-visible fusion, which achieves crossmodal person re-identification through semantically-guided feature enhancement and a dynamic fusion mechanism. First, a text descriptor is designed to generate high-level semantic descriptions for both infrared and visible images, thereby enhancing the modality invariance of feature representation. Second, a Transformer encoder based on the KAN, termed KANformer, is constructed to improve the extraction of weaktexture features in infrared images, while a contrastive loss is introduced to promote inter-modal alignment. Finally, a dynamic convolutional kernel fusion module is proposed to adaptively fuse dual-modal features and effectively suppress moda l discrepancies. Experiments on the SYSU-MM01 dataset demonstrate that the proposed method achieves a mean average precision (mAP) of 70.69% and a Rank-1 accuracy of 75.01%. It also exhibits excellent performance on the RegDB dataset. The proposed method provides an effective technical pathway for infrared-visible cooperative perception and holds promising application prospects in intelligent security and all-weather surveillance systems.