Abstract:
Multimodal fusion target detection has wide applications in all-weather target detection. However, current fusion methods mainly rely on simple accumulation operations and neglect the impact of modal feature differences on cross-modal fusion performance. To address this problem, this study introduces a multimodal target detection algorithm, known as the Cross-Scale Cross-Attention Fusion Module (CSMF), which incorporates cross-scale feature guidance and cross-attention fusion. The CSCM was employed to fuse visible and infrared images. Through cross-scale feature guidance, it learns the cross-attention relationships between the modalities, resulting in efficiently fused features. Moreover, a receptive field enhancement module (RFE) and an adaptive spatial pyramid pooling module (ASPPF) were introduced to enrich feature representations. At the data level, a cross-modal data augmentation method, CDS, was employed to effectively capture cross-modal correlations in local regions. The experimental results revealed that the CSMF algorithm outperformed several existing mainstream methods on the FLIR and LLVIP datasets, improving the mAP50 by 5.3% and 1.6%, respectively, compared with the original algorithms, with a significant enhancement in detection performance within complex environments.