Abstract:
To address the problems of low pixel occupancy, strong background interference, and insufficient single-modality information in visual detection of low-altitude, slow-speed, and small-sized UAVs, an improved YOLO11n detection algorithm based on visible-infrared dual-modal fusion is proposed. A hierarchical hybrid fusion module HCF is designed, introducing lightweight cross-attention at shallow layers for fine-grained interaction and stable concatenation at deep layers for semantic robustness. The DBSPPF module replaces the original SPPF to enhance multi-scale context modeling, and a P2 detection head is added to construct a P2–P5 four-scale framework. Experimental results on a self-built dual-modal dataset show that the proposed method achieves a precision of 98.83%, recall of 93.18%. With 4.84 M parameters and 8.3 GFLOPs, the proposed method reaches an inference speed of 71.35 FPS, achieving a favorable balance among detection performance, model complexity, and real-time capability. Optimal performance is also achieved under low-illumination and complex-background scenarios.