Abstract:
In dense pedestrian detection, the number of targets is large and is often accompanied by interference factors such as occlusion and complex backgrounds, which can easily cause problems such as insufficient detection accuracy, missing detections, and false detections. To address these challenges, a YOLOv10-n modified model, MSD-YOLO (C2f-MCLU, STNet, Dyhead-DCNv4), was developed. The proposed C2f-MCLU module effectively improved the capability of feature expression by establishing a close interdependence between the channel dimension and spatial position. A bidirectional fusion pyramid structure with a small target enhancement was designed to reconstruct the neck network such that the model could extract more subtle features. A Dyhead-DCNv4 detection head was constructed to further improve the recognition ability of severely occluded individuals. The experimental results showed that compared with YOLOv10-n, the accuracy of the improved model on the CrowdHuman and WiderPerson datasets increased by 3.3% and 1.5%, respectively. Furthermore, with only 3.0 M parameters and a computational cost of 10.6 GFLOPs, the network satisfied the requirements for high-precision deployment in resource-constrained environments.