Object detection algorithm based on infrared-visible dual-modality feature fusion
Zedong Huang, Kangkang Du, Xiaohuang Hu, Jianfeng Li
CFM-YOLO, an infrared–visible dual-modality detection algorithm based on YOLOv11, is proposed to improve UAV object detection under adverse illumination and complex aerial backgrounds. The network is redesigned from three aspects: cross-modal feature extraction, lightweight feature fusion, and small-object-oriented detection. First, a Cross-Modality Fusion Mamba (CFM) module is introduced to promote channel-level interaction between visible and infrared features and to model long-range spatial dependencies with selective state-space modeling. Second, a lightweight feature fusion network is used to improve multi-scale information transmission while limiting redundant computation. Third, a P2 detection layer, Ghost convolution, and Focal-WIoU loss are incorporated to enhance small-object localization and alleviate the effect of imbalanced bounding-box samples. Quantitative experiments on the DroneVehicle dataset show that CFM-YOLO achieves 81.6% mAP@0.5 and 60.3% mAP@0.5:0.95, improving over the YOLOv11n-dual/base baseline by 8.4 and 5.5 percentage points, respectively. Qualitative results on the LLVIP dataset further indicate that the proposed method can reduce several missed detections in low-light pedestrian scenes. These results suggest that CFM-YOLO provides a competitive trade-off between detection accuracy and computational cost for UAV-based infrared–visible object detection.