YOLO-SWIFT: A Small-object Wavelet-Integrated Feature Transform Network for Object Detection in UAV Aerial Imagery
Main Article Content
Abstract
Accurate detection of small-scale targets in UAV aerial imagery remains challenging due to severe scale variation, background interference, and feature degradation caused by repeated downsampling. To address these issues, this study proposes YOLO-SWIFT, a wavelet-integrated feature transformation network designed for small-object detection. First, a Position-Aware Coordinate Downsampling module is developed to preserve critical spatial information during feature compression through coordinate attention and skip connections. Second, a High-Resolution Feature Aggregation Network is introduced to establish an additional high-resolution detection branch for enhanced cross-scale feature fusion. Third, a Wavelet Bottleneck Enhancement module incorporating multi-level wavelet decomposition and a High-Frequency Retention pathway is designed to improve fine-detail representation while expanding the effective receptive field. Finally, an Adaptive Scale-aware Regression IoU loss function is proposed to dynamically balance localization and shape-consistency constraints for small targets. Experimental evaluation on the VisDrone2019 benchmark demonstrates that YOLO-SWIFT achieves 38.7% mAP50 and 22.5% mAP50:95. The proposed framework provides an effective solution for intelligent aerial sensing and offers potential applications in electromagnetic imaging, remote sensing interpretation, and autonomous surveillance systems.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
T. Kattenborn, J. Leitloff, F. Schiefer, and S. Hinz, “Review on convolutional neural networks (CNN) in vegetation remote sensing,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 173, pp. 24-49, 2021, doi: 10.1016/j.isprsjprs.2020.12.010.
B. F. Spencer, V. Hoskere, and Y. Narazaki, “Advances in computer vision-based civil infrastructure inspection and monitoring,” Engineering, vol. 5, no. 2, pp. 199-222, 2019, doi: 10.1016/j.eng.2018.11.030.
B. Mishra, D. Garg, P. Narang, and V. Mishra, “Drone-surveillance for search and rescue in natural disaster,” Computer Communications, vol. 156, pp. 1-10, 2020, doi: 10.1016/j.comcom.2020.03.012.
C. Lyu, W. Zhang, H. Huang, et al., “RTMDet: An empirical study of designing real-time object detectors,” arXiv preprint arXiv:2212.07784, 2022.
X. Zhu, W. Su, L. Lu, B. Li, X. Wang, and J. Dai, “Deformable DETR: Deformable transformers for end-to-end object detection,” In International Conference on Learning Representations, 2020.
Y. Tang, K. Han, J. Guo, C. Xu, C. Xu, and Y. Wang, “GhostNetV2: Enhance cheap operation with long-range attention,” Advances in Neural Information Processing Systems, vol. 35, pp. 9969-9982, 2023, doi: 10.52202/068431-0724.
C. Wang, W. He, Y. Nie, J. Guo, C. Liu, K. Han, and Y. Wang, “Gold-YOLO: Efficient object detector via gather-and-distribute mechanism,” Advances in Neural Information Processing Systems, vol. 36, pp. 51094-51106, 2023, doi: 10.52202/075280-2224.
Z. Chen, C. Yang, Q. Li, F. Zhao, Z. J. Zha, and F. Wu, “Disentangle your dense object detector,” In Proceedings of the 31st ACM International Conference on Multimedia, pp. 4939-4950, 2023, doi: 10.1145/3474085.3475351.
X. Zhu, S. Lyu, X. Wang, and Q. Zhao, “TPH-YOLOv5: Improved YOLOv5 based on transformer prediction head for object detection on drone-captured scenarios,” In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pp. 2778-2788, 2021, doi: 10.1109/ICCVW54120.2021.00312.
P. Zhu, L. Wen, X. Bian, H. Ling, and Q. Hu, “Vision meets drones: A challenge,” arXiv preprint arXiv:1804.07437, 2018.
T. Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2117-2125, 2017, doi: 10.1109/CVPR.2017.106.
S. Woo, J. Park, J. Y. Lee, and I. S. Kweon, “CBAM: Convolutional block attention module,” In Proceedings of the European Conference on Computer Vision, pp. 3-19, 2018, doi: 10.1007/978-3-030-01234-2_1.
Z. Liu, Y. Lin, Y. Cao, et al., “Swin transformer: Hierarchical vision transformer using shifted windows,” In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10012-10022, 2021, doi: 10.1109/ICCV48922.2021.00986.
A. Howard, M. Sandler, G. Chu, et al., “Searching for MobileNetV3,” In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1314-1324, 2019, doi: 10.1109/ICCV.2019.00140.
S. E. Finder, R. Amoyal, E. Treister, and O. Freifeld, “Wavelet convolutions for large receptive fields,” In European Conference on Computer Vision, pp. 351-367. Springer, 2024, doi: 10.1007/978-3-031-72949-2_21.
X. Li, W. Wang, L. Wu, et al., “Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection,” Advances in Neural Information Processing Systems, vol. 33, pp. 21002-21012, 2020.
S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 8759-8768, 2018, doi: 10.1109/CVPR.2018.00913.
Han Y, Wang C, Luo H,et al.LRDS-YOLO enhances small object detection in UAV aerial images with a lightweight and efficient design[J].SCIENTIFIC REPORTS, vol. 15, no. 1, 2025, doi: 10.1038/s41598-025-07021-6.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7132-7141, 2018, doi: 10.1109/CVPR.2018.00745.
P. Liu, H. Zhang, K. Zhang, L. Lin, and W. Zuo, “Multi-level wavelet-CNN for image restoration,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 773-782, 2019, doi: 10.1109/CVPRW.2018.00121.
C. Y. Wang, A. Bochkovskiy, and H. Y. M. Liao, “YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7464-7475, 2023, doi: 10.1109/CVPR52729.2023.00721.
A. Wang, H. Chen, L. Liu, et al., “YOLOv10: Real-time end-to-end object detection,” arXiv preprint arXiv:2405.14458, 2024.