Combining DETR Multi-scale Perception Structure to Improve Target Boundary Extraction Accuracy in Image Data
Main Article Content
Abstract
Accurate target boundary extraction is essential for high-precision image interpretation in intelligent sensing systems and provides important technical support for electromagnetic imaging, remote sensing, and vision-assisted signal perception applications. Owing to the limited feature representation capability of the original Detection Transformer (DETR) when processing objects at different scales, target boundary extraction accuracy remains insufficient, particularly in scenarios involving fine-grained contour localization. To address this issue, an improved DETR framework based on a multi-scale perception structure is proposed. A multi-scale encoding architecture integrated with a feature pyramid network is employed to capture features at multiple resolutions, enhancing boundary-aware spatial representations through an edge completion mechanism. A scale-aware multi-head attention module is incorporated into the encoder to preserve scale consistency during global feature modeling and reduce semantic drift. Furthermore, a boundary regression head combined with positional embedding jointly exploits spatial and semantic information to improve contour localization, while an IoU-aware loss function optimizes prediction accuracy in overlapping regions. Experimental results demonstrate that the proposed method increases the average IoU from 0.81 to 0.86 and improves target boundary AP from 44.9% to 48.1%. Compared with Deformable-DETR and Elastic-DETR, it also achieves lower boundary error ratios, confirming that the proposed multi-scale perception mechanism effectively enhances high-precision boundary fitting and offers practical value for intelligent visual sensing systems in engineering applications.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
O. Akcay, A. C. Kınacı, E. O. Avsar, and U. Aydar, “Boundary Extraction Based on Dual Stream Deep Learning Model in High Resolution Remote Sensing Images,” Journal of Advanced Research in Natural and Applied Sciences, vol. 7, no. 3, pp. 358-368, 2021, doi: 10.28979/jarnas.911130.
L. Ma, Y. Li, J. Li, J. M. Junior, W. N. Goncalves, and M. A. Chapman, “Boundarynet: Extraction and completion of road boundaries with deep learning using mobile laser scanning point clouds and satellite imagery,” IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 6, pp. 5638-5654, 2022, doi: 10.1109/TITS.2021.3055366.
A. Abdollahi, B. Pradhan, and A. Alamri, “SC-RoadDeepNet: A new shape and connectivity-preserving road extraction deep learning-based network from remote sensing data,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-15, 2022, doi: 10.1109/TGRS.2022.3143855.
Z. Wang, Y. L. Li, X. Chen, H. Zhao, and S. Wang, “Uni3detr: Unified 3d detection transformer,” in Oh A, Naumann T, Globerson A, Saenko K, Hardt M, Levine S, editors. Advances in Neural Information Processing Systems 36. Proceedings of the 37th Conference on Neural Information Processing Systems, 10-16 December 2023, New Orleans, Louisiana, USA. New York, NY, USA: Curran Associates, Inc., 2023, pp. 39876-39896, doi: 10.48550/arXiv.2310.05699.
L. Dai, H. Liu, H. Tang, Z. Wu, and P. Song, “AO2-DETR: Arbitrary-oriented object detection transformer,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 5, pp. 2342-2356, 2022, doi: 10.1109/TCSVT.2022.3222906.
M. A. Munir, S. H. Khan, M. H. Khan, M. Ali, and F. S. Khan, “Cal-DETR: Calibrated detection transformer,” in Oh A, Naumann T, Globerson A, Saenko K, Hardt M, Levine S, editors. Advances in Neural Information Processing Systems 36. Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS 2023), 10-16 December 2023, New Orleans, Louisiana, USA. New York, NY, USA: Curran Associates, Inc., 2023, pp. 71619-71631, doi: https://doi.org/10.48550/arXiv.2310.05699.
Y. Zeng, Y. Chen, X. Yang, Q. Li, and J. Yan, “ARS-DETR: Aspect ratio-sensitive detection transformer for aerial oriented object detection,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1-15, 2024, doi: 10.1109/TGRS.2024.3364713.
X. Song, Z. Hua, and J. Li, “Remote sensing image change detection transformer network based on dual-feature mixed attention,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-16, 2022, doi: 10.1109/TGRS.2022.3209972.
Y. Liu, P. Nand, M. A. Hossain, M. Nguyen, and W. Q. Yan, “Sign language recognition from digital videos using feature pyramid network with detection transformer,” Multimedia Tools and Applications, vol. 82, no. 14, pp. 21673-21685, 2023, doi: 10.1007/s11042-023-14646-0.
Q. Li, G. Zhong, C. Xie, and R. Hedjam, “Weak edge identification network for ocean front detection,” IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1-5, 2022, doi: 10.1109/lgrs.2021.3051203.
G. Singh and K. Singh, “Chroma key foreground forgery detection under various attacks in digital video based on frame edge identification,” Multimedia Tools and Applications, vol. 81, no. 1, pp. 1419-1446, 2022, doi: 10.1007/s11042-021-11380-3.
M. S. Sivapriya and S. Suresh, “ViT-DexiNet: A vision transformer-based edge detection operator for small object detection in SAR images,” International Journal of Remote Sensing, vol. 44, no. 22, pp. 7057-7084, 2023, doi: 10.1080/01431161.2023.2277167.
Y. Xie, Z. Tu, T. Yang, Y. Zhang, and X. Zhou, “EdgeFormer: Local patch-based edge detection transformer on point clouds,” Pattern Analysis and Applications, vol. 28, no. 1, pp. 1-13, 2025, doi: 10.1007/s10044-024-01386-6.
Z. Liu and J. Cheng, “Cb-fpn: Object detection feature pyramid network based on context information and bidirectional efficient fusion,” Pattern Analysis and Applications, vol. 26, no. 3, pp. 1441-1452, 2023, doi: 10.1007/s10044-023-01173-9.
Q. Liu, J. Bi, J. Zhang, X. Bu, and N. Hanajima, “B-FPN SSD: An SSD algorithm based on a bidirectional feature fusion pyramid,” The Visual Computer, vol. 39, no. 12, pp. 6265-6277, 2023, doi: 10.1007/s00371-022-02727-4.
X. Fu, Z. Yuan, T. Yu, and Y. Ge, “DA-FPN: Deformable convolution and feature alignment for object detection,” Electronics, vol. 12, no. 6, pp. 1354-1368, 2023, doi: 10.3390/electronics12061354.
H. Wu, J. Liu, T. Jiang, Q. Zou, S. Qi, Z. Cui, et al., “AttentionMGT-DTA: A multi-modal drug-target affinity prediction using graph transformer and attention mechanism,” Neural Networks, vol. 169, pp. 623-636, 2024, doi: 10.1016/j.neunet.2023.11.018.
Y. Ding and M. Jia, “Convolutional transformer: An enhanced attention mechanism architecture for remaining useful life estimation of bearings,” IEEE Transactions on Instrumentation and Measurement, vol. 71, pp. 1-10, 2022, doi: 10.1109/tim.2022.3181933.
F. P. An and J. E. Liu, “Medical image segmentation algorithm based on multilayer boundary perception-self attention deep learning model,” Multimedia Tools and Applications, vol. 80, no. 10, pp. 15017-15039, 2021, doi: 10.1007/s11042-021-10515-w.
H. Wang, Y. Jin, H. Ke, and X. Zhang, “DDH-YOLOv5: Improved YOLOv5 based on Double IoU-aware Decoupled Head for object detection,” Journal of Real-Time Image Processing, vol. 19, no. 6, pp. 1023-1033, 2022, doi: 10.1007/s11554-022-01241-z.
Y. Shin, J. Park, H. Song, S. Yoon, B. S. Lee, and J. G. Lee, “Exploiting representation curvature for boundary detection in time series,” in Globerson A, Mackey L, Belgrave D, Fan A, Paquet U, Tomczak J, Zhang C, editors. Advances in Neural Information Processing Systems 37. Proceedings of the 38th Conference on Neural Information Processing Systems, 10-15 December 2024, Vancouver, Canada. New York: Curran Associates, Inc., 2024, pp. 5974-5995, doi: 10.52202/079017-0194.
Y. Ai, R. Song, C. Huang, C. Cui, B. Tian, and L. Chen, “A real-time road boundary detection approach in surface mine based on meta random forest,” IEEE Transactions on Intelligent Vehicles, vol. 9, no. 1, pp. 1989-2001, 2024, doi: 10.1109/tiv.2023.3296767.
M. Ahmed, N. El-Sheimy, H. Leung, and A. Moussa, “Enhancing object detection in remote sensing: A hybrid yolov7 and transformer approach with automatic model selection,” Remote Sensing, vol. 16, no. 1, pp. 51-67, 2024, doi: 10.3390/rs16010051.
K. M. Elgamily, M. A. Mohamed, A. M. Abou-Taleb, and M. M. Ata, “Enhanced object detection in remote sensing images by applying metaheuristic and hybrid metaheuristic optimizers to YOLOv7 and YOLOv8,” Scientific Reports, vol. 15, no. 1, pp. 7226-7256, 2025, doi: 10.1038/s41598-025-89124-8.
H. Nguyen, T. Q. Ngo, H. T. T. Uyen, and M. K. Duong, “Enhanced object recognition from remote sensing images based on hybrid convolution and transformer structure,” Earth Science Informatics, vol. 18, no. 2, pp. 228, 2025, doi: 10.1007/s12145-025-01751-x.
Y. Zang, C. Fu, D. Yang, H. Li, C. Ding, and Q. Liu, “Transformer fusion and histogram layer multispectral pedestrian detection network,” Signal, Image and Video Processing, vol. 17, no. 7, pp. 3545-3553, 2023, doi: 10.1007/s11760-023-02579-y.
K. Huang, M. Wen, C. Wang, and L. Ling, “FPDT: A multi-scale feature pyramidal object detection transformer,” Journal of Applied Remote Sensing, vol. 17, no. 2, pp. 026510-026510, 2023, doi: 10.1117/1.jrs.17.026510.
X. Zhang, Z. Chen, J. Zhang, T. Liu, and D. Tao, “Learning general and specific embedding with transformer for few-shot object detection,” International Journal of Computer Vision, vol. 133, no. 2, pp. 968-984, 2025, doi: 10.1007/s11263-024-02199-0.
J. Liu, Y. Xia, J. Feng, and P. Bai, “A Novel Building Extraction Network via Multi-Scale Foreground Modeling and Gated Boundary Refinement,” Remote Sensing, vol. 15, no. 24, pp. 5638-5665, 2023, doi: 10.3390/rs15245638.
Q. Yuan, “Building rooftop extraction from high resolution aerial images using multiscale global perceptron with spatial context refinement,” Scientific Reports, vol. 15, no. 1, pp. 6499-6513, 2025, doi: 10.1038/s41598-025-91206-6.
H. Zhang, X. Zheng, N. Zheng, and W. Shi, “A multiscale and multipath network with boundary enhancement for building footprint extraction from remotely sensed imagery,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, vol. 15, pp. 8856-8869, 2022, doi: 10.1109/jstars.2022.3214485.
G. Yan, H. Jing, H. Li, H. Guo, and S. He, “Enhancing building segmentation in remote sensing images: Advanced multi-scale boundary refinement with MBR-HRNet,” Remote Sensing, vol. 15, no. 15, pp. 3766-3784, 2023, doi: 10.3390/rs15153766.
J. Chang, X. He, P. Li, T. Tian, X. Cheng, M. Qiao, et al., “Multi-scale attention network for building extraction from high-resolution remote sensing images,” Sensors, vol. 24, no. 3, pp. 1010-1026, 2024, doi: 10.3390/s24031010.
Z. Wang, Y. Zhou, F. Wang, S. Wang, G. Qin, W. Zou, et al., “A multi-scale edge constraint network for the fine extraction of buildings from remote sensing images,” Remote Sensing, vol. 15, no. 4, pp. 927-946, 2023, doi: 10.3390/rs15040927.