Modeling Dynamic Obstacle Avoidance Strategy of Drone Swarms Combined with Multi-Agent Reinforcement Learning
Main Article Content
Abstract
This paper proposes the Locally-decoupled and Embedding-enhanced Multi-Agent Deep Deterministic Policy Gradient (LDE-MADDPG) algorithm to address poor scalability and delayed response in drone swarm dynamic obstacle avoidance under complex cooperative environments. Such autonomous coordination capabilities are also important for distributed sensing, wireless networking, and electromagnetic information exchange in future intelligent aerial systems. The algorithm introduces three key innovations beyond standard MADDPG: a Graph Attention Network module that encodes variable-length observations into fixed-dimensional embeddings for swarm-size generalization; a dual-path critic with a global branch guiding policy updates and a local branch specializing in obstacle avoidance evaluation; and a hierarchical reward integrating multi-objective signals. Evaluated across eight static and dynamic obstacle scenarios, LDE-MADDPG achieves significantly lower collision rates (2.1%–4.2% in static scenarios and 3.8%–7.2% in dynamic scenarios) than state-of-the-art baselines and reaches a 97.5% mission completion rate in 100 random scenarios. The proposed framework demonstrates robust scalability and real-time coordination capability for dynamic environments, while providing a reliable decision-making paradigm for intelligent multi-agent systems operating in communication-intensive and electromagnetically complex application scenarios.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
G. Kumar, A. Anwar, A. Dikshit, A. Poddar, U. Soni, and W. K. Song, “Obstacle avoidance for a swarm of unmanned aerial vehicles operating on particle swarm optimization: a swarm intelligence approach for search and rescue missions,” Journal of the Brazilian Society of Mechanical Sciences and Engineering, vol. 44, no. 2, Art. no. 56, 2022, doi: 10.1007/s40430-022-03362-9.
D. Marek, P. Biernacki, J. Szyguła, M. Paszkuta, and M. Szczygiel, “Collision avoidance mechanism for swarms of drones,” Sensors, vol. 25, no. 4, pp. 1141-1141, 2025, doi: 10.3390/s25041141.
L. Zhao, B. Chen, and F. Hu, “Research on cooperative obstacle avoidance decision making of unmanned aerial vehicle swarms in complex environments under end-edge-cloud collaboration model,” Drones, vol. 8, no. 9, pp. 461-461, 2024, doi: 10.3390/drones8090461.
R. A. Saeed, M. Omri, S. Abdel-Khalek, E. S. Ali, and M. F. Alotaibi, “Optimal path planning for drones based on swarm intelligence algorithm,” Neural Computing and Applications, vol. 34, no. 12, pp. 10133-10155, 2022, doi: 10.1007/s00521-022-06998-9.
G. Jia and J. Wang, “A review of research methods for UAV swarm mission planning,” Systems Engineering & Electronics, vol. 43, no. 1, pp. 99-99, 2021, doi: 10.3969/j.issn.1001-506x.2021.01.13.
X. Fu and J. Pan, “Distributed formation control of UAV swarm to avoid dynamic obstacles,” Systems Engineering & Electronics, vol. 44, no. 2, pp. 529-529, 2022, doi: 10.12305/j.issn.1001506x.2022.02.22.
D. Yan, W. Zhang, H. Chen, and J. Shi, “Research on sliding mode consensus formation control of multi-UAV with time delay and interference constraints,” Journal of Northwestern Polytechnical University, vol. 38, no. 2, pp. 420-426, 2020, doi: 10.1051/jnwpu/20203820420.
Y. Wu and T. Liang, “UAV formation control based on improved consensus algorithm,” Acta Aeronautica Sinica, vol. 41, no. 9, pp. 323848-323848, 2020, doi: 10.7527/S10006893.2020.23848.
Z. Xue and T. Gonsalves, “Vision based drone obstacle avoidance by deep reinforcement learning,” Ai, vol. 2, no. 3, pp. 366-380, 2021, doi: 10.3390/ai2030023.
A. Novikov, S. Yakovlev, and I. Gushchin, “Exploring the possibilities of MADDPG for UAV swarm control by simulating in Pac-Man environment,” Radioelectronic and Computer Systems, vol. 2025, no. 1, pp. 327-337, 2025, doi: 10.32620/reks.2025.1.21.
J. Liu, S. Wei, B. Li, T. Wang, W. Qi, X. Han, et al., “Dual-timescale hierarchical MADDPG for Multi-UAV cooperative search,” Journal of King Saud University Computer and Information Sciences, vol. 37, no. 6, pp. 1-17, 2025, doi: 10.1007/s44443-025-00156-6.
E. Zhao, N. Zhou, C. Liu, H. Su, Y. Liu, J. Cong, et al., “Time-aware MADDPG with LSTM for multi-agent obstacle avoidance: A comparative study,” Complex & Intelligent Systems, vol. 10, no. 3, pp. 4141-4155, 2024, doi: 10.1007/s40747-024-01389-0.
J. Li, Z. Yan, K. Yan, Y. Zhao, R. Tan, C. Liang, et al., “UAV ground target detection algorithm based on attention and channel rearrangement,” Journal of Ordnance Equipment Engineering, vol. 45, no. 3, pp. 306-306, 2024, doi: 10.11809/bqzbgcxb2024.03.040.
X. He, X. Shi, J. Hu, and Y. Wang, “Multi-robot navigation with graph attention neural network and hierarchical motion planning,” Journal of Intelligent & Robotic Systems, vol. 109, no. 2, pp. 25-25, 2023, doi: 10.1007/S10846-023-01959-3.
Q. Wang, D. Zhuang, and H. Xie, “Identification of influential nodes for drone swarm based on graph neural networks,” Neural Processing Letters, vol. 53, no. 6, pp. 4073-4096, 2021, doi: 10.1007/S11063-021-10583-X.
X. Zhang, J. Zheng, T. Su, H. Liu, and Q. Gao, “Planning method for cooperative search and tracking mission of UAV swarm,” Radar Science and Technology, vol. 20, no. 5, pp. 480-491, 2022, doi: 10.3969/j.issn.1672-2337.2022.05.002.
H. Dong, J. Yang, S. Li, J. Wang, and Z. Duan, “Research progress of robot motion control based on deep reinforcement learning,” Control and Decision, vol. 37, no. 2, pp. 278-292, 2022, doi: 10.13195/j.kzyjc.2020.1382.
Y. Jia, Y. Song, B. Xiong, J. Cheng, W. Zhang, S. Yang, et al., “Hierarchical perception-improving for decentralized multi-robot motion planning in complex scenarios,” IEEE Transactions on Intelligent Transportation Systems, vol. 25, no. 7, pp. 6486-6500, 2024, doi: 10.1109/TITS.2023.3344518.
W. Wang, Y. Chen, Y. Zhang, Y. Chen, and Y. Du, “Collaborative Search Algorithm for Multi-UAVs Under Interference Conditions: A Multi-Agent Deep Reinforcement Learning Approach,” Drones, vol. 9, no. 6, pp. 445-445, 2025, doi: 10.3390/drones9060445.
D. Wei, L. Zhang, Q. Liu, H. Chen, and J. Huang, “UAV Swarm Cooperative Dynamic Target Search: A MAPPO-Based Discrete Optimal Control Method,” Drones, vol. 8, no. 6, pp. 214-214, 2024, doi: 10.3390/drones8060214.
P. Zhang and G. Li, “A cooperative dynamic target search approach for multi-UAV systems utilizing the MAPPO algorithm,” Discover Artificial Intelligence, vol. 5, no. 1, pp. 153-153, 2025, doi: 10.1007/S44163-025-00411-9.
C. Ning, J. Fan, and S. Sun, “A review of research on multi-UAV collaborative planning,” Journal of Computer Engineering & Applications, vol. 61, no. 1, pp. 42-42, 2025, doi: 10.3778/j.issn.1002-8331.2405-0110.
Y. Sun, C. Yan, X. Xiang, D. Tang, H. Zhou, J. Jiang, et al., “Multi-UAV cooperative capture method based on hierarchical reinforcement learning,” Control Theory & Applications/Kongzhi Lilun Yu Yinyong, vol. 42, no. 1, pp. 96-96, 2025, doi: 10.7641/CTA.2024.30439.
J. Wu, C. Luo, Y. Luo, and K. Li, “Distributed UAV swarm formation and collision avoidance strategies over fixed and switching topologies,” IEEE transactions on cybernetics, vol. 52, no. 10, pp. 10969-10979, 2021, doi: 10.1109/TCYB.2021.3132587.
S. Argiliana, E. Ekawati, and F. Mukhlish, “Adaptive Strategies for Dynamic Obstacle Avoidance and Formation Control in Multi-Agent Drone Systems: A Review,” Journal of Robotics and Control (JRC), vol. 6, no. 4, pp. 1710-1720, 2025, doi: 10.18196/jrc.v6i4.26243.
R. Fan, J. Wang, W. Han, and B. Xu, “UAV swarm control based on hybrid bionic swarm intelligence,” Guidance, Navigation and Control, vol. 3, no. 02, pp. 2350008-2350008, 2023, doi: 10.1142/S2737480723500085.
H. Muller, V. Niculescu, and T. Polonelli, “Robust and efficient depth-based obstacle avoidance for autonomous miniaturized uavs,” IEEE Transactions on Robotics, vol. 39, no. 6, pp. 4935-4951, 2023, doi: 10.1109/TRO.2023.3315710.
M. H. Harun, S. S. Abdullah, M. S. M. Aras, and M. B. Bahar, “Collision avoidance control for Unmanned Autonomous Vehicles (UAV): Recent advancements and future prospects,” Indian Journal of Geo-Marine Sciences (IJMS), vol. 50, no. 11, pp. 873-883, 2022.
C. C. Ekechi, T. Elfouly, A. Alouani, and T. Khattab, “A Survey on UAV Control with Multi-Agent Reinforcement Learning,” Drones, vol. 9, no. 7, pp. 484-484, 2025, doi: 10.3390/drones9070484.
X. Wei, W. Cui, X. Huang, L. Yang, X. Geng, Z. Tao, et al., “Hierarchical RNNs with graph policy and attention for drone swarm,” Journal of Computational Design and Engineering, vol. 11, no. 2, pp. 314-326, 2024, doi: 10.1093/jcde/qwae031.
M. M. Alam, S. A. Trina, T. Hossain, M. S. Ahmed, and M. Y. Arafat, “Variations in Multi-Agent Actor–Critic Frameworks for Joint Optimizations in UAV Swarm Networks: Recent Evolution, Challenges, and Directions,” Drones, vol. 9, no. 2, pp. 153-153, 2025, doi: 10.3390/drones9020153.
R. Tang, J. Tang, M. S. A. Talip, N. K. Aridas, and X. Xu, “Enhanced multi agent coordination algorithm for drone swarm patrolling in durian orchards,” Scientific Reports, vol. 15, no. 1, pp. 9139-9139, 2025, doi: 10.1038/S41598-025-88145-7.
P. Cao, L. Lei, S. Cai, G. Shen, X. Liu, X. Wang, et al., “Computational intelligence algorithms for UAV swarm networking and collaboration: A comprehensive survey and future directions,” IEEE Communications Surveys & Tutorials, vol. 26, no. 4, pp. 2684-2728, 2024, doi: 10.1109/COMST.2024.3395358.
S. A. Mustafa and A. S. M. Kakshar, “Analysis and Prospect of Existing Path Planning Algorithms for Multi-Drone Systems,” QALAAI ZANIST SCIENTIFIC JOURNAL, vol. 10, no. 1, pp. 1508-1543, 2025.