• 中国科技核心期刊
  • Scopus收录期刊
  • DOAJ收录期刊
  • JST收录期刊
  • Euro Pub收录期刊
Turn off MathJax
Article Contents
FU Songchen, BAI Letian, ZHAO Shaojing, MENG Aomeng, LIANG Hong, LI Ta. A Delay-Robust Multi-Agent Reinforcement Learning Approach for Cooperative Target Encirclement[J]. Journal of Unmanned Undersea Systems. doi: 10.11993/j.issn.2096-3920.2025-0173
Citation: FU Songchen, BAI Letian, ZHAO Shaojing, MENG Aomeng, LIANG Hong, LI Ta. A Delay-Robust Multi-Agent Reinforcement Learning Approach for Cooperative Target Encirclement[J]. Journal of Unmanned Undersea Systems. doi: 10.11993/j.issn.2096-3920.2025-0173

A Delay-Robust Multi-Agent Reinforcement Learning Approach for Cooperative Target Encirclement

doi: 10.11993/j.issn.2096-3920.2025-0173
  • Received Date: 2025-12-26
  • Accepted Date: 2026-01-27
  • Rev Recd Date: 2026-01-06
  • Available Online: 2026-07-20
  • To address delayed observations caused by acoustic communication in cooperative operations of multiple unmanned underwater vehicles (UUVs), a delay-robust multi-agent reinforcement learning based cooperative encirclement method is proposed. First, the mechanism of delayed observations is analyzed under common communication scenarios. Second, a turbulent flow field is modeled using the two-dimensional Navier–Stokes equations to construct a simulation environment that reflects realistic task settings. Then, a delay-robust multi-agent reinforcement learning method is proposed, and its constituent modules are described in detail. On this basis, the reward functions for the encirclement task, the network architectures, and the training procedures are designed, and ablation studies are conducted under different tasks and delay conditions. Experimental results demonstrate that the proposed method effectively handles delayed observations and maintains strong performance under varying delay levels, approaching the theoretical upper bound of the no-delay case in some tasks. Furthermore, the ablation results verify the effectiveness of each module in mitigating delayed observations, providing new theoretical support and a practical methodology for multi-UUV cooperative strategies under delayed observations.

     

  • loading
  • [1]
    胡桥, 赵振轶, 冯豪博, 等. AUV智能集群协同任务研究进展[J]. 水下无人系统学报, 2023, 31(2): 189-200.

    Hu Q, Zhao Z Y, Feng H B, et al. Progress of AUV intelligent swarm collaborative task[J]. Journal of Unmanned Undersea Systems, 2023, 31(2): 189-200.
    [2]
    陈昭, 丁一杰, 张治强. 无人潜航器发展历程及运用优势研究[J]. 舰船科学技术, 2024, 46(23): 98-102.

    Chen Z, Ding Y J, Zhang Z Q. Research on the development history and application advantages of unmanned underwater vehicle[J]. Ship Science and Technology, 2024, 46(23): 98-102.
    [3]
    韦良才, 李东升. 无人潜航器集群作业的应用[J]. 电子技术, 2024, 53(9): 310-311.

    Wei L C, Li D S. Application of unmanned underwater vehicle cluster operation[J]. Electronic Technology, 2024, 53(9): 310-311.
    [4]
    王科翔, 卢野, 李银城, 等. 面向海战场下无人集群作战发展研究综述[J]. 现代防御技术, 2025, 53(3): 11-22.

    Wang K X, Lu Y, Li Y C, et al. Review of unmanned swarm operations development for maritime battlefields[J]. Modern Defence Technology, 2025, 53(3): 11-22.
    [5]
    王光诚. 基于多智能体强化学习的无人水下航行器决策方法研究[D]. 吉林大学, 2024: 1-2.
    [6]
    林学航. 基于强化学习的无人潜航器三维路径规划方法研究[D]. 哈尔滨工程大学, 2024: 5.
    [7]
    赵剑楠, 覃琪琪, 李云, 等. 基于分布式强化学习的AUV水下三维洋流目标跟踪控制算法[J]. 电信科学, 2025, 41(10): 88-101.

    Zhao J N, Qin Q Q, Li Y, et al. Distributed reinforcement learning-based AUV 3D underwater current target tracking control algorithm[J]. Telecommunication Science, 2025, 41(10): 88-101.
    [8]
    杨晓康. 基于深度强化学习的多智能体协同围捕控制[D]. 杭州电子科技大学, 2025: 2.
    [9]
    王景璟, 魏维, 王永越, 等. 水下自主潜航器集群协同围捕技术[J]. 指挥与控制学报, 2023, 9(6): 637-650.

    Wang J J, Wei W, Wang Y Y, et al. Autonomous underwater vehicle swarm collaborative hunting technologies[J]. Journal of Command and Control, 2023, 9(6): 637-650.
    [10]
    赵少靖, 付松琛, 白乐天, 等. 基于自适应多目标优化的UUV全覆盖路径规划方法[J]. 水下无人系统学报, 2025, 33(3): 459-472.

    Zhao S J, Fu S C, Bai L T, et al. Adaptive multi-objective optimization-based coverage path planning method for UUVs[J]. Journal of Unmanned Undersea Systems, 2025, 33(3): 459-472.
    [11]
    刘千里, 张松, 李乐琦. 水下无人通信载荷技术综述与应用前景[J]. 舰船电子工程, 2025, 45(1): 7-12.

    Liu QL, Zhang S, Li L Q. Overview of underwater unmanned communication payload technology and application[J]. Ship Electronic Engineering, 2025, 45(1): 7-12.
    [12]
    董佳明. 基于多智能体强化学习的多UUV博弈围捕方法研究[D]. 哈尔滨工程大学, 2024: 6-7.
    [13]
    Wu G, Xu T, Sun Y, et al. Review of multiple unmanned surface vessels collaborative search and hunting based on swarm intelligence[J]. International journal of advanced robotic systems, 2022, 19(2): 17298806221091885.
    [14]
    Zhao Z, Hu Q, Feng H, et al. A cooperative hunting method for multi-AUV swarm in underwater weak information environment with obstacles[J]. Journal of marine science and engineering, 2022, 10(9): 1266. doi: 10.3390/jmse10091266
    [15]
    Li J, Lu H, Zhang H, et al. Dynamic target hunting under autonomous underwater vehicle (AUV) motion planning based on improved dynamic window approach (DWA)[J]. Journal of Marine Science and Engineering, 2025, 13(2): 221. doi: 10.3390/jmse13020221
    [16]
    Cao X, Liu W, Ren L. Underwater target capture based on heterogeneous unmanned system collaboration[J]. IEEE Transactions on Intelligent Vehicles, 2024.
    [17]
    Jiang B, Wang Y, Kong F, et al. Multi-autonomous underwater vehicle trajectory planning in ocean current based on hierarchical hunting and evolutionary learning[J]. ACM Transactions on Intelligent Systems and Technology, 2025, 16(5): 1-26. doi: 10.1145/3757928
    [18]
    Zhang M, Chen H, Cai W. Collaborative hunting method of multi-AUV in 3-D IoUT: Searching, tracking, and encirclement keeping[J]. IEEE Internet of Things Journal, 2025, 12(8): 10958-10973. doi: 10.1109/JIOT.2024.3514632
    [19]
    Wang Z, Du J, Jiang C, et al. Task scheduling for distributed AUV network target hunting and searching: An energy-efficient AoI-aware DMAPPO approach[J]. IEEE Internet of Things Journal, 2022, 10(9): 8271-8285.
    [20]
    Hou X, Xing T, Wang J, et al. Age of information-aware multi-objective optimization for heterogeneous UAV-USV-UUV Networks in Underwater Target Hunting[J]. IEEE Transactions on Mobile Computing, 2025.
    [21]
    Chen J, Wang Y, Zhang Y, et al. Extrinsic-and-intrinsic reward-based multi-agent reinforcement learning for multi-UAV cooperative target encirclement[J]. IEEE Transactions on Intelligent Transportation Systems, 2025.
    [22]
    Wei W, Wang J, Du J, et al. Differential game-based deep reinforcement learning in underwater target hunting task[J]. IEEE Transactions on Neural Networks and Learning Systems, 2023.
    [23]
    伍雨辰. 基于深度强化学习的水声传感器网络时隙可变MAC协议[D]. 天津大学, 2022: 2.
    [24]
    Rashid T, Samvelyan M, De Witt C S, et al. Monotonic value function factorisation for deep multi-agent reinforcement learning[J]. Journal of Machine Learning Research, 2020, 21(178): 1-51.
    [25]
    Oliehoek F A, Amato C. A concise introduction to decentralized POMDPs[M]. Cham, Switzerland: Springer International Publishing, 2016: 14-15.
    [26]
    Mnih V, Kavukcuoglu K, Silver D, et al. Human-level control through deep reinforcement learning[J]. Nature, 2015, 518: 529-533. doi: 10.1038/nature14236
    [27]
    Tampuu A, Matiisen T, Kodelja D, et al. Multiagent cooperation and competition with deep reinforcement learning[J]. PLOS One, 2017, 12(4): e0172395. doi: 10.1371/journal.pone.0172395
    [28]
    Sunehag P, Lever G, Gruslys A, et al. Value-decomposition networks for cooperative multi-agent learning based on team reward[C]//Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems. Stockholm, 2018: 2085-2087.
    [29]
    Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Advances in neural information processing systems, 2017, 30: 6000-6010. doi: 10.1007/978-3-031-84300-6_13
    [30]
    Dey R, Salem F M. Gate-variants of gated recurrent unit (GRU) neural networks[C]//2017 IEEE 60th international midwest symposium on circuits and systems (MWSCAS). Massachusetts, 2017: 1597-1600.
    [31]
    Ha D, Dai A, Le Q V. Hypernetworks[EB/OL]. 2016[2025-12-22]. https://arxiv.org/abs/1609.09106.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(7)  / Tables(6)

    Article Metrics

    Article Views(109) PDF Downloads(24) Cited by()
    Proportional views
    Related
    Service
    Subscribe

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return