Multi-Agent Collaborative Search Path Planning Based on Reinforcement Learning
-
摘要: 针对多平台协同搜索路径规划优化问题, 以最大化累积探测概率为目标, 构建多智能体强化学习模型, 提出一种基于深度Q网络的多平台协同搜索算法Joint-DQN。该算法通过设计经验知识共享机制提升了多平台间的协作效率与稳定性; 引入冲突检测机制并给予惩罚, 有效解决了多智能体协作研究中路径冲突频繁的问题; 同时设计了复合奖励函数, 提升了搜索覆盖率并降低重复搜索率。仿真实验结果表明, 该算法在静态与动态目标场景下均能有效引导搜索平台沿目标最可能存在方向高效移动, 在高效搜索的同时快速提升累积探测概率, 为多平台协同搜索路径规划提供理论支撑和指导。Abstract: To address the problem of optimizing multi-agent collaborative search path planning with maximizing cumulative detection probability, a multi-agent reinforcement learning model is developed. A multi-agent collaborative search algorithm based on Deep Q-Network, Joint-DQN, is proposed. It enhances the collaboration efficiency and stability among multiple platforms by designing an experience knowledge sharing mechanism; introduces a conflict detection mechanism and imposes penalties, effectively addressing the frequent path conflicts in multi-agent collaboration research; and designs a composite reward function to improve search coverage and reduce the rate of duplicate searches. Simulation experimental results demonstrate that this algorithm can effectively guide search platforms to avoid obstacles while efficiently moving in the direction where the target is most likely to be found, both in static and dynamic target scenarios. It rapidly enhances the cumulative detection probability while conducting efficient search, providing theoretical support and valuable guidance for multi-agent collaborative search path planning.
-
表 1 实验参数列表
Table 1. List of experimental parameters
参数 值 搜索平台数量/台 4 POD 0.8 POC 服从二维高斯分布 障碍物密度 5% α 0.01 γ 0.99 ε 0.3 总迭代次数/轮 500 最大探索步数/步 20 单次训练样本/个 128 经验回放池容量/个 1000 τ 0.005 表 2 算法性能对比
Table 2. Comparison of algorithm performance
性能评估指标 JDQN IDQN QMIX VDN CDP 1.0 1.0 0.9978 0.8937 平均冲突次数 4.3 6.4 4.3 2.7 搜索覆盖率/% 18.57 17.26 13.08 11.13 重复搜索率/% 28.88 33.50 51.38 59.38 结束搜索平均步数 14.8 14.4 18.2 19.3 -
[1] 王中, 温志文, 蔡卫军. 水下无人航行器编队协同搜索研究综述[J]. 舰船科学技术, 2024, 46(2): 57-62. doi: 10.3404/j.issn.1672-7649.2024.02.010Wang Z, Wen Z W, Cai W J. Review of research on unmanned underwater vehicle formation cooperative search[J]. Ship Science and Technology, 2024, 46(2): 57-62. doi: 10.3404/j.issn.1672-7649.2024.02.010 [2] 范学满, 薛昌友, 张会. 基于多种群遗传算法的多UUV任务分配方法[J]. 水下无人系统学报, 2022, 30(5): 621-630. doi: 10.11993/j.issn.2096-3920.202107001Fan X M, Xue C Y, Zhang H. Task assignment method for multiple UUVs based on multi-population genetic algorithm[J]. Journal of Unmanned Undersea Systems, 2022, 30(5): 621-630. doi: 10.11993/j.issn.2096-3920.202107001 [3] Ni J, Yang L, Wu L, et al. An improved spinal neural system-based approach for heterogeneous auvs cooperative hunting[J]. International Journal of Fuzzy Systems, 2018, 20(2): 672-686. doi: 10.1007/s40815-017-0395-x [4] 曹迟, 史文涛, 王百合, 等. 无人水下航行器反潜作战模型仿真[J]. 水下无人系统学报, 2025, 33(1): 156-163. doi: 10.11993/j.issn.2096-3920.2024-0116Cao C, Shi W T, Wang B H, et al. Simulation of anti-submarine warfare model of unmanned undersea vehicles[J]. Journal of Unmanned Undersea Systems, 2025, 33(1): 156-163. doi: 10.11993/j.issn.2096-3920.2024-0116 [5] 李小新, 陈华林. 水下机器人科学考察技术应用[J]. 船舶工程, 2021, 43(9): 8-13.Li X X, Chen H L. Application of underwater robot scientific investigation technology[J]. Ship Engineering, 2021, 43(9): 8-13. [6] 于传合, 高卫华, 李树奎, 等. 面向水下目标搜救的潜器探测模型构建与性能分析[J]. 控制与决策, 2025, 40(1): 170-179. doi: 10.13195/j.kzyjc.2024.0324Yu C H, Gao W H, Li S K, et al. Detection model construction and performance analysis for target search and rescue via autonomous underwater vehicles[J]. Control and Decision, 2025, 40(1): 170-179. doi: 10.13195/j.kzyjc.2024.0324 [7] Koopman B O. The theory of search. II. Target detection[J]. Operations research, 1956, 4(5): 503-531. doi: 10.1287/opre.4.5.503 [8] Kan Y C. Optimal search of a moving target[J]. Operations Research, 1977, 25(5): 864-870. [9] Hohzaki R, Iida K. Optimal ambushing search for a moving target[J]. European Journal of Operational Research, 2001, 133(1): 120-129. doi: 10.1016/S0377-2217(00)00187-9 [10] 轩永波, 黄长强, 吴文超, 等. 运动目标的多无人机编队覆盖搜索决策[J]. 系统工程与电子技术, 2013, 35(3): 539-544. doi: 10.3969/j.issn.1001-506X.2013.03.15Xuan Y B, Huang C Q, Wu W C, et al. Coverage search strategies for moving targets using multiple unmanned aerial vehicle teams[J]. Systems Engineering and Electronics, 2013, 35(3): 539-544. doi: 10.3969/j.issn.1001-506X.2013.03.15 [11] 曾国奇, 白宇, 林伟, 等. 地面运动目标的多UAV协同搜索方法[J]. 系统工程与电子技术, 2018, 40(7): 1498-1505.Zeng G Q, Bai Y, Lin W, et al. Multi-UAV cooperative search method for ground moving targets[J]. Systems Engineering and Electronics, 2018, 40(7): 1498-1505. [12] 岳伟, 席云, 关显赫. 基于多蚁群协同搜索算法的多AUV路径规划[J]. 水下无人系统学报, 2020, 28(5): 505-511.Yue W, Xi Y, Guan X H. Path planning of multi-AUVs based on multi-ant colony cooperative search algorithm[J]. Journal of Unmanned Undersea Systems, 2020, 28(5): 505-511. [13] Pham T H, Bestaoui Y, Mammar S. Aerial robot coverage path planning approach with concave obstacles in precision agriculture[C]//2017 Workshop on Research, Education and Development of Unmanned Aerial Systems (RED-UAS). IEEE, 2017: 43-48. [14] 丁文俊, 柴亚军, 杨宇贤, 等. 基于空海异构无人平台的水下目标搜索与跟踪[J]. 水下无人系统学报, 2024, 32(2): 237-249.Ding W J, Chai Y J, Yang Y X, et al. Underwater target search and tracking based on air-sea heterogeneous unmanned platform[J]. Journal of Unmanned Undersea Systems, 2024, 32(2): 237-249. [15] 刘和祥, 边信黔, 李娟, 等. 基于前视声纳信息的AUV局部路径规划研究[J]. 微计算机信息, 2007(23): 243-245.Liu H X, Bian X Q, Li J, et al. Study of local path planning based on forward looking sonar for AUV[J]. Microcomputer Information, 2007(23): 243-245. [16] Noguchi Y, Kuranaga Y, Maki T. Adaptive navigation of a high speed autonomous underwater vehicle using low-cost sensors for low-altitude survey[C]//2017 IEEE Underwater Technology(UT), IEEE, 2017: 1-4. [17] Yang Q Q, Gao Y Y, Guo Y, et al. Target search path planning for naval battle field based on deep reinforcement learning[J]. Systems Engineering and Electronics, 2022, 44(11): 3486-3495. [18] 杨鹏程, 杨清清, 高盈盈, 等. 基于强化学习的海上移动目标搜索路径规划[J]. 系统工程与电子技术, 2026, 48(2): 515-523.Yang P C, Yang Q Q, Gao Y Y, et al. Path planning for maritime moving target search based on reinforcement learning[J]. Systems Engineering and Electronics, 2026, 48(2): 515-523. [19] Li J , Jiang X , Zhang H , et al. Multi-joint adaptive control enhanced reinforcement learning for unmanned ship[J]. Ocean Engineering, 2025, 318: 120121. [20] Zhang C, Wu C, Chen H. Coverage path planning for maritime search based on reinforcement learning with path expansion strategy[C]//2024 7th International Conference on Robotics, Control and Automation Engineering(RCAE), IEEE, 2024: 531-536. [21] Sun Y, Zhang R, Liang W, et al. Multi-agent cooperative search based on reinforcement learning[C]//2020 3rd International Conference on Unmanned Systems(ICUS). IEEE, 2020: 891-896. [22] 邢博闻, 张昭夷, 王世明, 等. 基于深度强化学习的多无人艇协同目标搜索算法[J]. 兵器装备工程学报, 2023, 44(11): 118-125.Xing B W, Zhang Z Y, Wang S M, et al. Multi-USV cooperative target search algorithm based on deep reinforcement learning[J]. Journal of Ordnance Equipment Engineering, 2023, 44(11): 118-125. [23] Guo S Y, Zhang X G, Zheng Y S, et al. An Autonomous path planning model for unmanned ships based on deep reinforcement learning[J]. Sensors, 2020, 20(2): 426. [24] Loane E P, Richardson H R, Boylan E S. Theory of cumulative detection probability[R]. Washington DC: U. S. Navy Bureau of Ordnance, 1964. [25] Tan M. Multi-agent reinforcement learning: Independent vs. cooperative agents[J]. Machine Learning Proceedings, 1993: 330-337. [26] Rashid T, Samvelyan M, De Witt C S, et al. Monotonic value function factorisation for deep multi-agent reinforcement learning[J]. Journal of Machine Learning Research, 2020, 21(178): 1-51. [27] Sunehag P, Lever G, Gruslys A, et al. Value-decomposition networks for cooperative multi-agent learning based on team reward[C]//Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems. 2018: 2085-2087. -

下载: