• 中国科技核心期刊
  • Scopus收录期刊
  • DOAJ收录期刊
  • JST收录期刊
  • Euro Pub收录期刊

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

基于强化学习的多平台协同搜索路径规划

贺聪炜 谢勇 陈于涛

贺聪炜, 谢勇, 陈于涛. 基于强化学习的多平台协同搜索路径规划[J]. 水下无人系统学报, xxxx, x(x): x-xx doi: 10.11993/j.issn.2096-3920.2025-0163
引用本文: 贺聪炜, 谢勇, 陈于涛. 基于强化学习的多平台协同搜索路径规划[J]. 水下无人系统学报, xxxx, x(x): x-xx doi: 10.11993/j.issn.2096-3920.2025-0163
HE Congwei, XIE Yong, CHEN Yutao. Multi-Agent Collaborative Search Path Planning Based on Reinforcement Learning[J]. Journal of Unmanned Undersea Systems. doi: 10.11993/j.issn.2096-3920.2025-0163
Citation: HE Congwei, XIE Yong, CHEN Yutao. Multi-Agent Collaborative Search Path Planning Based on Reinforcement Learning[J]. Journal of Unmanned Undersea Systems. doi: 10.11993/j.issn.2096-3920.2025-0163

基于强化学习的多平台协同搜索路径规划

doi: 10.11993/j.issn.2096-3920.2025-0163
基金项目: 国家自然科学基金面上项目(62573203)项目.
详细信息
    作者简介:

    贺聪炜(1999-), 男, 硕士, 主要研究方向为系统优化

  • 中图分类号: TJ630; U676.8

Multi-Agent Collaborative Search Path Planning Based on Reinforcement Learning

  • 摘要: 针对多平台协同搜索路径规划优化问题, 以最大化累积探测概率为目标, 构建多智能体强化学习模型, 提出一种基于深度Q网络的多平台协同搜索算法Joint-DQN。该算法通过设计经验知识共享机制提升了多平台间的协作效率与稳定性; 引入冲突检测机制并给予惩罚, 有效解决了多智能体协作研究中路径冲突频繁的问题; 同时设计了复合奖励函数, 提升了搜索覆盖率并降低重复搜索率。仿真实验结果表明, 该算法在静态与动态目标场景下均能有效引导搜索平台沿目标最可能存在方向高效移动, 在高效搜索的同时快速提升累积探测概率, 为多平台协同搜索路径规划提供理论支撑和指导。

     

  • 图  1  海洋搜索环境

    Figure  1.  Ocean search environment

    图  2  栅格地图环境

    Figure  2.  Grid map environment

    图  3  动作空间

    Figure  3.  Action space

    图  4  JDQN算法框架

    Figure  4.  JDQN algorithm framework

    图  5  4种常见冲突类型

    Figure  5.  Four common types of conflicts

    图  6  静态场景下两种算法路径规划结果

    Figure  6.  Path planning results of two algorithms in static scene

    图  7  动态场景下两种算法路径规划结果

    Figure  7.  Path planning results of two algorithms in dynamic scene

    图  8  不同算法的损失函数值箱线图

    Figure  8.  Box graph of loss function values with different algorithms

    图  9  JDQN与JDQN-S算法的结束搜索平均步数对比

    Figure  9.  Comparison of average steps of end search between JDQN and JDQN-S algorithms

    图  10  JDQN与JDQN-C算法的CDP曲线对比

    Figure  10.  Comparison of CDP curves between JDQN and JDQN-C algorithms

    图  11  JDQN与JDQN-E算法的搜索覆盖率对比

    Figure  11.  Comparison of coverage rate between JDQN and JDQN-E algorithms

    图  12  JDQN与JDQN-E算法的重复搜索率对比

    Figure  12.  Comparison of duplicate rate between JDQN and JDQN-E algorithms

    图  13  IDQN与DQN-J算法的搜索覆盖率对比

    Figure  13.  Comparison of coverage rate between IDQN and DQN-J algorithms

    图  14  IDQN与DQN-J算法的重复搜索率对比

    Figure  14.  Comparison of duplicate rate between IDQN and DQN-J algorithms

    图  15  不同算法的冲突发生次数对比

    Figure  15.  Comparison of the number of conflicts between different algorithms

    图  16  不同算法的CDP曲线对比

    Figure  16.  Comparison of CDP curves with different algorithms

    表  1  实验参数列表

    Table  1.   List of experimental parameters

    参数
    搜索平台数量/台4
    POD0.8
    POC服从二维高斯分布
    障碍物密度5%
    α0.01
    γ0.99
    ε0.3
    总迭代次数/轮500
    最大探索步数/步20
    单次训练样本/个128
    经验回放池容量/个1000
    τ0.005
    下载: 导出CSV

    表  2  算法性能对比

    Table  2.   Comparison of algorithm performance

    性能评估指标JDQNIDQNQMIXVDN
    CDP1.01.00.99780.8937
    平均冲突次数4.36.44.32.7
    搜索覆盖率/%18.5717.2613.0811.13
    重复搜索率/%28.8833.5051.3859.38
    结束搜索平均步数14.814.418.219.3
    下载: 导出CSV
  • [1] 王中, 温志文, 蔡卫军. 水下无人航行器编队协同搜索研究综述[J]. 舰船科学技术, 2024, 46(2): 57-62. doi: 10.3404/j.issn.1672-7649.2024.02.010

    Wang Z, Wen Z W, Cai W J. Review of research on unmanned underwater vehicle formation cooperative search[J]. Ship Science and Technology, 2024, 46(2): 57-62. doi: 10.3404/j.issn.1672-7649.2024.02.010
    [2] 范学满, 薛昌友, 张会. 基于多种群遗传算法的多UUV任务分配方法[J]. 水下无人系统学报, 2022, 30(5): 621-630. doi: 10.11993/j.issn.2096-3920.202107001

    Fan X M, Xue C Y, Zhang H. Task assignment method for multiple UUVs based on multi-population genetic algorithm[J]. Journal of Unmanned Undersea Systems, 2022, 30(5): 621-630. doi: 10.11993/j.issn.2096-3920.202107001
    [3] Ni J, Yang L, Wu L, et al. An improved spinal neural system-based approach for heterogeneous auvs cooperative hunting[J]. International Journal of Fuzzy Systems, 2018, 20(2): 672-686. doi: 10.1007/s40815-017-0395-x
    [4] 曹迟, 史文涛, 王百合, 等. 无人水下航行器反潜作战模型仿真[J]. 水下无人系统学报, 2025, 33(1): 156-163. doi: 10.11993/j.issn.2096-3920.2024-0116

    Cao C, Shi W T, Wang B H, et al. Simulation of anti-submarine warfare model of unmanned undersea vehicles[J]. Journal of Unmanned Undersea Systems, 2025, 33(1): 156-163. doi: 10.11993/j.issn.2096-3920.2024-0116
    [5] 李小新, 陈华林. 水下机器人科学考察技术应用[J]. 船舶工程, 2021, 43(9): 8-13.

    Li X X, Chen H L. Application of underwater robot scientific investigation technology[J]. Ship Engineering, 2021, 43(9): 8-13.
    [6] 于传合, 高卫华, 李树奎, 等. 面向水下目标搜救的潜器探测模型构建与性能分析[J]. 控制与决策, 2025, 40(1): 170-179. doi: 10.13195/j.kzyjc.2024.0324

    Yu C H, Gao W H, Li S K, et al. Detection model construction and performance analysis for target search and rescue via autonomous underwater vehicles[J]. Control and Decision, 2025, 40(1): 170-179. doi: 10.13195/j.kzyjc.2024.0324
    [7] Koopman B O. The theory of search. II. Target detection[J]. Operations research, 1956, 4(5): 503-531. doi: 10.1287/opre.4.5.503
    [8] Kan Y C. Optimal search of a moving target[J]. Operations Research, 1977, 25(5): 864-870.
    [9] Hohzaki R, Iida K. Optimal ambushing search for a moving target[J]. European Journal of Operational Research, 2001, 133(1): 120-129. doi: 10.1016/S0377-2217(00)00187-9
    [10] 轩永波, 黄长强, 吴文超, 等. 运动目标的多无人机编队覆盖搜索决策[J]. 系统工程与电子技术, 2013, 35(3): 539-544. doi: 10.3969/j.issn.1001-506X.2013.03.15

    Xuan Y B, Huang C Q, Wu W C, et al. Coverage search strategies for moving targets using multiple unmanned aerial vehicle teams[J]. Systems Engineering and Electronics, 2013, 35(3): 539-544. doi: 10.3969/j.issn.1001-506X.2013.03.15
    [11] 曾国奇, 白宇, 林伟, 等. 地面运动目标的多UAV协同搜索方法[J]. 系统工程与电子技术, 2018, 40(7): 1498-1505.

    Zeng G Q, Bai Y, Lin W, et al. Multi-UAV cooperative search method for ground moving targets[J]. Systems Engineering and Electronics, 2018, 40(7): 1498-1505.
    [12] 岳伟, 席云, 关显赫. 基于多蚁群协同搜索算法的多AUV路径规划[J]. 水下无人系统学报, 2020, 28(5): 505-511.

    Yue W, Xi Y, Guan X H. Path planning of multi-AUVs based on multi-ant colony cooperative search algorithm[J]. Journal of Unmanned Undersea Systems, 2020, 28(5): 505-511.
    [13] Pham T H, Bestaoui Y, Mammar S. Aerial robot coverage path planning approach with concave obstacles in precision agriculture[C]//2017 Workshop on Research, Education and Development of Unmanned Aerial Systems (RED-UAS). IEEE, 2017: 43-48.
    [14] 丁文俊, 柴亚军, 杨宇贤, 等. 基于空海异构无人平台的水下目标搜索与跟踪[J]. 水下无人系统学报, 2024, 32(2): 237-249.

    Ding W J, Chai Y J, Yang Y X, et al. Underwater target search and tracking based on air-sea heterogeneous unmanned platform[J]. Journal of Unmanned Undersea Systems, 2024, 32(2): 237-249.
    [15] 刘和祥, 边信黔, 李娟, 等. 基于前视声纳信息的AUV局部路径规划研究[J]. 微计算机信息, 2007(23): 243-245.

    Liu H X, Bian X Q, Li J, et al. Study of local path planning based on forward looking sonar for AUV[J]. Microcomputer Information, 2007(23): 243-245.
    [16] Noguchi Y, Kuranaga Y, Maki T. Adaptive navigation of a high speed autonomous underwater vehicle using low-cost sensors for low-altitude survey[C]//2017 IEEE Underwater Technology(UT), IEEE, 2017: 1-4.
    [17] Yang Q Q, Gao Y Y, Guo Y, et al. Target search path planning for naval battle field based on deep reinforcement learning[J]. Systems Engineering and Electronics, 2022, 44(11): 3486-3495.
    [18] 杨鹏程, 杨清清, 高盈盈, 等. 基于强化学习的海上移动目标搜索路径规划[J]. 系统工程与电子技术, 2026, 48(2): 515-523.

    Yang P C, Yang Q Q, Gao Y Y, et al. Path planning for maritime moving target search based on reinforcement learning[J]. Systems Engineering and Electronics, 2026, 48(2): 515-523.
    [19] Li J , Jiang X , Zhang H , et al. Multi-joint adaptive control enhanced reinforcement learning for unmanned ship[J]. Ocean Engineering, 2025, 318: 120121.
    [20] Zhang C, Wu C, Chen H. Coverage path planning for maritime search based on reinforcement learning with path expansion strategy[C]//2024 7th International Conference on Robotics, Control and Automation Engineering(RCAE), IEEE, 2024: 531-536.
    [21] Sun Y, Zhang R, Liang W, et al. Multi-agent cooperative search based on reinforcement learning[C]//2020 3rd International Conference on Unmanned Systems(ICUS). IEEE, 2020: 891-896.
    [22] 邢博闻, 张昭夷, 王世明, 等. 基于深度强化学习的多无人艇协同目标搜索算法[J]. 兵器装备工程学报, 2023, 44(11): 118-125.

    Xing B W, Zhang Z Y, Wang S M, et al. Multi-USV cooperative target search algorithm based on deep reinforcement learning[J]. Journal of Ordnance Equipment Engineering, 2023, 44(11): 118-125.
    [23] Guo S Y, Zhang X G, Zheng Y S, et al. An Autonomous path planning model for unmanned ships based on deep reinforcement learning[J]. Sensors, 2020, 20(2): 426.
    [24] Loane E P, Richardson H R, Boylan E S. Theory of cumulative detection probability[R]. Washington DC: U. S. Navy Bureau of Ordnance, 1964.
    [25] Tan M. Multi-agent reinforcement learning: Independent vs. cooperative agents[J]. Machine Learning Proceedings, 1993: 330-337.
    [26] Rashid T, Samvelyan M, De Witt C S, et al. Monotonic value function factorisation for deep multi-agent reinforcement learning[J]. Journal of Machine Learning Research, 2020, 21(178): 1-51.
    [27] Sunehag P, Lever G, Gruslys A, et al. Value-decomposition networks for cooperative multi-agent learning based on team reward[C]//Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems. 2018: 2085-2087.
  • 加载中
计量
  • 文章访问数:  94
  • HTML全文浏览量:  41
  • PDF下载量:  17
  • 被引次数: 0
出版历程
  • 收稿日期:  2025-12-04
  • 修回日期:  2026-01-29
  • 录用日期:  2026-02-02
  • 网络出版日期:  2026-07-23
图(16) / 表(2)

目录

    /

    返回文章
    返回
    服务号
    订阅号