• 中国科技核心期刊
  • Scopus收录期刊
  • DOAJ收录期刊
  • JST收录期刊
  • Euro Pub收录期刊
Turn off MathJax
Article Contents
PANG Zhouqi, LIN Xiaobo, GENG Shicheng, HAO Chengpeng. AUV traversal planning method based on AMDQN[J]. Journal of Unmanned Undersea Systems. doi: 10.11993/j.issn.2096-3920.2026-0028
Citation: PANG Zhouqi, LIN Xiaobo, GENG Shicheng, HAO Chengpeng. AUV traversal planning method based on AMDQN[J]. Journal of Unmanned Undersea Systems. doi: 10.11993/j.issn.2096-3920.2026-0028

AUV traversal planning method based on AMDQN

doi: 10.11993/j.issn.2096-3920.2026-0028
  • Received Date: 2026-01-26
  • Accepted Date: 2026-02-26
  • Rev Recd Date: 2026-02-13
  • Available Online: 2026-07-23
  • To improve the traversal path planning capability of autonomous underwater vehicle(AUV), this paper proposes an AUV traversal path planning algorithm based on advanced M-DQN (AMDQN). First, this paper establishes a traversal planning environment suitable for practical scenarios. On the one hand, the ray coverage method is used instead of the traditional rectangular grid modeling method to improve modeling accuracy; on the other hand, the optional action set of the AUV is constrained to limit its maneuverability. Second, each component of the algorithm is carefully designed under the above environment. Specifically, in the state space design, this paper fuses vague global environmental information, accurate local environmental information centered on the AUV, and the AUV’s own position information to systematically represent the environment. In the reward function design, “rule penalty” and “edge guidance” are introduced, enabling the AUV to stably improve coverage along environmental edges. For the parameter update method, an adaptive temperature parameter update strategy is designed based on M-DQN, while the multi-step reward and a “dueling” network architecture are introduced to alleviate training variance. During the training process, “soft reset” of fully connected layer parameters is adopted to mitigate the local optimum problem. Finally, simulation results show that the proposed algorithm has stronger environmental adaptability and higher traversal coverage compared with traditional methods and unimproved reinforcement learning algorithms.

     

  • loading
  • [1]
    苗润龙, 庞硕, 姜大鹏, 等. 海洋自主航行器多海湾区域完全遍历路径规划[J]. 测绘学报, 2019, 48(2): 256-264.

    Miao R L, Pang S, Jiang D P, et al. Complete coverage path planning for autonomous marine vehicle used in multi-bay areas[J]. Acta Geodaetica et Cartographica Sinica, 2019, 48(2): 256-264.
    [2]
    Tang G, Tang C, Zhou H, et al. R-DFS: A coverage path planning approach based on region optimal decomposition[J]. Remote Sensing, 2021, 13(8): 1525.
    [3]
    Xu A, Viriyasuthee C, Rekleitis I. Optimal complete terrain coverage using an unmanned aerial vehicle[C]//2011 IEEE International Conference on Robotics and Automation. IEEE, 2011: 2513-2519.
    [4]
    熊亿民. 基于改进蚁群算法的全向移动机器人全遍历路径规划[J]. 计算机系统应用, 2021, 30(6): 209-214.

    Xiong Y M. Full Traversal Path Planning of Omnidirectional Mobile Robot Based on Improved Ant Colony Algorithm[J]. Computer Systems and Applications, 2021, 30(6): 209-214.
    [5]
    池志猛, 李智刚, 赵洋, 等. 针对深海采矿车的遍历路径规划方法研究[J]. 矿业研究与开发, 2022, 42(7): 160-166.

    Chi Z M, Li Z G, Zhao Y, et al. Research on traversal path planning method for deep-sea mining vehicles[J]. Mining Research and Development, 2022, 42(7): 160-166.
    [6]
    Bouman A, et al. Adaptive coverage path planning for efficient exploration of unknown environments[C]//2022 IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS). IEEE, 2022: 11916-11923.
    [7]
    温志文, 杨春武, 蔡卫军, 等. 复杂环境下UUV完全遍历路径规划方法[J]. 鱼雷技术, 2017, 25(1): 22-26.

    Wen Z W, Yang C W, Cai W J, et al. A complete coverage path planning method of UUV under complex environment[J]. Journal of Unmanned Undersea Systems, 2017, 25(1): 22-26.
    [8]
    史方青, 黄华, 张昊, 等. 玻璃幕墙清洗机器人内螺旋完全遍历路径规划研究[J]. 哈尔滨工程大学学报, 2024, 45(6): 1170-1178.

    Shi F Q, Huang H, Zhang H, et al. Study on the path planning of the internal spiral complete traversal for glass-curtain wall-cleaning robots[J]. Journal of Harbin Engineering University, 2024, 45(6): 1170-1178.
    [9]
    吕霞付, 程启忠, 李森浩, 等. 基于改进A*算法的无人船完全遍历路径规划[J]. 水下无人系统学报, 2019, 27(6): 695-703.

    LÜ X F, Cheng Q Z, Li S H, et al. Unmanned surface vehicle full traversal path planning based on improved A* algorithm[J]. Journal of Unmanned Undersea Systems, 2019, 27(6): 695-703.
    [10]
    时天龙. 基于A*算法的无人船完全遍历路径规划研究[D]. 大连: 大连海事大学, 2024: 32-38.
    [11]
    王岩, 李春书, 张福龙. 一种改进的爬壁机器人全遍历路径规划[J]. 计算机仿真, 2023, 40(7): 447-452.

    Wang Y, Li C S, Zhang F L. An improved full traversal path planning for wall climbing robot[J]. Computer Simulation, 2023, 40(7): 447-452.
    [12]
    孟浩德, 吴征天, 吴闻笛, 等. 基于记忆模拟退火算法的扫地机器人遍历路径规划[J]. 计算机与数字工程, 2024, 52(3): 821-826, 857.

    Meng H D, Wu Z T, Wu W D et al. Traversal path planning of sweeping robot based on memory simulated annealing algorithm[J]. Computer and Digital Engineering. 2024, 52(3): 821-826, 857.
    [13]
    Lin B, Han G H, Song C C, et al. Traversal path planning and simulation of robot based on radiation scanning[J]. Journal of System Simulation, 2021, 33(1): 83-90.
    [14]
    Carvalho J P, Aguiar A P. Deep reinforcement learning for zero-shot coverage path planning with mobile robots[J]. IEEE/CAA Journal of Automatica Sinica, 2025, 12(8): 1594-1609.
    [15]
    Mnih V, Kavukcuoglu K, Silver D, et al. Human-level control through deep reinforcement learning[J]. Nature, 2015, 518(7540): 529-533.
    [16]
    Chen Y, Lu Z M, Cui J L, et al. A complete coverage path planning algorithm for lawn mowing robots based on deep reinforcement learning[J]. Sensors, 2025, 25(2): 416.
    [17]
    Xing B, Wang X, Yang L, et al. An algorithm of complete coverage path planning for unmanned surface vehicle based on reinforcement learning[J]. Journal of Marine Science and Engineering, 2023, 11(3): 645.
    [18]
    Ai B, Jia M, Xu H, et al. Coverage path planning for maritime search and rescue using reinforcement learning[J]. Ocean Engineering, 2021, 241: 110098.
    [19]
    Heydari J, Saha O, Ganapathy V. Reinforcement learning-based coverage path planning with implicit cellular decomposition[EB/OL]. (2021-10-18)[2026-07-21]. https://doi.org/10.48550/arXiv.2110.09018.
    [20]
    Van H H, Guez A, Silver D. Deep reinforcement learning with double Q-learning[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2016, 30(1): 2094-2100.
    [21]
    Theile M, Bayerlein H, Nai R, et al. UAV coverage path planning under varying power constraints using deep reinforcement learning[C]//2020 IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS). IEEE, 2020: 1444-1449.
    [22]
    Zheng T, Jin Y, Zhao H, et al. Deep reinforcement learning based coverage path planning in unknown environments[C]//2024 6th International Conference on Frontier Technologies of Information and Computer (ICFTIC). IEEE, 2024: 1608-1611.
    [23]
    Carvalho J P, Aguiar A P. A reinforcement learning based online coverage path planning algorithm[C]//2023 IEEE International Conference on Autonomous Robot Systems and Competitions(ICARSC). IEEE, 2023: 81-86.
    [24]
    Haarnoja T, Zhou A, Abbeel P, et al. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor[C]//International Conference on Machine Learning. PMLR, 2018: 1861-1870.
    [25]
    Jonnarth A, Zhao J, Felsberg M. Learning coverage paths in unknown environments with deep reinforcement learning[EB/OL]. (2024-06-07)[2026-07-21]. https://doi.org/10.48550/arXiv.2306.16978.
    [26]
    Schulman J, Wolski F, Dhariwal P, et al. Proximal policy optimization algorithms[EB/OL]. (2017-07-20)[2026-07-21]. https://doi.org/10.48550/arXiv.1707.06347.
    [27]
    Wijegunawardana I D, Samarakoon S M B P, Muthugala M A V J, et al. Risk-aware complete coverage path planning using reinforcement learning[J]. IEEE Transactions on Systems, Man, and Cybernetics: Systems, 2025, 55(4): 2476-2488.
    [28]
    Vieillard N, Pietquin O, Geist M. Munchausen reinforcement learning[C]//Advances in Neural Information Processing Systems. 2020, 33: 4235-4246.
    [29]
    Von B M. An open-source benchmark simulator: control of a BlueROV2 underwater robot[J]. Journal of Marine Science and Engineering, 2022, 12(10): 1898.
  • 加载中

Catalog

    通讯作者: 陈斌, bchen63@163.com
    • 1. 

      沈阳化工大学材料科学与工程学院 沈阳 110142

    1. 本站搜索
    2. 百度学术搜索
    3. 万方数据库搜索
    4. CNKI搜索

    Figures(17)  / Tables(10)

    Article Metrics

    Article Views(81) PDF Downloads(15) Cited by()
    Proportional views
    Related
    Service
    Subscribe

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return