Safety–Efficiency Balanced Navigation for Unmanned Tracked Vehicles in Uneven Terrain Using Prior-Based Ensemble Deep Reinforcement Learning
Yiming Xu, Songhai Zhu, Dianhao Zhang, Yinda Fang, Mien Van
This paper proposes a novel navigation approach for Unmanned Tracked Vehicles (UTVs) using prior-based ensemble deep reinforcement learning, which fuses the policy of the ensemble Deep Reinforcement Learning (DRL) and Dynamic Window Approach (DWA) to enhance both exploration efficiency and deployment safety in unstructured off-road environments. First, by integrating kinematic analysis, we introduce a novel state and an action space that account for rugged terrain features and track–ground interactions. Local elevation information and vehicle pose changes over consecutive time steps are used as inputs to the DRL model, enabling the UTVs to implicitly learn policies for safe navigation in complex terrains while minimizing the impact of slipping disturbances. Then, we introduce an ensemble Soft Actor–Critic (SAC) learning framework, which introduces the DWA as a behavioral prior, referred to as the SAC-based Hybrid Policy (SAC-HP). Ensemble SAC uses multiple policy networks to effectively reduce the variance of DRL outputs. We combine the DRL actions with the DWA method by reconstructing the hybrid Gaussian distribution of both. Experimental results indicate that the proposed SAC-HP converges faster than traditional SAC models, which enables efficient large-scale navigation tasks. Additionally, a penalty term in the reward function about energy optimization is proposed to reduce velocity oscillations, ensuring fast convergence and smooth robot movement. Scenarios with obstacles and rugged terrain have been considered to prove the SAC-HP’s efficiency, robustness, and smoothness when compared with the state of the art.