Transferable Deep Reinforcement Learning With Edge‐Contour‐Depth Fusion for Autonomous Wireless Capsule Endoscopy Navigation
Haoxuan Wu, Haitao Gao, Qingyang Liu, Sishen Yuan, Haiyang Fang, Mingwu Su, Baijia Liang, Yongzun Yang, Long Bai, Wenzhen Dong, Dihong Xie, Shijian Su, Jiewen Lai, Shing Shin Cheng, Zhen Li, Xiuli Zuo, Hongliang Ren
ABSTRACT Wireless capsule endoscopy (WCE) enables painless, minimally invasive visualization of the gastrointestinal tract. Still, its diagnostic potential is limited by incomplete mucosal coverage and poor transferability of existing navigation methods across patient anatomies. We propose a transferable, anatomical landmark‐guided deep reinforcement learning framework for robust autonomous gastric navigation. Leveraging a lightweight edge‐contour‐depth fusion module, our policy operates on stable, low‐dimensional landmark coordinates rather than high‐dimensional video streams. This design effectively bridges the sim‐to‐real visual gap and ensures robustness across diverse anatomies, enabling low‐cost deployment by reducing computational overhead. In simulations across eight patient‐derived models, the method achieves >97% coverage within 50 s, significantly outperforming vanilla Proximal Policy Optimization, Soft Actor‐Critic, and Deep Q‐Network agents by enhancing coverage and minimizing variance. To ensure deployment reliability, a two‐stage sim‐to‐real pipeline supported by an adaptive dynamic programming controller actively mitigates physical disturbances, including actuator latency and peristalsis. Ex vivo experiments across five independent scans demonstrate high coverage stability, achieving a mean coverage of 87% and a 53% reduction in procedure time compared with expert manual control. This study establishes a scalable paradigm for autonomous, high‑coverage endoscopic navigation, advancing the clinical deployment of intelligent WCE systems for GI diagnostics.