Federated Deep Reinforcement Learning-Based Energy Efficiency Optimization for Closed-Loop Operations and Maintenance in Autonomous Communication Networks
Haitao Li, Donglei Xu, Quanfeng Yao, Lei Zhang, Hui Wan, Xianjun Peng, Ruchun Wang, Ming Xue
With the evolution toward 5G-Advanced and 6G autonomous communication networks, achieving energy-efficient closed-loop operations and maintenance (O&M) has become increasingly challenging due to large-scale deployment, heterogeneous network elements, and highly dynamic traffic conditions. To overcome the limitations of centralized optimization and standalone deep reinforcement learning (DRL) methods in terms of scalability, privacy protection, and adaptability, a federated deep reinforcement learning (FDRL)–based optimization framework is developed. The closed-loop O&M energy optimization problem is formulated as a decentralized partially observable Markov decision process, in which distributed edge nodes operate as local DRL agents that perceive real-time network states and perform sequential O&M decision-making. Each agent adopts a segmented actor–critic architecture to jointly handle mixed discrete–continuous action spaces, enabling coordinated optimization of key O&M controls such as resource activation, transmission power adjustment, and sleep scheduling under quality-of-service constraints. A federated learning mechanism is integrated to support privacy-preserving collaboration, where local model updates are periodically aggregated at a federation server through a heterogeneity-aware adaptive weighting strategy to mitigate the impact of non-independent and identically distributed data across nodes. By embedding federated training into a perception–decision–execution–feedback closed-loop O&M architecture, the proposed approach enables continual online adaptation to non-stationary network environments. Experimental results based on realistic autonomous network scenarios demonstrate significant improvements in energy efficiency, convergence stability, and scalability compared with centralized DRL and heuristic methods, validating the effectiveness of the proposed framework for next-generation autonomous communication networks.