Reinforcement-Learning–Based Collaborative Optimization of Load Regulation and Opportunistic Maintenance of Wind Turbines
Abstract. To address the difficulty in coordinating component-degradation regulation strategies with opportunistic maintenance timings over the full lifecycle of wind turbines, this study proposes a reinforcement-learning–based collaborative optimization method for wind-turbine load regulation and opportunistic maintenance (RL-OppOM).The proposed method establishes a full-lifecycle simulation environment for wind turbines by coupling wind conditions, power-load characteristics, component reliability, and maintenance restoration, and formulates graded power derating, multicomponent maintenance combinations, and operational feasibility constraints within a unified Markov decision process. A lifecycle risk-aware reward function is developed and a factorized dual-clip masked proximal policy optimization algorithm is proposed. By incorporating policy factorization, action masking, and dual clipping, the algorithm reduces the complexity of policy learning and improves training stability. Simulation validation is conducted using SCADA data from a wind farm in northern China. The results show that RL-OppOM achieves a comprehensive cost of 7.320 M CNY, which is 13.80 % and 11.65 % lower than the costs of CBM and OppM, respectively, while increasing the net profit by 2.36 % and 1.93 %. The mean and minimum turbine-health indicators are 0.6413 and 0.4491, respectively, and the number of high-risk operating days is zero.