<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" specific-use="SMUR" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">WESD</journal-id>
<journal-title-group>
<journal-title>Wind Energy Science Discussions</journal-title>
<abbrev-journal-title abbrev-type="publisher">WESD</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Wind Energ. Sci. Discuss.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">2366-7621</issn>
<publisher><publisher-name></publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/wes-2026-157</article-id>
<title-group>
<article-title>Reinforcement-Learning&amp;ndash;Based Collaborative Optimization of Load Regulation and Opportunistic Maintenance of Wind Turbines</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Lin</surname>
<given-names>Jie</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Shi</surname>
<given-names>Jing</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Zhu</surname>
<given-names>Jianghao</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Li</surname>
<given-names>Xubin</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Chen</surname>
<given-names>Wei</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>School of Electrical Engineering and Information Engineering, Lanzhou University of Technology. Lanzhou, 730050. China</addr-line>
</aff>
<pub-date pub-type="epub">
<day>10</day>
<month>09</month>
<year>2026</year>
</pub-date>
<volume>2026</volume>
<fpage>1</fpage>
<lpage>32</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Jie Lin et al.</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://wes.copernicus.org/preprints/wes-2026-157/">This article is available from https://wes.copernicus.org/preprints/wes-2026-157/</self-uri>
<self-uri xlink:href="https://wes.copernicus.org/preprints/wes-2026-157/wes-2026-157.pdf">The full text article is available as a PDF file from https://wes.copernicus.org/preprints/wes-2026-157/wes-2026-157.pdf</self-uri>
<abstract>
<p>To address the difficulty in coordinating component-degradation regulation strategies with opportunistic maintenance timings over the full lifecycle of wind turbines, this study proposes a reinforcement-learning&amp;ndash;based collaborative optimization method for wind-turbine load regulation and opportunistic maintenance (RL-OppOM).The proposed method establishes a full-lifecycle simulation environment for wind turbines by coupling wind conditions, power-load characteristics, component reliability, and maintenance restoration, and formulates graded power derating, multicomponent maintenance combinations, and operational feasibility constraints within a unified Markov decision process. A lifecycle risk-aware reward function is developed and a factorized dual-clip masked proximal policy optimization algorithm is proposed. By incorporating policy factorization, action masking, and dual clipping, the algorithm reduces the complexity of policy learning and improves training stability. Simulation validation is conducted using SCADA data from a wind farm in northern China. The results show that RL-OppOM achieves a comprehensive cost of 7.320 M CNY, which is 13.80 % and 11.65 % lower than the costs of CBM and OppM, respectively, while increasing the net profit by 2.36 % and 1.93 %. The mean and minimum turbine-health indicators are 0.6413 and 0.4491, respectively, and the number of high-risk operating days is zero.</p>
</abstract>
<counts><page-count count="32"/></counts>
<funding-group>
<award-group id="gs1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>52667011</award-id>
<award-id>62666044</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body/>
<back>
</back>
</article>