Articles | Volume 11, issue 9
https://doi.org/10.5194/wes-11-3377-2026
© Author(s) 2026. This work is distributed under the Creative Commons Attribution 4.0 License.
Classification of leading-edge-erosion severity via machine learning surrogate models
Download
- Final revised paper (published on 10 Sep 2026)
- Preprint (discussion started on 28 Jan 2026)
Interactive discussion
Status: closed
Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor
| : Report abuse
- RC1: 'Comment on wes-2025-289', Anonymous Referee #1, 19 Mar 2026
- RC2: 'Comment on wes-2025-289', Anonymous Referee #2, 23 May 2026
- AC1: 'Comment on wes-2025-289', Aidan Gettemy, 30 Jun 2026
Peer review completion
AR – Author's response | RR – Referee report | ED – Editor decision | EF – Editorial file upload
AR by Aidan Gettemy on behalf of the Authors (30 Jun 2026)
Author's response
Author's tracked changes
Manuscript
ED: Referee Nomination & Report Request started (07 Jul 2026) by Julie Teuwen
RR by Anonymous Referee #1 (08 Jul 2026)
RR by Anonymous Referee #2 (14 Jul 2026)
ED: Publish as is (16 Jul 2026) by Julie Teuwen
ED: Publish as is (02 Aug 2026) by Athanasios Kolios (Chief editor)
AR by Aidan Gettemy on behalf of the Authors (12 Aug 2026)
Author's response
Manuscript
The manuscript addresses an important problem in wind turbine monitoring by exploring the use of surrogate models to generate training data for erosion classification. The approach is interesting and the paper is generally well written, but several aspects of the data generation, methodology, and positioning of the contribution could be strengthened to better reflect the complexity of the real monitoring problem and to clarify the novelty of the proposed framework.
Major Points
- The overall contribution is not yet clearly positioned with respect to the existing literature on surrogate modeling and wind turbine condition monitoring. Gaussian-process surrogates, sensitivity analysis, and random forest classifiers are all well-established techniques. The manuscript should clarify more explicitly what methodological advance is introduced beyond applying these tools to a specific erosion-monitoring scenario.
- The claim of novelty regarding the PPzGP surrogate is not sufficiently demonstrated. While combining parallel partial emulation and range-censored Gaussian processes is technically interesting, the manuscript does not clearly show why this combination enables capabilities that standard GP surrogates would not provide for this problem.
- The erosion model used to generate the data is highly simplified. Blade erosion is represented through a parametric scaling of lift and drag coefficients across six blade regions. While this may be suitable for a proof-of-concept study, the manuscript should discuss the limitations of this representation and justify why it captures the key aerodynamic effects of real leading-edge erosion.
- The erosion process is modeled as discrete severity classes rather than a continuous degradation process. In reality erosion evolves gradually and spatially across the blade surface. The use of five artificial classes may simplify the classification task and should be justified more clearly.
- The classification problem may be artificially easy because the erosion perturbations are directly embedded in the aerodynamic coefficients and the classifier is trained on outputs that are strongly linked to those coefficients (e.g., lift and drag sensor statistics). This raises the possibility that the model is learning the synthetic perturbation rather than identifying erosion signatures that would be observable in practice.
- The strong performance of a relatively simple random forest classifier suggests that the generated dataset may be too clean or too easily separable. In practice, leading-edge erosion detection is known to be challenging due to turbulence, operational variability, sensor noise, and confounding effects. The simulations appear to lack these disturbances, which may make the classification task unrealistically simple.
- The simulations assume uniform wind conditions rather than turbulent inflow. Since turbulence strongly influences turbine loads and vibration signals, the use of uniform wind fields likely underestimates the variability present in real monitoring data. Including turbulent wind realizations would significantly improve realism.
- The operational variability of the turbine is limited. Real turbines experience controller transitions, yaw adjustments, and varying operating regimes that influence measured signals. The current simulation setup may not capture these effects.
- The sensor configuration used in the study may not reflect practical monitoring systems. In particular, lift and drag pressure sensors are rarely available in operational wind turbines. The manuscript should discuss the feasibility of the assumed sensing setup or consider signals more commonly available in SCADA or structural monitoring systems.
- The feature extraction strategy reduces time-series signals to simple statistical moments (mean, standard deviation, skewness, kurtosis). This discards potentially important information contained in the temporal and spectral structure of the signals. The authors should justify this choice or explore richer feature representations.
- The surrogate model is trained on a relatively small number of simulations compared to the dimensionality and nonlinearity of the system. Although the reported prediction errors are moderate, the manuscript should discuss potential model bias and the limits of extrapolation.
- The reported classification improvement between simulator-trained and surrogate-trained models (83% vs 87%) is relatively modest. It would be helpful to provide statistical analysis or repeated experiments to assess whether this improvement is significant.
Minor Points
- The manuscript frequently refers to digital twins, but the work primarily demonstrates surrogate modeling and classification using simulated data. Since essential digital twin elements such as data assimilation, state estimation, or online updating are not included, the connection to digital twins should be described more cautiously.
- The introduction is somewhat lengthy and could be shortened. Several sections summarizing wind-energy background material could be condensed to focus more directly on the methodological contribution.
- Some terminology is used interchangeably throughout the manuscript (e.g., emulator, surrogate model). Consistent terminology would improve clarity.
- Figures illustrating emulator predictions are informative but could be improved in readability, particularly through larger axis labels and clearer legends.
- The manuscript would benefit from a clearer discussion of the gap between simulation-based validation and deployment on real turbine monitoring data.
- Minor typographical issues and formatting inconsistencies appear throughout the manuscript and should be corrected during revision.