the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Real-Time IoT, LLM, AI-Supported Wind Turbine Failure Prediction System
Abstract. Wind and solar energy are two popular alternative energy sources. However, wind turbines are larger and more complex than solar panels. Accordingly, wind turbines are more exposed to environmental factors and therefore more prone to mechanical failures. Our work improves the reliability and efficiency of wind energy systems by presenting an artificial intelligence (AI)-based system that predicts mechanical gearbox and electrical failures. Historical data are aggregated with real-time sensor data to train the prediction model. Apart from related mechanical and environmental data, sensors provide real-time vibration, internal nacelle temperature, wind speed, noise and smoke levels. The developed system integrates AI and Large Language Model (LLM)-based interfaces for real-time interactive monitoring of turbines. The user interface of the developed system allows users to receive informative responses on performance, detected risks, predicted failures, and energy production levels. The developed model has been validated using 5-fold cross-validation based on Accuracy, Precision, Recall, F1-Score, and ROC-AUC. The model achieves approximately 89.68 % Accuracy, 90.08 % F1- Score, 95.65 % ROC-AUC, and novel metric 65.13 %, Overall Performance. The performance results demonstrate the promising potential of AI- and LLM-integrated systems for wind energy applications. Prototype data, labeled via the XGBoost model trained with SCADA data, was retrained using the LightGBM algorithm, achieving 98.37 % Accuracy, 99.16 % F1-Score, 98.95 % ROC-AUC and 94.65 % Overall Performance; the analysis proved that the newly added gas and sound sensors significantly improved the fault prediction performance of the system.
- Preprint
(1470 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
-
RC1: 'Comment on wes-2026-95', Anonymous Referee #1, 12 Aug 2026
The comment was uploaded in the form of a supplement: https://wes.copernicus.org/preprints/wes-2026-95/wes-2026-95-RC1-supplement.pdfCitation: https://doi.org/
10.5194/wes-2026-95-RC1 -
AC1: 'Reply on RC1', Sara Mahyanbakhshayesh, 29 Aug 2026
We would like to sincerely thank you for your constructive and insightful feedback. Although our revision process took some time, we have dedicated significant effort to carefully addressing each of your points, which we believe has greatly improved the quality of our manuscript.
Please find our detailed, point-by-point response in the attached supplementary file.
Additionally, since we are addressing your valuable comments first, please kindly note that the line numbers referenced in our response correspond to our current working draft. These locations may shift slightly as we continue to incorporate feedback from the other reviewers.
Thank you again for your time and contribution to our work.-
RC2: 'Reply on AC1', Anonymous Referee #1, 29 Aug 2026
I have not accessed the revised manuscript yet, but I have gone through the response document, and most of the points seem to have been addressed carefully. I nevertheless remain concerned about the rationale and evaluation of the proposed teacher–student procedure.
Very high student accuracy is not surprising when the student is trained using labels generated by the teacher, because the student is effectively learning to reproduce the teacher's decisions. The resulting error therefore mixes the teacher's labeling error on the prototype, the change from XGBoost to LightGBM, and the contribution of the additional sensors. It does not clearly demonstrate that vibration, gas, or acoustic measurements improve actual fault detection.
A more convincing design would reserve a subset of independently verified prototype labels completely unseen by both models and use this subset to evaluate both the transferred teacher and the student against the same true labels. If the student performs better, the improvement could provide evidence that the additional prototype-specific features contribute useful information; if it performs worse, this would suggest that they do not provide a reliable benefit.
Please clarify whether independently verified prototype labels exist for any portion of the data. If they do, why are they not used for independent evaluation? If they do not, then I would also like clarification on how the ground-truth labels used to train the original SCADA teacher were obtained, and how the authors establish that the teacher-generated prototype labels are sufficiently reliable to serve as targets for the student.
Citation: https://doi.org/10.5194/wes-2026-95-RC2 -
AC2: 'Reply on RC2', Sara Mahyanbakhshayesh, 05 Sep 2026
Thank you very much for your time, constructive comments, and valuable suggestions.
We are grateful for your detailed feedback, which has significantly helped us improve the quality of our manuscript.
To ensure clarity and readability, we have prepared our point-by-point responses in a separate document. For each comment, we have indicated the exact location in the revised manuscript by line number, and all changes have been clearly highlighted in yellow in the revised manuscript file.
We hope that the revised version now fully addresses all the points raised.
Thank you once again for your consideration. We look forward to hearing from you.
Sincerely,
Sara Mahyanbakhshayesh
-
AC2: 'Reply on RC2', Sara Mahyanbakhshayesh, 05 Sep 2026
-
RC2: 'Reply on AC1', Anonymous Referee #1, 29 Aug 2026
-
AC1: 'Reply on RC1', Sara Mahyanbakhshayesh, 29 Aug 2026
-
RC3: 'Comment on wes-2026-95', Anonymous Referee #2, 31 Aug 2026
(1) The novelty of the predictive modeling component is not sufficiently clear. The manuscript mainly evaluates conventional machine learning models, with the best performance obtained from tree-based methods such as XGBoost and LightGBM. The authors should clarify why more advanced models for multivariate time-series analysis are not considered and better distinguish the methodological novelty from system-level integration.
(2) The role of the LLM and RAG components requires much clearer explanation. It is unclear what knowledge base is used for RAG, how retrieval is performed, and what retrieved information is provided to the LLM. Based on the current description, the LLM appears mainly to serve as a natural-language interface for presenting sensor and prediction results rather than providing a new diagnostic or reasoning capability.
(3) The manuscript lacks a clear unified framework or end-to-end workflow that connects the individual components. Although Figures 4 and 5 show the general system architecture and communication flow, the overall research methodology remains fragmented across SCADA preprocessing, model training, prototype sensing, pseudo-label generation, real-time inference, explainability, and the LLM interface. The manuscript would benefit from a unified methodological figure showing the complete workflow.
(4) The literature review is relatively limited and does not sufficiently establish the research gap. Rather than mainly listing individual studies and their performance, the authors should provide a more systematic overview of SCADA-based prediction, multimodal sensing, IoT-based monitoring, explainable AI, and recent LLM/RAG applications in wind-turbine maintenance. A comparison table could also help distinguish the proposed work from existing studies.
(5) The claim that the newly added gas and sound sensors significantly improve fault prediction is not sufficiently supported by the current experiments. Feature importance and SHAP results indicate that the model uses these features, but they do not demonstrate that adding these sensors improves predictive performance. The authors should conduct an ablation study comparing model performance with and without the newly added sensors to directly quantify their contribution.
(6) The current manuscript reads more like a technical project report than a research paper.
Citation: https://doi.org/10.5194/wes-2026-95-RC3 -
AC3: 'Reply on RC3', Sara Mahyanbakhshayesh, 05 Sep 2026
We sincerely appreciate the time and effort you have devoted to reviewing our manuscript, as well as your insightful comments and constructive suggestions.
Your detailed feedback has been extremely valuable and has contributed greatly to improving the clarity, quality, and overall presentation of our manuscript.
For the sake of clarity, we have provided a detailed, point-by-point response to each of the comments in a separate document. The corresponding revisions have been identified by line numbers, and all modifications made in the revised manuscript have been highlighted in yellow for ease of reference.
We believe that we have carefully addressed all of the issues and suggestions raised during the review process and hope that the revised manuscript meets your expectations.
Thank you once again for your valuable contribution and careful consideration of our work. We greatly appreciate your time and look forward to your further evaluation.
Sincerely,
Sara Mahyanbakhshayesh
-
AC3: 'Reply on RC3', Sara Mahyanbakhshayesh, 05 Sep 2026
Data sets
Wind Turbine SCADA Data For Early Fault Detection Azizjon Kasimov https://www.kaggle.com/datasets/azizkasimov/wind-turbine-scada-data-for-early-fault-detection
Model code and software
SWTFPS_Wind-Turbine-Failure-Prediction Sara Mahyanbakhshayesh, İlhan Gamze Saygı, and Berdan Ak https://github.com/SaraMahyan/SWTFPS_Wind-Turbine-Failure-Prediction
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 280 | 109 | 67 | 456 | 75 | 67 |
- HTML: 280
- PDF: 109
- XML: 67
- Total: 456
- BibTeX: 75
- EndNote: 67
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1