Preprints
https://doi.org/10.5194/wes-2026-145
https://doi.org/10.5194/wes-2026-145
31 Aug 2026
 | 31 Aug 2026
Status: this preprint is currently under review for the journal WES.

Chronologically calibrated PCA for wind-turbine SCADA health monitoring: a leakage-controlled CARE v6 benchmark

Jui-Hung Liu, Jien-Chen Chen, and Chun-Chieh Wang

Abstract. Reliable wind-turbine prognostics and health management requires anomaly detectors that limit false-alarm burden without concealing missed events. This study evaluates chronological calibration of principal component analysis (PCA) reconstruction error on CARE to Compare v6, which contains 5 242 948 real ten-minute supervisory control and data acquisition (SCADA) records from 36 turbines in three wind farms. Turbine identities were partitioned before evaluation into 23 development turbines (58 cases) and 13 locked-test turbines (37 cases). Within each case, PCA was fitted on an earlier normal block, and a disjoint chronological block calibrated the alert threshold. Development-only selection chose global conformal PCA with α = 0.005. On the locked test set it achieved a CARE-style score of 0.506 (farm–turbine cluster-bootstrap 95 % confidence interval 0.378–0.632), compared with 0.252 for a fixed-contamination isolation forest comparator; the paired difference was 0.249 (95 % confidence interval 0.026–0.466). Mean normal-case accuracy increased from 0.252 to 0.865, and false-positive events decreased from 17 to 4. However, only 5 of 18 anomaly events were detected, including none of the five events in Farm A. A development-tuned isolation forest scored 0.542 and a conformalized isolation forest scored 0.502; neither paired comparison distinguished the selected PCA method. Ordinary empirical calibration generated exactly the same locked-test alerts as the finite-sample conformal rank rule. The evidence supports chronological, identity-disjoint calibration as a transparent reference for SCADA health monitoring, but it does not support universal superiority of PCA or of its conformal rank correction. We report the negative and equivalence findings in full so that the benchmark can be reused and contested.

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Share
Jui-Hung Liu, Jien-Chen Chen, and Chun-Chieh Wang

Status: open (until 28 Sep 2026)

Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor | : Report abuse
Jui-Hung Liu, Jien-Chen Chen, and Chun-Chieh Wang
Jui-Hung Liu, Jien-Chen Chen, and Chun-Chieh Wang
Metrics will be available soon.
Latest update: 01 Sep 2026
Download
Short summary
Wind turbines log sensor data every ten minutes, and software is built to spot faults early. We asked how well such software works when tested fairly, on turbines it has never seen. Using public data from 36 turbines, we fixed every choice in advance, then looked once at the held-out turbines. Our method raised far fewer false alarms than the standard comparison, but missed most faults, and a physics-based variant found none. We publish the full protocol and code so others can check it.
Share
Altmetrics