Transformer-Based Probabilistic Wind Power Prediction and Power Ramping Verification with Error Tolerance
Abstract. Accurate probabilistic wind power forecasts and reliable detection of ramping events are important for the stable operation of wind-integrated power systems. This study implements a Self-Attentive Ensemble Transformer for postprocessing operational wind power ensembles in the Belgian Offshore Zone. This postprocessing delivers corrected ensemble members as output rather than a predictive distribution. The Transformer model achieves near-zero bias in the ensemble mean and a lower Continuous Ranked Probability Score across all lead times than power-curve derived raw ensembles and shows some improvements to Member-by-Member benchmark methods. This model also shows the ability to correct for power overestimation associated with the wake effect. For power ramping, this paper focuses on the predictability of 15 % and 30 % hourly ramping events. The Transformer produces event frequencies that closely align with the observations over all ramping thresholds. Considering that frequent temporal shifts and intensity underestimations limit a fair evaluation of ramping prediction, we introduce an error-tolerant probabilistic verification framework with buffer concepts and score the model by Buffer Brier Skill Score (BBSS). Incorporating a temporal buffer and magnitude tolerance substantially mitigates penalties for minor prediction errors. This reveals that the models are capable of detecting ramping signals, even though they lack strict precision. While uncertainty is inherent in ramping forecasts, the Transformer demonstrates better-calibrated probabilities than the raw ensembles. Overall, the Transformer method improves probabilistic power prediction and provides an informative probabilistic representation of ramping events.
The manuscript introduces some interesting contributions with a transformer-based approach for probabilistic wind power forecast from weather ensemble forecasts as well as introducing a new verification method for ramp events. The motivations for such methods are clearly described and argumented. It is some valuable work and it is believed that some of the following minor comments would help strengthen the presented results and the choice of some assumptions.
Line 8, in the abstract, it would be great to also quantify how "close" "...the transformer produces event frequencies that closely align with the observations over all ramping thresholds..."
Line 97: The models are trained on a relatively short time window with hourly data over 1 year. It would be interesting to compare or comment on how such models could compare against other non-linear approaches (for instance, based on simpler and potentially less computationally demanding ML techniques such as XGboost or others)
Line 125: Could the authors justify the model parameters ? The total feature dimension c, the number of attention heads, the internal MLP hidden dimension, the dropout rate, the learning rate. All of that seems relevant, and some variations would probably result in similar performances but it would be great to get some justification.Â
Line 155: Would it be possible to get some information or some comments about how the approach performs at wind farm levels also, instead of the aggregated level ? Performance similarities and differences  ? Also, it would be interesting to get some comparisons or comments on the forecast error variability across other dimensions such as seasons ?
Line 157: It was previously said in the section 2 - data, that the EPS-PC was bias corrected. How does that compare to the observed 7-11% positive bias ?
Line 161: Would "partly" be more adequate instead of "effectively" ? It seems that the diurnal is still present (to a lower extent, but still)
Line 165-166: It is believed that quantifying how much more accurate the Transformer method is compared to MBM would help to highlight the improvement (in % for instance, for intraday).
Line 175: Could the authors justify the peaks at both ends ?Â
Line 190: Could the authors provide a metric to justify that it "aligns much closer to zero"Â (and compare it to another model)
Figure 5: Add the green lines to the legend would avoid some confusion
Line 215: Hourly resolution is chosen (by default), it would be valuable to have some comments on the different resolutions for the calculation of such ramps.
Line 234: The MBM seems to align better with the observations than the Transformer method for intraday +/-30% and for day-ahead +/-15%?
Figure 7:Â Â Could the authors consider adding a reference distribution corresponding to cases (5th bargroup) where no ramping event is observed? Such a comparison could help better assess the ability of the different methods to identify and represent ramping events.
Line 242: It is not striking to me how much more compressed the ensemble spread is ? Could the authors revise or support it with some metric introductions ?
Equation (13): How would the Brier Score be linked to BBSS with 0 buffer ?Â
Figure 10: It is believed that it would be beneficial to also calculate the Brier Score in parallel and comment on both score variations (along with the new introduced method) to see the complementarities and the additional value of the new method.
Line 361: Could any reason be provided for the difference in predictability for the two directions ?
Line 385: Could the authors provide a metric for the enhanced reliability ? On what is it based ?
Line 389:Â That is true the EPS-PC significantly underestimates the ramp event frequencies, but is it not the case for all methods? Is this statement supported by Figure 7 or the BBSS ?
Line 390-391: What is the improvement (in % for instance) for the Transformer and MBM compared to EPS-PC ?
Line 392-394: Could you clarify the statement with quantified values ?
Line 405: Could you also quantify how much more it enhances the baseline ?Â
Â