the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
A Two-Step Spatio-Temporal Framework for Turbine-Height Wind Estimation at Unmonitored Sites from Sparse Meteorological Data
Abstract. Accurate estimates of wind speeds at wind turbine hub heights are crucial for both wind resource assessment and day-to-day management of electricity grids with high renewable penetration. In the absence of direct measurements, parametric models are commonly used to extrapolate wind speeds from observed heights to turbine heights. Recent literature has proposed extensions to allow for spatially or temporally varying vertical wind gradients, that is, the rate at which wind speed changes with height. However, these approaches typically assume that reference height and hub height measurements are available at the same locations, which limits their applicability in operational settings where meteorological stations and wind farms are spatially separated. In this paper, we develop a two-step spatio-temporal framework to estimate turbine height wind speeds using only open-access observations from sparse meteorological stations. First, a non-parametric generalized additive model is trained on reanalysis data to perform vertical height extrapolation. Second, a spatial Gaussian process model interpolates these hub-height estimates to wind farm locations while explicitly propagating uncertainty from the height extrapolation stage. The proposed framework enables the construction of high-resolution, sub-hourly turbine-height wind speed time series and spatial wind maps using data available in real time, capabilities not provided by existing reanalysis products. We further provide calibrated uncertainty estimates that account for both vertical extrapolation and spatial interpolation errors. The approach is validated using hub-height measurements from seven operational wind farms in Ireland, demonstrating improved accuracy relative to ERA5 reanalysis while relying solely on real-time, open-access data.
- Preprint
(4161 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 17 Aug 2026)
-
RC1: 'Comment on wes-2026-101', Anonymous Referee #1, 28 Jul 2026
reply
-
AC1: 'Reply on RC1', Eamonn Organ, 28 Jul 2026
reply
We thank the reviewer for their thoughtful and constructive comments. We address several points requiring clarification below for the benefit of the ongoing discussion. The manuscript will be revised accordingly.
1) Time periods are unclear:
The statistical model is trained using reanalysis data for the full year of 2018 using the New European Wind Atlas (NEWA), which is the most recently available year.
The trained model is then applied to meteorological observations from the full year of 2023, generating hub-height wind speed predictions for the full year 2023. These predictions are subsequently validated against observed wind farm wind speeds from the same period.The non-overlapping time periods, (beyond the data limitations, as other more recent reanalysis datasets exist), is designed to reflect the operational use case as in practice all reanalysis datasets have a lag, so we wish to evaluate the ability of a model trained on historical reanalysis applied to current meteorological data.
Nevertheless, we agree that additional analyses spanning multiple years would provide a more robust assessment of inter-annual variability and seasonal performance. Additional analyses are currently being undertaken and will be incorporated into the revised manuscript.
The manuscript will be updated to more clearly describe the training, prediction, and validation periods.
2) Data Recovery
The seven wind farms each contain at least 98% of the expected 10-minute observations during 2023, with three of the seven farms achieving complete (100%) data coverage.
These data recovery statistics will be reported in the revised manuscript.
3) Bias
The bias reported in the manuscript is currently defined as:
Bias = Observation − Prediction
We agree that the alternative convention,
Bias = Prediction − Observation
is more commonly adopted within the literature and provides a more intuitive interpretation of positive and negative values. The manuscript will be updated accordingly and the bias definition will be explicitly stated.
4) Farm Location
The wind farm coordinates used within the analysis correspond to latitude and longitude values provided by the asset manager. Inspection suggests that these coordinates generally correspond to the location of a turbine within the wind farm.
While turbine-level wind speed observations are available, turbine identifiers and turbine coordinates were not provided. Consequently, the precise spatial relationship between the reported wind speeds and the supplied farm coordinates cannot be determined.
We agree that the choice of validation location warrants further investigation. Additional analyses are being undertaken to compare predictions based on a single representative location within a wind farm against approaches that utilise information aggregated across the full wind farm footprint. These results will be incorporated into the revised manuscript.
5) Additional analyses
The reviewer raises several important questions regarding seasonal performance, diurnal variability, wind regime dependence, and the factors influencing differences in performance between wind farms. More extensive analyses addressing these aspects are available and will be incorporated into an expanded Results and Discussion section in the revised manuscript.
We again thank the reviewer for their detailed comments and helpful suggestions.
Eamonn Organ,
On behalf of all co-authors
Citation: https://doi.org/10.5194/wes-2026-101-AC1
-
AC1: 'Reply on RC1', Eamonn Organ, 28 Jul 2026
reply
-
RC2: 'Comment on wes-2026-101', Anonymous Referee #2, 02 Aug 2026
reply
Review of: “A Two-Step Spatio-Temporal Framework for Turbine-Height Wind Estimation at Unmonitored Sites from Sparse Meteorological Data”
General Comments:
The manuscript presents a practical and relevant two-step framework combining vertical extrapolation (via GAMs) and spatial interpolation (via Gaussian Processes) to estimate hub-height wind speeds from sparse surface observations. The objective of relying strictly on open-access, real-time data streams for operational contexts is commendable.
I have read, and I agree with the substantive concerns raised by Referee #1 (RC1), particularly regarding the need to clarify the temporal structure of the datasets, evaluate model performance across seasonal and diurnal cycles, and address potential confounding factors when utilizing farm-average wind speeds and wake-loss adjustments. The author’s initial responses (AC1) are encouraging, and I look forward to seeing these additions incorporated into the revised manuscript.
To further strengthen the manuscript, I request that the authors address the following two major and one minor comments:
Major Comments:
- Handling local topography and coastal heterogeneity in the Spatial GP Model: Ireland features varied terrain and distinct coastal-to-inland transitions. Standard Gaussian Process Models risk over-smoothing local topographic acceleration or abrupt changes in boundary layer dynamics near coastlines. Could the authors clarify how terrain elevation and distance to coast are accounted for within the spatial GP covariance structure or feature set? If these are not explicitly included, please discuss how topographic smoothing impacts performance at elevated or coastal wind farm sites relative to inland stations.
- Sub-Hourly temporal variability & Downscaling Mechanics: 10-meter meteorological observations are typically recorded as hourly or 10-minute spot/averaged measurements, while reanalysis datasets (like NEWA) are provided at coarse hourly intervals. In Section 2.1, the authors perform linear interpolation on hourly station observations to bring them to sub-hourly/10-minute resolutions. Linear interpolation of wind speed smooths out turbulent fluctuations and sub-hourly wind variability, which artificially suppresses the true variance of hub-height wind estimates. Could the authors address how linear interpolation impacts sub-hourly variance and error metrics compared to sub-sampling at top-of-hour intervals? Furthermore, how does the framework handle non-stationary variance during rapidly changing weather conditions such as during the passage of a cold front?
Minor Comments:
- While the authors explained in AC1 why 2018 NEWA data was chosen to mirror operational lag, please briefly comment in the text on whether model sensitivity was tested across multiple reanalysis years to ensure the GAM parameters aren't overly tuned to 2018 atmospheric patterns.
Citation: https://doi.org/10.5194/wes-2026-101-RC2
Data sets
Processed meteorological and reanalysis data Eamonn Organ https://github.com/EamonnO22
Model code and software
R Code Eamonn Organ https://github.com/EamonnO22
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 155 | 39 | 13 | 207 | 10 | 12 |
- HTML: 155
- PDF: 39
- XML: 13
- Total: 207
- BibTeX: 10
- EndNote: 12
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Organ et al. (2026) propose a new framework for geographically broad wind resource assessment at hub height using 10 m observations and simulated wind datasets. The concept is interesting in that it utilizes the only wind measurements that are widespread and makes them applicable to higher heights and extensive locations. While the proposed technique is intriguing, the manuscript contains many gaps that leave the reader uncertain about its success and applicability.
From what I can tell, you are training with 10 m observations using the year 2023 (Line 112) and NEWA using the year 2018 (Line 136) to predict hub height winds at the wind farms over an undefined (as far as I could tell?) period of time for validation. A year or multiple years? Or less? All I can find are several figures showing 5 days, but I’m assuming the metrics in the results tables reflect longer periods of time than that? What is the data recovery and how does it vary from wind farm to wind farm? Would you expect the performance to improve if you used the consistent time periods across all datasets? Were 2018 and 2023 average, high, or low wind resource years in Ireland?
Models are fitted for each calendar month, but there is no discussion in the results on how the proposed technique performs seasonally (or diurnally). As you mention that grid operators could be interested in this technique, such analyses are critical.
Validation using the fastest wind speed in the farm is a simplistic approach but could be justified with more information. Assuming the farms are large(?), what are the impacts on the results given that your downscaled model is 250 m resolution? What location in the farm are you validating against? The center, despite the implication that the fastest wind speeds will occur at turbines at the edges of the farms typically?
On a similar note, how close are the wind farms in your sample to other wind farms? Figure 7 implies that the fastest wind speed in your farms could still be waked by other farms if close enough.
I think the wake analysis using the average wind farm speed should be removed entirely. Calculating error metrics to define the performance of your technique based on multiple unknowns is concerning. Line 428, in particular, hints of cherry picking when you don’t establish an observation-based wake loss: “the 15% wake-loss adjustment yields the lowest RMSE and a mean residual close to zero, indicating stronger agreement with observed farm average wind speeds.” This is acknowledged as a limitation in the conclusions section, which is appreciated, but doesn’t warrant discussion in the results section without a more robust wake assessment. Why isn’t ERA5 treated to the same wake loss assessment?
Given the uninspiring validation of the technique (Line 417: “The proposed model achieves lower RMSE than ERA5 at three out of seven wind farms”) and the very short results section (particularly if the wake assessment is removed as suggested), I don’t think this manuscript is ready for publication in its current form. It’s fine to try a technique and end up with unsatisfactory or inconsistent performance. But in order to be of value to the wind energy community, there needs to be a robust investigation into why the performance is what it is that is lacking in this manuscript. Is the technique struggling with faster/slower wind speeds? What is the performance like as a function of hub height or proximity to the coast or 10 m observation density or wind farm density? Does it perform better in winter versus summer, or night versus day?
Specific comments:
Please define bias somewhere. I think you are using observations – model, but other studies use model – observations to make it easy to show that the model is overperforming (positive bias) or underperforming (negative bias).
Line 41: “In the absence of wind measurements at turbine hub height, a common alternative is to vertically extrapolate wind speeds from observations at 10 m or other near-surface levels.” Please include some references to support “it is common” and to identify which audiences do this. Large-scale wind farm folks will utilize tall towers and lidars. Many small project developers rely on simulated wind resource datasets.
Section 2.1: I recommend removing the 5 observations that only have hourly resolution instead of introducing more error by linear interpolation. Given the fluctuations that occur in wind speed measurements over finer resolutions, these hourly observations are not comparable with other 18.
Line 127: What do you mean by “most reliable?” Smallest magnitude bias? Best correlation to capture fluctuations in the wind?
Line 163: “Each wind farm comprises of multiple turbines…” Three? Dozens? Hundreds?
Figure 11: It would be helpful to swap the axes in the profile plots so that height is the y-axis.
Figure 12: This figure confused me for a bit, but I eventually figured out that the attributes portrayed in the two legends, color and marker size, characterize the same metric. It’s confusing having the two legends, especially when the ordering of 0.15-0.30 m/s reverses between the two. Can you remove the color bar and color the markers accordingly? Or just use the colors and not the sizes?
Table 2: It would be interesting to include the bias.
Line 398: Again, I think you’re hurting rather than helping your analysis here by interpolating hourly data to 10-minute resolution. Can you simply subsample to the top of the hour instead?
Line 409: Single sentence paragraph, should be merged with the next paragraph.