the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
The IEA Wind Task 57 inflow reconstruction benchmark for a single turbine in simple terrain: real-world and synthetic case studies
Abstract. In experiments to validate wind turbine design codes, the full inflow field moving into the rotor is never measured. Instead, it is reconstructed from spatially limited measurements by using an atmospheric model. As such, the inflow represents a source of uncertainty when validating turbine models. Here, we characterize the behavior and accuracy of modern inflow reconstruction techniques. We compare eight inflow models for nine ~10 minute reference inflows, three from a real-world experiment with a 2.8 MW turbine and six from a synthetic field campaign. We document the models' differences in time series behavior and statistical characteristics like mean profiles, turbulence intensity, and power spectra. Across all case studies, the Superstatistical Mann model had the smallest root mean square error (average of 0.93 m s-1), and TurbSim had the largest (average of 1.19 m s-1). PyConTurb performed similarly to the inflows based on the Mann model. Notably, error time series showed synchronized spikes across models, often corresponding to physically coherent features that were not observed in the hub-height measurements. This study points toward areas for future inflow reconstruction model development, and it provides the foundation for future work that will examine turbine load validation errors in conjunction with inflow errors.
Competing interests: The lead author (Alex Rybchuk) both organized the benchmark and submitted two models (TurbSim and LER) to the benchmark, which has the potential for conflict of interest.
Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.- Preprint
(29673 KB) - Metadata XML
-
Supplement
(22116 KB) - BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on wes-2026-77', Anonymous Referee #1, 19 May 2026
-
CC1: 'Comment on wes-2026-77', J. Gordon Leishman, 21 Jun 2026
The manuscript addresses an important problem, namely, uncertainty in reconstructed inflow fields used for wind turbine model validation. The comparison of several inflow reconstruction methods using both real-world measurements and synthetic LES-based cases is a reasonable benchmark exercise. However, the paper is not fully convincing in its definition and use of “ground truth.” For real-world cases, the reference inflow is derived from lidar measurements, gridding, induction correction, time shifting, Taylor’s frozen-turbulence assumption, and other post-processing steps. The authors acknowledge this limitation, but the analysis still relies on the processed SpinnerLidar product to rank quantitative models. This weakens the significance of small differences among methods, especially given the reported discrepancies among instruments.
The synthetic cases partly address this limitation by providing a fully known reference field, but they introduce a different issue. The benchmark then becomes partly a test of reconstruction performance in an idealized LES atmosphere rather than in the real atmosphere. The paper should more clearly separate conclusions supported by real-world data from those supported by synthetic cases. It should also give a sharper discussion of practical significance. Ranking inflow models by RMSE is useful, but the central question is whether these inflow differences materially affect turbine-load validation, controller validation, or design-code assessment. Since that connection is left mainly to future work, the present manuscript is incomplete.
The paper is also longer and more descriptive than necessary. Some of the case setup, model descriptions, measurement details, and repeated discussion of inflow statistics could be compressed. The main text should focus more directly on reference-data uncertainty, robustness of model ranking, and the practical implications of the observed errors.
Finally, the manuscript discloses that the lead author both organized the benchmark and submitted two models to it. This is a relevant point that should be clarified more fully, because benchmark organization and model participation are not fully independent roles. The manuscript would benefit from a short explanation of how impartiality was maintained, including whether the protocol, case definitions, evaluation metrics, post-processing procedures, and analysis scripts were fixed before model submission. Overall, I recommend a major revision.
Disclaimer: this community comment is written by an individual and does not necessarily reflect the opinion of their employer.Citation: https://doi.org/10.5194/wes-2026-77-CC1 -
RC2: 'Comment on wes-2026-77', Anonymous Referee #2, 19 Aug 2026
GENERAL ASSESSMENT
This manuscript presents a valuable international benchmark of wind-turbine inflow reconstruction methods using complementary real-world lidar measurements and fully observed synthetic large-eddy simulation cases. The combination of real and synthetic datasets is a particular strength. The manuscript is generally well written and is commendably transparent about several limitations of the measurements, models, and relatively small set of cases.
The real-world reference contains substantial uncertainty arising from spatial separation, induction correction, advection, Taylor's frozen-turbulence hypothesis, line-of-sight reconstruction, probe-volume averaging, gridding, synchronization, and signal processing. These uncertainties may be comparable to the differences reported among models. The measurement campaign and the associated measurement operators therefore require more detailed documentation and uncertainty analysis.
Major revision is recommended to improve the quality and readability.
MAJOR COMMENTS
1. The headline ranking is confounded by unequal spatial resolution and smoothing
SSM and LER were generated at 10 m resolution and linearly interpolated to 5 m, whereas the Kaimal- and conventional Mann-based fields were generated at 5 m. The models with the lowest RMSE are also described as producing the smoothest fields. Pixelwise RMSE can favor smooth predictions because small-scale variability and phase errors increase squared error.
All methods should be evaluated at a common effective resolution, preferably by filtering or coarsening every reconstruction and reference field to 10 m before comparison. Results should also be presented over several spatial scales. Otherwise, the conclusion that SSM and LER are more accurate may partly reflect spatial filtering rather than superior reconstruction.
2. The real-world reference uncertainty must be quantified
The SpinnerLidar-derived field is not a direct measurement of the true three-dimensional inflow. It involves observations near 1D rather than at the 3D reconstruction plane, advection using an average velocity, Taylor's frozen-turbulence hypothesis, induction removal, line-of-sight velocity reconstruction, neglected turbine tilt, spatial gridding, temporal processing, and gaps between data files.
Flow evolution over the approximately 2D separation can create an error even for a perfect reconstruction at 3D. The manuscript also shows disagreements of approximately 1 m/s among instruments, which are comparable to the reported reconstruction RMSE. The authors should quantify the uncertainty of the processed reference and perform sensitivity analyses for the advection velocity, time shift, induction correction, and reconstruction assumptions. Real-world RMSE should not be pooled or interpreted together with LES RMSE as though the two references had equal fidelity. The SpinnerLidar-derived field should be described as a processed reference reconstruction rather than unqualified ground truth.
3. Model rankings need uncertainty estimates and statistical assessment
The conclusion that SSM performs best is based on nine short and potentially correlated cases, with only three cases in each broad category. Differences among leading methods are small. For example, the average difference between SSM and LER is approximately 0.04 m/s, which may not be practically or statistically meaningful relative to case-to-case and reference uncertainty.
The authors should report paired case-level differences, uncertainty intervals, and sensitivity to the averaging and weighting procedure. A paired or hierarchical bootstrap could account for case and ensemble variability. The manuscript should distinguish consistent ranking differences from differences that are small relative to uncertainty.
4. Potential distribution advantage for LER in the synthetic cases
The LER neural networks were trained using AMR-Wind simulations with the same large-scale forcing used to generate the synthetic benchmark cases. Although the training and evaluation simulations may be separate, the stable and unstable evaluation cases appear closely matched to LER's training distribution.
The authors should clarify the independence of the training and evaluation datasets, including random seeds, simulated periods, boundary conditions, surface properties, and forcing. LER's strong performance in unstable LES cases should be described as in-distribution performance unless generalization to meaningfully different LES conditions is demonstrated.
5. The measurement campaign and instrument configurations require a more complete description
The manuscript should provide a self-contained description of the RAAW measurement campaign. For each instrument - the nacelle-mounted scanning lidar, SpinnerLidar, profiling lidar, and meteorological mast - the authors should report the make and model, pulsed or continuous-wave operating principle, wavelength, coordinates relative to the turbine, measurement heights and upstream distances, scan trajectory and angular range, sampling and scan-completion frequencies, range-gate or focal-distance settings, probe-volume dimensions or weighting functions, temporal and spatial averaging, velocity-retrieval assumptions, data availability, quality-control criteria, synchronization, and coordinate transformations.
A plan-view and side-view schematic should show the turbine, meteorological mast, profiling lidar, nacelle-lidar measurement plane at 3D, SpinnerLidar scan near 1D, wind direction, scan trajectories, and effective measurement volumes. Spatial separation and different measurement volumes directly affect the apparent reconstruction error and therefore must be clearly documented.
6. Induction effects at the 1D SpinnerLidar reference plane require stronger treatment
The SpinnerLidar measurements are collected approximately one rotor diameter upstream, where rotor induction cannot generally be assumed negligible. The manuscript states that induction effects are estimated and removed, but the correction is not described sufficiently to demonstrate that the corrected field represents freestream inflow.
The authors should provide the induction model and assumptions, turbine operating variables used by the correction, treatment of radial and azimuthal induction variation, treatment of shear and yaw misalignment, validation of the corrected velocities, sensitivity to model parameters, and uncertainty propagated into the reported RMSE. The combined uncertainty from induction correction, advection from 1D to 3D, and turbulent evolution should be quantified. Recent induction-aware nacelle-lidar reconstruction studies should be discussed when explaining and validating this correction.
7. Probe-volume averaging and spatial filtering should be included in the comparison
The SpinnerLidar is a continuous-wave lidar, so its observation represents a weighted average over a finite probe volume rather than a point velocity. Probe-volume averaging can attenuate small-scale turbulence, modify spectra, and smooth coherent structures.
The manuscript should explain how continuous-wave range weighting is treated when converting SpinnerLidar line-of-sight observations into gridded u, v, and w reference fields. It should also describe the probe-volume or range-gate characteristics of the nacelle lidar measuring near 3D. The reconstructed model fields should ideally be passed through instrument-specific virtual-lidar operators, including line-of-sight projection, scan trajectory, probe-volume weighting, gridding, and temporal sampling, before comparison with measurements.
The SpinnerLidar measures near 1D and its field is subsequently shifted to the 3D comparison plane; it does not directly measure at 3D. This geometry and the corresponding measurement operators should be stated precisely.
8. Instrument disagreement in Figures 1 and 2 needs further investigation
Figures 1 and 2 show substantial differences among the nacelle lidar, SpinnerLidar, profiling lidar, and meteorological mast. The manuscript attributes these differences to distinct sampled air masses and signal-processing procedures, but the explanation remains qualitative.
Figure 1 contains instantaneous disagreements exceeding 3 m/s. Exact measurement positions, scan times, effective averaging volumes, and temporal alignment should be reported directly in or near the figure. The authors should determine how much of the disagreement arises from spatial separation, probe-volume filtering, advection, flow evolution, measurement uncertainty, and data processing.
For Figure 2, the profiling lidar and meteorological mast are described as co-located. If so, the relatively large differences in their profiles and estimated shear require closer examination. The authors should report uncertainty bars, common measurement heights, averaging windows, data availability, and the wind-speed retrieval method for each instrument. Rather than stating only that differences of approximately 1 m/s are expected, the authors should demonstrate that the observed disagreement lies within documented measurement and representativeness uncertainty. These instruments supply benchmark constraints, so their disagreement can bias the apparent performance of every model.
9. The fitted Davenport coherence model should be validated over multiple separations
Table 1 lists one pair of Davenport coherence parameters for all three real-world cases. Coherence may depend on atmospheric stability, velocity component, vertical and lateral separation, height, frequency, and averaging period. Parameters fitted using one pair of measurement heights may not represent coherence across the full rotor disk.
The authors should add figures comparing measured and fitted coherence curves for several vertical and, where available, lateral separations. The frequency range used for fitting, uncertainty of the fitted parameters, goodness-of-fit measures, and sensitivity of reconstructed fields to the selected parameters should be reported. If one parameter pair cannot describe all separations adequately, multiple parameter groups or a separation-dependent formulation should be considered. The authors should also justify applying parameters fitted from four hours of observations identically to all three approximately 10-minute unstable cases.
10. The literature review should include recent induction-aware and extreme-turbulence studies
The literature review should be updated to discuss recent studies that directly address limitations relevant to this benchmark. These include nacelle-lidar reconstruction of freestream inflow from measurements inside the turbine induction zone, methods that jointly estimate wind speed, shear, veer, and induction, and nacelle-lidar characterization of extreme turbulence for constrained-field generation and aeroelastic load analysis.
Relevant recent examples include:
- "Identification of wind inflow characteristics from nacelle lidar measurements in the induction zone of a 9 MW wind turbine," Renewable Energy, DOI: 10.1016/j.renene.2025.124523.
- "Reconstructing the upwind field of wind turbines using LiDAR data," Sustainable Energy Technologies and Assessments, DOI: 10.1016/j.seta.2025.104382.
- "Extreme turbulence effects on wind turbine loads: A case study for the North China Plain using nacelle lidar," Renewable Energy, DOI: 10.1016/j.renene.2026.125774.
The authors should explain how the present benchmark differs from and advances these induction-aware and extreme-turbulence reconstruction studies.
MINOR COMMENTS
1. The conclusion that accuracy generally improves with model complexity should be qualified because complexity is confounded with inputs, resolution, and implementation.2. Explain why SSM uses an intermittency coefficient derived from unstable observations for stable and synthetic cases.
3. Report the SSM realization rejection rate associated with the 10% standard-deviation acceptance criterion.
4. Clarify whether machine-learning processing of the nacelle-lidar constraint introduces an LES-like spectral shape that could favor LER or other smooth reconstruction methods.
5. Explain whether the constraint preprocessing used for PCT2 and PCT3 is included when interpreting their performance relative to other PyConTurb submissions.
6. Add uncertainty bars or ranges to the instrument comparisons in Figures 1 and 2.
7. Add the exact instrument coordinates and scan planes to the captions or a dedicated measurement-layout figure.
8. Correct the duplicated wording "differences between hub-height differences" in the conclusions.
9. The reference "Hannesdóttir, 2026" is incomplete and should include full bibliographic information.SUGGESTED DECISION STATEMENTThis manuscript presents an important and thoughtfully designed inflow-reconstruction benchmark using complementary real-world and synthetic datasets. However, the current model ranking is confounded by substantial differences in constraint information, preprocessing, spatial resolution, ensemble size, parameter estimation, and realization selection. In addition, uncertainty in the real-world reference may be comparable to the reported differences among models, while the primary accuracy metric considers only the streamwise velocity. The measurement geometry, induction correction, probe-volume effects, coherence fitting, and inter-instrument disagreement also require more complete documentation and uncertainty analysis. I recommend major revision to introduce fairer or stratified comparisons, quantify reference and ranking uncertainty, adopt ensemble-size-robust and multicomponent metrics, and qualify conclusions regarding model accuracy.
Citation: https://doi.org/10.5194/wes-2026-77-RC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 327 | 155 | 22 | 504 | 78 | 37 | 25 |
- HTML: 327
- PDF: 155
- XML: 22
- Total: 504
- Supplement: 78
- BibTeX: 37
- EndNote: 25
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The motivation of the paper is three folded it is stated in the beginning of the paper, but there is no clear answer in the conclusion part of the paper. It would be an improvement if the main objective was clearer and not divided into three parts. The paper gives a very detailed list and figures of the differences, but brings in so many details that it is hard to follow. An overview in the conclusion part pinpointing the “common strength and weaknesses” stated in line 531 would be helpful so that it is clear what is found in the study. Here it is also stated that “context on how large inflow errors can be …” I don’t think that there is inflow errors? Is this in the simulations or in the measurement or is it the differences? I think that there should put in a bit more work on the conclusion to show that the objective (or motivation) is responded to in a clear manner.
I find the paper to be very interesting, and think it would be of broad international interest, and the simulations and comparisons are of good quality. However, I am missing information on who are the participants? Is there a reason why the institutes are not mentioned? There are also a lot of times where “we” are used, and for clarity of the paper, the usage of “we” should maybe be omitted. An example is in line 161 “… participants reconstructed inflows …” and then in the next sentence “ .. we compared to data …” To the reader it seems like all the authors compared the data, and the participants are not a part of this group? For clarity try to keep “we” out of the Methods section.
In many sections is referred to Supplementary Material. Where is this found? I also see that there is a to appendices, is it necessary with both Supplementary Material and appendices?
I also miss an illustration or information about at which height is the met mast measuring, what is the participants given as information (constraint information?), and exactly what measurements are the “ground truth”. Consider making some of the information tabular so it is easier to find.
In line 324 and onward the model ensembles and how often they fall within the ground truth is discussed, but there is no mentioning that the number of ensembles are different. Is this taken into account when discussing the results anywhere? In Figure 7 the black line is referred to as Spinner, should this have been ground truth?
In line 454, failure modes are mentioned. Is this the correct use of words? Are the models failing, or are they missing some features that exists in real life? Is it possible for these models to generate the coherent gust or the veer, would it really help to run more ensembles?
Minor corrections:
Define Davenport coherence parameters in line 153 and the parameters in line 155
Line 177, Fig 1 does not contain a gap
Line 275, give a clear definition of the 2-dimensional wind angle
Line 292 Based on previous
It would be helpful with grid lines in figure 14