Validation beyond probability sampling

The preceding sections describe the recommended framework for map validation based on probability sampling and design-based accuracy assessment. In practice, however, this framework is not always feasible. Large-area mapping projects may lack purpose-designed probability samples, existing reference data may have been collected for different objectives, or certain products (particularly continuous-field variables) may require specialized validation strategies. In these situations, validation typically relies on one or more of three approaches: alternative inferential frameworks, existing reference datasets, and product-specific validation protocols. Although these approaches provide valuable evidence of map quality, they generally do not provide the same statistical guarantees as design-based accuracy assessment.

Alternative inferential frameworks. One alternative is to replace the design-based framework with model-based inference. Whereas design-based inference derives its validity from the randomness of the sampling process, model-based inference relies on a statistical model that describes how map errors relate to predictors such as landscape characteristics or data availability. This enables the use of non-probability reference data, such as opportunistic field plots or spatially sparse inventories, by extrapolating accuracy estimates to unsampled locations through the model [1]. The trade-off is that validity depends entirely on whether the assumed model is correct, which may be difficult to verify. Consequently, model-based inference has seen only limited adoption in operational map accuracy assessment [2].

Existing reference datasets. In practice, validation more commonly relies on existing reference datasets than on model-based inference. National Forest Inventories (NFIs) represent a valuable source of high-quality, field-verified reference data. When the inventory sampling design aligns with the target population and class definitions of the mapped product, for example, national land cover assessments using compatible forest definitions, NFI data can support probability-sampling-based validation. However, in practice, NFIs are also used to assess products with different thematic definitions, spatial extents, or target classes, such as global land cover products. In these situations, the original sampling design may no longer correspond to the validation objective, and the resulting assessment should be interpreted with appropriate caution.

In addition, NFI data are typically not openly accessible and often require formal collaboration for use in validation, particularly when applied to global products. Nevertheless, when appropriately matched to the validation objective and sampling framework, NFI-based assessments provide valuable and independent evidence of map quality and remain among the most reliable sources of reference data currently available.

Product-specific validation protocols. Some geospatial products require dedicated validation protocols that extend beyond the general recommendations described above. Biomass mapping provides a representative example. The CEOS WGCV Land Product Validation subgroup has developed dedicated protocols that address the particular challenges of validating continuous-field products, including the use of plot-level field measurements, allometric uncertainty, and spatial scaling from plot to pixel [3]. These protocols recommend combining field inventory data with airborne LiDAR transects to bridge the scale gap between ground plots and satellite-derived estimates. Even under such protocols, however, spatial coverage remains fundamentally limited, and the resulting accuracy characterization applies only to the sampled regions.

Regardless of the validation strategy adopted, the principles of transparency and rigorous reporting remain unchanged. Reference data sources and their limitations should be clearly documented, the distinction between verification and validation should be made explicit, and results should be interpreted as indicative rather than as unbiased estimates of map accuracy.

The preceding sections have traced the full arc of the EO mapping pipeline, from the processing of satellite observations through to a validated, disseminated product. At each stage, we have highlighted choices that are easy to get wrong and hard to diagnose after the fact—from the silent differences between data providers, through the spatial structure of training splits, to the gap between algorithm-derived uncertainty and genuine accuracy assessment. We conclude by summarizing the key recommendations that emerge from this analysis, examining the structural and institutional challenges that the community must confront as ML-based mapping matures from a research activity into an operational capability, and acknowledging the limitations of the present work.

[1]
S. V. Stehman and G. M. Foody, Key issues in rigorous accuracy assessment of land cover products,” Remote Sensing of Environment, vol. 231, p. 111199, 2019, doi: 10.1016/j.rse.2019.05.018.
[2]
A. Tyukavina, S. V. Stehman, A. H. Pickens, P. Potapov, and M. C. Hansen, Practical global sampling methods for estimating area and map accuracy of land cover and change,” Remote Sensing of Environment, vol. 324, p. 114714, 2025.
[3]
L. Duncanson et al., Global Aboveground Biomass Product Validation Best Practices Protocol,” in Best practice protocol for satellite derived land product validation, L. Duncanson, M. Disney, J. Armston, D. Minor, F. Camacho, and J. Nickeson, Eds., Land Product Validation Subgroup (WGCV/CEOS), 2020, p. 222. doi: 10.5067/doc/ceoswgcv/lpv/agb.001.