Design-based accuracy assessment

A design-based accuracy assessment is more than the calculation of accuracy metrics from a set of reference samples. Rather, it is a statistical framework in which the validity of the resulting accuracy estimates depends jointly on how reference samples are selected, how their reference labels are determined, and how the results are analyzed. Accordingly, three components form the foundation of a statistically rigorous accuracy assessment [1]:

These components must be considered jointly [2], [3], [4]. Across all stages, a clear and consistent problem definition is critical and should precede the validation process. In particular, the definitions used in sampling and labeling (such as what constitutes “forest” or the boundaries between land-cover classes) must be explicitly defined and aligned with the broader conceptual framing of the map product.

Sampling design

Sampling design determines how locations are selected for validation across the mapped domain. Design-based accuracy assessment derives its validity from the randomness of the sample selection process rather than from model assumptions [2]. Under this framework, every location within the study area must have a known, non-zero probability of being selected. This requirement is known as the inclusion probability criterion [4], [5].

A wide range of probability sampling strategies can be applied, including simple random, systematic, and stratified approaches. The choice of sampling design, as well as the definition of strata, directly influences the statistical properties of the resulting accuracy estimates [6]. While simple random sampling provides an unbiased representation of the mapped domain, it is often inefficient because common classes dominate the sample, leaving insufficient observations for rare but important classes. Consequently, stratified random sampling is widely recommended as an adequate general-purpose design [6], [7]. It allows better control over sample allocation across strata (e.g., land cover classes) while maintaining a probabilistic framework for unbiased inference.

Once a sampling strategy has been selected, an appropriate sample size must be determined. Sample size should balance statistical precision against the practical cost of reference labeling. Minimum sample sizes per class are generally recommended to ensure reliable class-specific accuracy estimates, and formal procedures for determining sample size under stratified designs are well established [3], [6].

A further consideration arises when validating maps containing rare classes, such as deforestation or urban expansion. Different sampling designs affect the balance between the precision of user’s and producer’s accuracy, particularly under severe class imbalance [8]. In such cases, buffer-based stratification methods can substantially improve estimation. Olofsson et al. [8] proposed introducing an additional stratum around mapped change areas, thereby increasing the likelihood of sampling near class boundaries where omission errors are most likely to occur and improving the estimation of rare-class areas.

Several operational implementations demonstrate how these sampling principles can be applied at large scales. A widely adopted approach is the stratified one-stage cluster approach [5], which combines stratification and clustering to balance statistical efficiency with the logistical constraints of human interpretation. The study area is first divided into strata based on attributes such as biomes, after which clusters, referred to as primary sampling units (PSUs), are randomly selected within each stratum. In a one-stage design, every pixel within a selected cluster is interpreted, and each sample is assigned a weight based on its inclusion probability, enabling unbiased estimation of overall accuracy and class-specific areas [9].

A representative example of this approach is the multi-purpose Global Land Cover Validation dataset developed for the Copernicus Global Land Service [9], [10]. The dataset comprises more than 21,000 PSUs globally and has been used to validate products including CGLS-LC100m and ESA WorldCover, demonstrating that CEOS WGCV Validation Stage 4, the highest validation stage currently achieved for satellite-derived land cover products and characterized by sustained validation across successive product updates, is practically feasible at the global scale.

Response design

Response design specifies how reference labels are assigned at the sampled locations. It encompasses the choice of spatial assessment unit (pixel, spatial window, or object-based), the source of reference information (field observations, high-resolution imagery, or crowdsourced data), the labeling protocol, and the definition of agreement between map and reference labels. These decisions should be explicitly documented, as they directly affect the interpretation and reproducibility of the validation results [4].

Reference information may be obtained from field observations, visual interpretation of high-resolution imagery, crowdsourced platforms, or existing reference datasets. Field observations generally provide the highest-quality reference data but are often unavailable at the spatial and temporal scales required for large-area validation. Visual interpretation of high-resolution imagery by trained experts is the most commonly adopted approach. Although widely regarded as the best practical alternative, it is not fully reproducible: interpreter disagreement rates of 30% or more have been reported [11], with agreement varying substantially across land-cover classes [12]. Crowdsourced platforms, such as Geo-Wiki, provide broader geographic coverage but generally at the expense of more variable labeling quality [13], [14], [15]. More generally, the suitability of any reference data source depends on the mapping objective, spatial scale, and thematic classes being evaluated [16].

Several practices help improve the reliability of reference labels and reduce uncertainty. Reference data should be produced by interpreters who are independent from the map makers and unaware of map labels at sample locations, to avoid confirmation bias. Using multiple interpreters with consensus-based labeling further improves reliability [4]. Spatial uncertainty arising from imperfect co-registration between the map and reference data can be mitigated by considering alternative labels within the spatial neighborhood of each sample unit, for instance by assigning the majority class within a \(3\times3\) kernel around each reference pixel [17], [18]. Finally, reference observations should be temporally consistent with the map epoch to avoid discrepancies caused by land cover change between the acquisition dates of the map and the reference data [9].

Analysis

Analysis refers to the statistical procedures used to derive accuracy metrics from the validation sample. For categorical maps, this typically involves constructing an error (confusion) matrix, which should be expressed in terms of area proportions rather than sample counts [4]. When sampling is stratified or otherwise non-uniform, unequal inclusion probabilities between strata must be accounted for, since sample sites are typically not allocated proportionally to strata areas [19], [20]. Sample weights, calculated as the inverse of inclusion probabilities, are therefore used to construct an area-weighted confusion matrix. This allows unbiased estimates of overall accuracy, class-specific accuracies, and area proportions to be derived.

For categorical maps, user’s and producer’s accuracy (or equivalently, precision and recall), along with other metrics quantifying omission and commission errors, should be reported for individual classes in addition to overall accuracy. This is particularly important when a target class occupies a small proportion of the map, as high overall accuracy can mask poor performance on the class of interest [4]. In addition to traditional accuracy metrics, alternative decompositions of classification error, such as quantity disagreement and allocation disagreement [21], can provide complementary insight into the nature of map errors and are commonly adopted in the literature. For continuous-field products (e.g., tree cover fraction), adaptations of the standard confusion matrix framework, along with continuous accuracy metrics such as mean absolute error (MAE) and root mean square error (RMSE), are available [4], [5].

Accuracy metrics should be accompanied by measures of uncertainty, such as standard errors or confidence intervals, to reflect the variability inherent in the sampling process [22]. In practice, standard errors of accuracy metrics are often not reported, making it difficult for users to assess whether the provided estimates are reliable. Bootstrapping with replacement is commonly used to derive confidence intervals because it does not rely on parametric assumptions about the error distribution, although it does not correct systematic sampling bias in the validation design.

Reporting and transparency

Transparent reporting of validation procedures is essential for ensuring credibility, reproducibility, and appropriate interpretation of map products [22]. At a minimum, studies should document the sampling design (including strategy, sample size, and allocation), response design (reference data sources, labeling protocol, and uncertainty considerations), and analysis methods (accuracy metrics, area estimation procedures, and uncertainty measures such as confidence intervals). Any deviations from probability-based sampling or limitations in reference data should be clearly acknowledged.

For products that are operationally updated, accuracy should be monitored for temporal stability, not only assessed at a single point in time. Tsendbazar et al. [9] propose to quantify stability by calculating an index from class accuracies at different moments in time, following the requirement that omission and commission error should each remain below 15% with comparable stability over time [23].

[1]
S. V. Stehman and R. L. Czaplewski, Design and analysis for thematic map accuracy assessment: fundamental principles,” Remote sensing of environment, vol. 64, no. 3, pp. 331–344, 1998.
[2]
S. V. Stehman, Statistical rigor and practical utility in thematic map accuracy assessment,” Photogrammetric Engineering and Remote Sensing, vol. 67, no. 6, pp. 727–734, 2001.
[3]
R. G. Congalton and K. Green, Assessing the Accuracy of Remotely Sensed Data: Principles and Practices, Third Edition. CRC Press, 2019. doi: 10.1201/9780429052729.
[4]
A. Tyukavina, S. V. Stehman, A. H. Pickens, P. Potapov, and M. C. Hansen, Practical global sampling methods for estimating area and map accuracy of land cover and change,” Remote Sensing of Environment, vol. 324, p. 114714, 2025.
[5]
B. Pengra, J. Long, D. Dahal, S. V. Stehman, and T. R. Loveland, A global reference database from very high resolution commercial satellite data and methodology for application to Landsat derived 30 m continuous field tree cover data,” Remote Sensing of Environment, vol. 165, pp. 234–248, Aug. 2015, doi: 10.1016/j.rse.2015.01.018.
[6]
P. Olofsson, G. M. Foody, M. Herold, S. V. Stehman, C. E. Woodcock, and M. A. Wulder, Good practices for estimating area and assessing accuracy of land change,” Remote Sensing of Environment, vol. 148, pp. 42–57, May 2014, doi: 10.1016/j.rse.2014.02.015.
[7]
S. V. Stehman, Sampling designs for accuracy assessment of land cover,” International Journal of Remote Sensing, vol. 30, no. 20, pp. 5243–5272, 2009, doi: 10.1080/01431160903131000.
[8]
P. Olofsson et al., Mitigating the effects of omission errors on area and area change estimates,” Remote Sensing of Environment, vol. 236, p. 111492, Jan. 2020, doi: 10.1016/j.rse.2019.111492.
[9]
N. Tsendbazar, M. Herold, L. Li, A. Tarko, and M. Duerauer, Towards operational validation of annual global land cover maps,” Remote Sensing of Environment, vol. 266, p. 112686, 2021, doi: 10.1016/j.rse.2021.112686.
[10]
N.-E. Tsendbazar et al., Developing and applying a multi-purpose land cover validation dataset for Africa,” Remote Sensing of Environment, vol. 219, pp. 298–309, Dec. 2018, doi: 10.1016/j.rse.2018.10.025.
[11]
R. L. Powell et al., Sources of error in accuracy assessment of thematic land-cover maps in the Brazilian Amazon,” Remote Sensing of Environment, vol. 90, no. 2, pp. 221–234, 2004, doi: 10.1016/j.rse.2003.12.007.
[12]
B. W. Pengra et al., Quality control and assessment of interpreter consistency of annual land cover reference data in an operational national monitoring program,” Remote Sensing of Environment, vol. 238, p. 111261, Mar. 2020, doi: 10.1016/j.rse.2019.111261.
[13]
S. Fritz et al., Geo-Wiki: An online platform for improving global land cover,” Environmental Modelling & Software, vol. 31, pp. 110–123, 2012, doi: 10.1016/j.envsoft.2011.11.015.
[14]
L. See et al., Crowdsourcing, Citizen Science or Volunteered Geographic Information? The Current State of Crowdsourced Geographic Information,” ISPRS International Journal of Geo-Information, vol. 5, no. 5, 2016, doi: 10.3390/ijgi5050055.
[15]
C. C. Fonte, L. Bastin, L. See, G. Foody, and F. Lupia, Usability of VGI for validation of land cover maps,” International Journal of Geographical Information Science, vol. 29, no. 7, pp. 1269–1291, 2015, doi: 10.1080/13658816.2015.1018266.
[16]
N. E. Tsendbazar, S. de Bruin, and M. Herold, Assessing global land cover reference datasets for different user communities,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 103, pp. 93–114, May 2015, doi: 10.1016/j.isprsjprs.2014.02.008.
[17]
P. Xu, N.-E. Tsendbazar, and M. Herold, Comparative validation of recent 10 m-resolution global land cover maps,” Remote Sensing of Environment, vol. 314, p. 114316, 2024, doi: 10.1016/j.rse.2024.114316.
[18]
N. E. Tsendbazar, S. de Bruin, B. Mora, L. Schouten, and M. Herold, Comparative assessment of thematic accuracy of GLC maps for specific applications using existing reference data,” International Journal of Applied Earth Observation and Geoinformation, vol. 44, pp. 124–135, Feb. 2016, doi: 10.1016/j.jag.2015.08.009.
[19]
P. Olofsson et al., A global land-cover validation data set, part I: fundamental design principles,” International Journal of Remote Sensing, vol. 33, no. 18, pp. 5768–5788, Mar. 2012, doi: 10.1080/01431161.2012.674230.
[20]
J. D. Wickham, S. V. Stehman, J. A. Fry, J. H. Smith, and C. G. Homer, Thematic accuracy of the NLCD 2001 land cover for the conterminous United States,” Remote Sensing of Environment, vol. 114, no. 6, pp. 1286–1296, 2010, doi: 10.1016/j.rse.2010.01.018.
[21]
R. G. Pontius and M. Millones, Death to Kappa: birth of quantity disagreement and allocation disagreement for accuracy assessment,” International Journal of Remote Sensing, vol. 32, no. 15, pp. 4407–4429, Aug. 2011, doi: 10.1080/01431161.2011.552923.
[22]
S. V. Stehman and G. M. Foody, Key issues in rigorous accuracy assessment of land cover products,” Remote Sensing of Environment, vol. 231, p. 111199, 2019, doi: 10.1016/j.rse.2019.05.018.
[23]
W. Lahoz, Systematic Observation Requirements for Satellite-Based Products for Climate, 2011 Update, Supplemental Details to the Satellite-Based Component of the Implementation Plan for the Global Observing System for Climate in Support of the UNFCCC (2010 Update),” Jan. 2011.