Common misconceptions in map validation

Despite the central role of validation in establishing the scientific credibility of geospatial products, accuracy assessment methodology is frequently under-reported, and reference datasets are rarely made publicly available [1], [2]. Four concepts are frequently conflated: model evaluation, algorithm-derived uncertainty, map-to-map comparison, and map validation. While the first three provide valuable information about model performance or product consistency, they do not constitute rigorous validation of the final map product. In this section, we focus specifically on map validation, namely the independent assessment of map quality using reference data, and clarify how it differs from these related but distinct concepts.

Model evaluation. Model evaluation is often confused with map validation. In many studies, standard ML practices (such as evaluating performance on a held-out test set) are presented as evidence of map quality. Model evaluation is indeed an important component of map development, as the final map is generated through spatial application of the model across a large geographic domain. Consequently, model performance and map quality are inherently related. However, they are not equivalent. Model evaluation primarily measures how well the model generalizes to a held-out test dataset, whereas map validation concerns the quality of the final spatial product across the full mapped domain. A held-out test set is typically not a probability sample of the map and may not represent the full range of landscapes, class distributions, and edge cases encountered in the final product. Even when the training sample is itself a probability sample and a portion is set aside for evaluation, this is generally not recommended: biases and errors present in the training data may transfer to the reference data [3]. More broadly, when map and validation data originate from the same source or interpretation process, their errors tend to be correlated, and correlated errors inflate quality estimates, while independent errors deflate them [4]. For these reasons, strong test-set performance of the model does not guarantee, and cannot substitute for, an independent assessment of map quality.

Algorithm-derived uncertainty. Metrics produced by ensemble variance, posterior probabilities, or model confidence scores quantify internal model uncertainty, not agreement with ground truth. These measures reflect what the model “thinks” about its own predictions and should not be interpreted as accuracy estimates [5].

Map-to-map comparison. Comparison against an existing map product is another practice that is frequently mistaken for validation. In principle, validation should rely on reference data of higher quality than the product being evaluated, such as field observations or carefully interpreted high-resolution imagery. Existing map products, however, are often the result of inference procedures, including machine learning models or rule-based gridded approaches, and therefore contain their own uncertainties and biases. While map-to-map comparisons can provide valuable qualitative insights and help identify systematic differences between products [6], they should not be interpreted as ground truth validation. What these comparisons measure is inter-product consistency, not map quality. They are more accurately described as consistency checks and should be reported as such.

[1]
S. V. Stehman and G. M. Foody, Key issues in rigorous accuracy assessment of land cover products,” Remote Sensing of Environment, vol. 231, p. 111199, 2019, doi: 10.1016/j.rse.2019.05.018.
[2]
G. M. Foody, Status of land cover classification accuracy assessment,” Remote Sensing of Environment, vol. 80, no. 1, pp. 185–201, Apr. 2002, doi: 10.1016/s0034-4257(01)00295-4.
[3]
A. Tyukavina, S. V. Stehman, A. H. Pickens, P. Potapov, and M. C. Hansen, Practical global sampling methods for estimating area and map accuracy of land cover and change,” Remote Sensing of Environment, vol. 324, p. 114714, 2025.
[4]
G. M. Foody, Assessing the accuracy of land cover change with imperfect ground reference data,” Remote Sensing of Environment, vol. 114, no. 10, pp. 2271–2285, Oct. 2010, doi: 10.1016/j.rse.2010.05.003.
[5]
A. Tyukavina et al., Land Cover and Change Map Accuracy Assessment and Area Estimation Good Practices Protocol. 2025. doi: 10.5067/doc/ceoswgcv/lpv/lc.001.
[6]
A. H. Strahler et al., Global Land Cover Validation: Recommendations for Evaluation and Accuracy Assessment of Global Land Cover Maps,” European Commission, Joint Research Centre, Institute for Environment; Sustainability, Ispra, Italy, GOFC-GOLD Report No. 25, 2006.