Looking forward

The rapid maturation of ML-based EO mapping has exposed several structural challenges that individual research groups cannot resolve alone and that will shape the field’s trajectory over the coming years. We highlight six directions that we consider both urgent and tractable.

Foundation models and transfer learning for EO

Geospatial foundation models pre-trained on large satellite archives [1], [2], [3], [4] promise to reduce the labeled-data requirements of downstream tasks and to learn representations that generalize across sensors, resolutions, and geographies. However, several open questions remain. The optimal pre-training data mix — in terms of sensor diversity, geographic coverage, temporal extent, and preprocessing level — is not yet established, and current models differ widely in these choices. Whether a single model can meaningfully span the spectral and geometric heterogeneity of the full EO data landscape, from sub-meter optical imagery to coarse-resolution SAR composites, or whether sensor-family-specific models will prove more effective, is an empirical question that the community is only beginning to address. Equally important is the development of systematic evaluation protocols: without standardized benchmarks that cover diverse geographies, tasks, and sensor configurations, it is difficult to compare foundation models on equal footing or to determine when fine-tuning outperforms training from scratch.

Multi-modal and multi-scale learning

Many of the map products surveyed in this paper rely on a single sensor or a fixed combination of two. Yet the EO data landscape offers a far richer array of complementary modalities — optical, SAR, LiDAR, thermal, gravimetric — each with distinct spatial, temporal, and spectral characteristics. Architectures capable of ingesting arbitrary subsets of available modalities at their native resolutions, and of gracefully handling missing inputs at inference time [5], [6], [7], represent a promising direction. The challenge extends beyond architecture design to data infrastructure: aligning heterogeneous sources in space and time at global scale, while preserving each modality’s native information content, remains an engineering problem that lacks standardized solutions.

Standardization of preprocessing and provenance tracking

A recurring finding throughout this paper is that nominally identical data products differ across platforms in ways that are poorly documented and difficult to detect. The absence of standardized, reproducible preprocessing pipelines — particularly for Sentinel-1 — means that two researchers starting from the same mission can arrive at materially different analysis-ready datasets without being aware of the divergence. Community-endorsed reference implementations for common preprocessing chains, accompanied by machine-readable provenance metadata that records every transformation applied to the data, would substantially improve the reproducibility and comparability of large-scale mapping efforts. Emerging specifications such as TACO [8] and initiatives around STAC extensions represent steps in this direction, but broader adoption and institutional support are needed.

Validation as infrastructure

Current validation practice is overwhelmingly project-specific: each mapping effort designs its own accuracy assessment, collects its own reference data, and reports results using its own metrics and stratification. This fragmentation makes cross-product comparison difficult and wastes collective resources. Establishing and maintaining globally distributed, multi-purpose validation datasets — updated regularly, designed around probability sampling, and made openly available — would transform validation from a per-project cost into shared infrastructure [9]. Such datasets should be designed with spatial error correlation in mind, sampling densely enough in representative regions to support variogram estimation and rigorous aggregate uncertainty reporting.

Uncertainty quantification as standard practice

Despite the methodological maturity of many UQ approaches (Section UQ Methods), uncertainty layers remain the exception rather than the norm in published map products. Making UQ a default component of every released map will require not only methodological advances — particularly in scalable methods that capture both aleatoric and epistemic uncertainty while remaining computationally tractable — but also community agreement on minimum reporting standards: which calibration diagnostics to include, how to communicate uncertainty to non-technical users, and how to handle the spatial aggregation problem that renders naive pixel-level uncertainty misleading at regional scales.

Open data infrastructure and long-term sustainability

The accessibility of EO data has improved enormously over the past decade, but the ecosystem remains fragile. Key re-distribution platforms have changed their pricing and access models with little advance notice, and the long-term availability of any commercial hosting service cannot be guaranteed. Building a durable infrastructure for large-scale EO mapping requires investment in platform-independent standards, community-governed data repositories, and institutional commitments to long-term data stewardship. The convergence around COG and Zarr as open formats is encouraging, but formats alone do not ensure sustainability without the governance structures and funding models to maintain the repositories that host them.

[1]
J. Jakubik et al., Foundation Models for Generalist Geospatial Artificial Intelligence.” arXiv, 2023. doi: 10.48550/ARXIV.2310.18660.
[2]
Y. Cong et al., SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite Imagery,” in Advances in neural information processing systems, A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022. Available: https://openreview.net/forum?id=WBhqzpF6KYH
[3]
C. F. Brown et al., AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data,” arXiv preprint arXiv:2507.22291, 2025, Available: https://arxiv.org/abs/2507.22291
[4]
Z. Feng et al., TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis,” arXiv preprint arXiv:2506.20380, 2025, Available: https://arxiv.org/abs/2506.20380
[5]
Z. Xiong et al., Neural Plasticity-Inspired Foundation Model for Observing the Earth Crossing Modalities,” arXiv preprint arXiv:2403.15356, 2024.
[6]
G. Astruc, N. Gonthier, C. Mallet, and L. Landrieu, AnySat: An Earth Observation Model for Any Resolutions, Scales, and Modalities,” arXiv preprint arXiv:2412.14123, 2024.
[7]
A. Fuller, K. Millard, and J. R. Green, CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders,” in Thirty-seventh conference on neural information processing systems, 2023.
[8]
C. Aybar et al., The Missing Piece: Standardising for AI-ready Earth Observation Datasets,” in ICML 2025 workshop TerraBytes, 2025. Available: https://openreview.net/forum?id=HV6F0dsGLK
[9]
N. Tsendbazar, M. Herold, L. Li, A. Tarko, and M. Duerauer, Towards operational validation of annual global land cover maps,” Remote Sensing of Environment, vol. 266, p. 112686, 2021, doi: 10.1016/j.rse.2021.112686.