OFICIAL Hugging Face Blog

The OlmoEarth Platform: Geospatial inference at planetary scale

What happened
Based on Hugging Face Blog · Jul 28, 2026

AI2’s OlmoEarth Platform provides infrastructure to run large-scale geospatial AI models for environmental monitoring, addressing challenges in data access, processing, and failure recovery across distributed systems.

The OlmoEarth Platform: Geospatial inference at planetary scale
Hugging Face Blog — Hugging Face
Key points
·
Governments, NGOs, and other mission-driven organizations are already adapting OlmoEarth for applications including deforestation monitoring, food security, and wildfire risk.
·
At Ai2, we know how to train and release powerful open models, and for organizations with strong engineering teams, an open model is all they need to run with.
·
But most organizations in the environmental space – the ones best placed to apply these models – don't have the infrastructure or engineering teams that can manage the full lifecycle: labeling data, fine-tuning models, and running large-scale inference.
·
We’ve spent more than a decade operating platforms like Skylight and EarthRanger, software that users around the world rely on every day, so it has to work every day.
Key numbers
·
5 hours using 19,600 CPUs and 994 GPUs, achieving a 155× speedup over serial compute.

The OlmoEarth Platform, developed by AI2, supports large-scale geospatial inference using foundation models pretrained on approximately 10 terabytes of multimodal satellite data. It enables organizations such as governments and NGOs to apply these models for tasks like deforestation monitoring, food security, and wildfire risk assessment without requiring extensive internal infrastructure or engineering teams.

Satellite imagery inference operates at a vastly different scale than typical machine learning tasks, often involving terabytes of data across multiple spectral bands, sensor types, and geographic regions. The platform addresses these challenges by dividing inference jobs into three hardware-matched stages—data acquisition, preprocessing, and model execution—optimizing resource use and reducing costs to fractions of a penny per square kilometer.

To efficiently locate and fetch satellite data, the platform maintains its own metadata index, updated in real time via notifications or polling, to avoid overwhelming external services like ESA’s or Microsoft Planetary Computer’s STAC APIs. It selects the best imagery sources and retrieves only the necessary data using cloud-optimized formats such as COG or Zarr, enabling windowed reads without full scene downloads.

The platform is designed to handle failures automatically by dynamically provisioning tasks as reentrant and idempotent virtual machine runners. It supports parallel processing across thousands of CPUs and GPUs, as demonstrated by a recent wildfire-risk map of North America generated in 30.5 hours using 19,600 CPUs and 994 GPUs, achieving a 155× speedup over serial compute.

Original source → Deals on Clipraptor.com →