← All posts / Research

From Petabytes to the Pale Moon: NASA and IBM Open-Source a Foundation Model for Lunar Science

Trained on 17 years of Lunar Reconnaissance Orbiter imagery, the NASA-IBM Lunar Foundation Model maps craters, spots young volcanism, and estimates polar ice stability — and it's free on Hugging Face.

From Petabytes to the Pale Moon: NASA and IBM Open-Source a Foundation Model for Lunar Science

The Moon just got its own foundation model. On September 10, 2026, NASA and IBM Research released the NASA-IBM Lunar Foundation Model, one of the first open-source AI models built specifically for lunar science. Hosted publicly on Hugging Face with the complete codebase on GitHub, the model is designed to help researchers rapidly analyze the lunar surface at a scale no human team could match — mapping craters, identifying unusual volcanic features, and estimating where ice may be hiding near the lunar poles.

For an agency sitting on petabytes of planetary data, the release marks a deliberate shift in how NASA thinks about scientific discovery: not just collecting more data, but making the existing archive machine-searchable and scientifically actionable.

A Model Raised on 17 Years of LRO Data

The workhorse behind the model is NASA’s Lunar Reconnaissance Orbiter (LRO), which has been imaging the Moon since 2009. The LRO dataset is larger than the data produced by all other NASA planetary missions combined, capturing a nearly seamless, high-resolution mosaic of the entire lunar surface.

The foundation model was trained on roughly 2 million image tiles drawn from that archive, including:

  • More than 1 million high-resolution Narrow Angle Camera images at 1-meter resolution
  • Nearly 964,000 multispectral images at 100-meter resolution

That corpus was supplemented with imagery and terrain data from other missions: NASA’s GRAIL (Gravity Recovery and Interior Laboratory), NASA’s Lunar Prospector, and JAXA’s SELENE (Selenological and Engineering Explorer).

Because the model is pre-trained on this vast, largely unlabeled dataset, planetary scientists can fine-tune it for specific tasks — crater mapping, ice prospectivity, volcanic feature detection — using only small amounts of labeled data. That is the core pitch of the foundation-model approach: instead of building a bespoke algorithm from scratch for each question, researchers adapt one general-purpose model to many.

What It Can Actually Do

NASA’s announcement walks through three concrete capabilities, each grounded in a real research problem.

Crater mapping. Every crater is formed by an impact, making crater counts and measurements essential for dating the lunar surface and reconstructing solar system history. The model maps craters more efficiently than manual methods, letting scientists spend their time interpreting results rather than tracing outlines. In one striking demonstration, the model detected existing craters around Einstein crater and highlighted a newly formed impact crater from a SpaceX rocket body — a surface change that was deliberately excluded from its pre-training data.

Young volcanism. Although the Moon is no longer volcanically active, it once was. The model accelerates the identification of irregular mare patches — unusually young-looking volcanic structures that challenge established timelines for lunar cooling. Mapping them at scale could help piece together a more accurate picture of the Moon’s thermal evolution.

Polar ice stability. Perhaps most consequential for exploration: the model estimates where ice patches are likely to remain stable, on and below the surface. The Moon’s permanently shadowed regions stay cold enough to trap and preserve ice for up to billions of years. In NASA’s benchmarks, the NASA-IBM model preserved fine-scale patterns in reference ice-prospectivity maps near the lunar south pole — including Mons Mouton — where a competing ConvNeXt baseline washed them out. Across all evaluated tasks, the model matched or exceeded strong baselines, with its clearest advantage on polar ice estimation.

Open Science, By Design

The release goes well beyond a checkpoint file. Alongside the model, the team published:

  • Machine-learning-ready pre-training datasets and benchmark collections
  • Integration into TerraTorch, IBM’s open-source toolkit for geospatial foundation models
  • A companion paper hosted on Hugging Face for reproducibility

The science team spanned NASA’s Impact AI team at Marshall Space Flight Center, the Planetary Science Division at NASA Headquarters, Goddard Space Flight Center, and Ames Research Center, alongside IBM Research and academic partners.

The Lunar Foundation Model joins a growing NASA-IBM portfolio. The Prithvi family covers Earth observation — disaster monitoring, flood mapping, crop yield prediction, hurricane prediction. The Surya model targets heliophysics, predicting space weather phenomena like solar flares that can disrupt power grids and satellites. The lunar model extends that playbook from Earth orbit to another world.

Why It Matters

The timing is not accidental. With the Artemis program targeting sustained lunar presence, questions like “where is the ice?” and “is this landing site safe?” have moved from academic curiosity to operational planning. A freely available model that any researcher, agency, or company can fine-tune lowers the barrier to answering them — and turns NASA’s 17-year observational record into a shared computational resource rather than a dusty archive.

Kevin Murphy, NASA’s chief science data officer and acting chief data and AI officer, framed it plainly: “NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job. We also have to make data easier for scientists to explore and use. The NASA-IBM Lunar Foundation Model shows what’s possible when we bring AI to NASA’s petabytes of scientific data. That’s a real opportunity we see with AI: turning large-scale data into new discoveries.”

It’s also a quietly notable datapoint in the open-versus-closed model debate. While frontier labs increasingly guard their weights, NASA and IBM are shipping a domain-specific foundation model, training data, benchmarks, and tooling as a package — betting that reproducible, open science accelerates a field that no single institution can dominate anyway.

The gap between “we have the imagery” and “we understand the surface” is exactly where foundation models earn their keep. The Moon is now one of the best-documented bodies in the solar system. With this release, it becomes one of the most computationally accessible too.