← All posts / Research

Two Million Moon Tiles and 1,100 GPU-Hours: Inside NASA and IBM's Open-Source Lunar Foundation Model

NASA and IBM have open-sourced a multimodal lunar foundation model trained on SomBench, a 39TB corpus spanning 17 years of LRO observations — cutting polar ice-mapping error by up to 22% while beating ImageNet baselines with half the training data.

Two Million Moon Tiles and 1,100 GPU-Hours: Inside NASA and IBM's Open-Source Lunar Foundation Model

On September 10, 2026, IBM and NASA quietly shipped one of the most unusual AI releases of the year: the NASA-IBM Lunar Foundation Model, a multimodal, multi-resolution foundation model trained not on web text or code, but on roughly two million co-registered image tiles of the Moon’s surface. The weights are live on Hugging Face under an Apache-2.0 license, the fine-tuning code sits in a NASA-IMPACT GitHub repository, and the training corpus — SomBench, the first open, machine-learning-ready lunar dataset of its kind — is published alongside it under CC BY 4.0.

It is one of the first publicly available foundation models purpose-built for lunar science, and it arrives at a moment when NASA and its partners are actively planning a sustained human presence on the lunar surface. The bet here is straightforward: instead of hand-crafting a new algorithm for every scientific question, researchers anywhere can start from a shared pre-trained model and fine-tune it for their task with small amounts of labeled data.

Why the Moon needed a foundation model

The Lunar Reconnaissance Orbiter has been observing the Moon since 2009, and the data it has produced is, by NASA’s own accounting, larger than all other NASA planetary missions combined. The result is petabytes of imagery and geophysical maps — high-resolution camera shots at 1 meter per pixel from the Narrow Angle Camera, multispectral context at 100 meters per pixel from the Wide Angle Camera, plus terrain, thermal, gravity, and radar products from instruments like Diviner, LOLA, Mini-RF, GRAIL, and JAXA’s SELENE/Kaguya.

The problem is that none of this was machine-learning-ready. There was no publicly available unified dataset that brought multi-modal, multi-resolution lunar data into a common framework. Scientists studying the Moon had to either sift through maps and images by hand or train small, task-specific models from scratch — computationally expensive approaches that often lacked the accuracy needed to identify fine geographic features.

SomBench fills that gap. It aggregates more than 30 spatially aligned layers from nine instruments across four missions, organized into two tracks: a WAC LowRes track of 963,609 tiles at 100 meters per pixel totaling 38 TB, and a NAC HighRes track of 1,000,113 tiles at 1 meter per pixel totaling 1.4 TB. Splits are assigned at the Lunar Transverse Mercator zone level — tiles straddling two zones are dropped to prevent spatial leakage — and the test split is held out for downstream evaluation. The full dataset is hosted on AWS.

A ViT-B with lunar-specific engineering

Architecturally, the model is compact by frontier standards: a ViT-B encoder-decoder with 768 dimensions, 12 layers, and 12 attention heads, trained from scratch on SomBench. What makes it interesting is not raw scale but the domain-specific design decisions layered on top of the TerraMind masked-token pretraining recipe.

The first is how the team handles illumination. Lunar surface appearance is governed more by illumination geometry — sun angles, solar-frame anchors, tile footprint — than by intrinsic surface variation. The same crater can look radically different under different lighting, so the model tokenizes per-tile acquisition geometry as explicit encoder inputs rather than letting the network silently confuse shadow for terrain.

The second is multi-resolution training. NAC-anchored and WAC-anchored tiles train together in a single mixed-batch loop at native resolution, so one set of weights covers both resolution families across a 100× scale gap. A FlexiViT patch embedding lets the checkpoint be fine-tuned at other patch sizes without retraining the backbone, and modality-wise tokenization means modalities can be dropped or added at fine-tuning time — a practical concern when downstream researchers may only have a subset of the eleven pretraining modalities available.

Inputs span nine dense image-like layers plus two sequence-like context modalities: per-tile optical metadata with eight fields, and static-map context with 28 fields. Pretraining itself was refreshingly modest by 2026 standards: 16 H100 GPUs, 150,000 steps, a global batch size of 1,536 in bf16 — roughly 1,100 GPU-hours. Nine modality-specific VQ-VAE tokenizers with FSQ quantization, a DDPM decoder, and a cross-entropy objective over discrete token vocabularies complete the recipe.

The benchmarks: where it wins and where it merely ties

According to the NASA-IBM technical paper, the model exceeds widely used methods by up to 23% on key geographic feature identification. The headline numbers decompose into three tasks:

Polar ice prospectivity is the widest margin. The model reduced RMSE by up to 22% compared to SwinV2-B (ImageNet) when identifying areas with high potential for lunar ice — 0.0293 versus 0.0377 for the best baseline. This matters because permanently shadowed regions near the lunar poles are among the most difficult environments to observe directly, yet they may contain ice indicating water and oxygen — resources essential for a future Moon base and for producing rocket fuel for Mars missions.

Crater detection shows the efficiency story. At meter-scale resolution the model matches state-of-the-art baselines like SwinV2-B. But at context-scale resolution (~100 meters), it outperforms SwinV2-B by nearly 19% using just half the training data — pretrained variants at 50% of the training data already match or exceed baselines trained on the full set. Crater counts are essential for dating the lunar surface and reconstructing solar system history, and crater mapping directly feeds NASA’s landing-site selection, hazard avoidance (steep slopes, boulders), and planning for long-term lunar infrastructure.

Irregular Mare Patch segmentation is the narrowest result: an IoU of 0.5709 against 0.5687 for ConvNeXtV2-B — a 3% edge achieved with imperfect labels. The authors candidly note that leaders on the NAC crater and IMP benchmarks should be treated as comparable, since margins are smaller than the spread across random seeds. All benchmarks ran through TerraTorch with loaders, splits, augmentations, loss, and metrics held fixed across backbones, reported as mean and standard deviation over five seeds against baselines including ResNet-50, ViT-B MAE, ConvNeXt-B, DaViT-B, DeepLabV3+, and SegFormer.

A particularly compelling demonstration in NASA’s announcement: the model detected a newly formed impact crater near Einstein crater created by a SpaceX rocket body. Because the post-impact image was excluded from pretraining, the test shows the model can be fine-tuned to recognize novel surface changes between observations — automating change detection across vast lunar datasets.

Honest limitations, stated up front

The model card is unusually explicit about what this system is not. It is not a scientific-grade generative product; it maintains no geodetic reference frame; it is not validated for operational decisions such as landing-site certification or hazard clearance. Ice-prospectivity outputs regress a knowledge-driven fuzzy-overlay map, not measured ice. Ablation contributions are not yet isolated, and NAC pretraining was restricted to 1,095 co-registered frames. These are the right caveats for a research artifact, and they frame the release as infrastructure for the scientific community rather than a finished product.

The bigger picture: Prithvi’s lunar sibling

The Lunar Foundation Model is the newest member of a growing family of open models from the NASA-IBM collaboration: the Prithvi geospatial family (disaster monitoring, flood mapping, crop yield prediction, hurricane prediction) and the Surya heliophysics model for space weather prediction. Prithvi already became the first AI geospatial foundation model deployed in orbit earlier this year.

The pattern is consistent and deliberate — open weights, open data, open code, domain-adapted architectures, and modest compute budgets that any university lab can reproduce. As NASA Chief Science Data Officer Kevin Murphy put it: “We also have to make data easier for scientists to explore and use. The NASA-IBM Lunar Foundation Model shows what’s possible when we bring AI to NASA’s petabytes of scientific data. That’s a real opportunity we see with AI: turning large-scale data into new discoveries.”

While frontier labs race toward ever-larger general models, this release is a reminder that some of the most practical AI progress happens in narrow domains with well-curated data — and that 1,100 GPU-hours, spent wisely on two million Moon tiles, can outperform a decade of hand-drawn maps.