More research

Towards Geo-Aware Generative Foundation Models in Earth Observation

Jeonghyeok Do Munchurl Kim†

Korea Advanced Institute of Science and Technology (KAIST), South Korea

† Corresponding authorehwjdgur0913@kaist.ac.krmkimee@kaist.ac.kr

NeurIPS 2026

Abstract

Existing generative models for earth observation (EO) predominantly rely on fine-tuning natural image priors, which limits their scalability and introduces perspective biases that conflict with geospatial constraints. To address this, we introduce GeoCore-9B, a 9-billion-parameter generative foundation model, which is the first of its scale to be trained from scratch exclusively on EO data. Unlike previous EO foundation models, GeoCore-9B is built upon a Flow Matching-based Diffusion Transformer (DiT) and natively conditions generation on text descriptions and continuous geospatial metadata, including ground sample distances, latitudes, and longitudes.

To overcome the convergence and spatial disorientation challenges of training at this scale, we propose a Geospatial Semantic Alignment loss. This objective distills structural Earth surface priors (e.g., terrain and urban areas) from a frozen specialist teacher network, constraining the diffusion latent trajectory during training without adding inference overhead.

Pre-trained on the global-scale Git-10M dataset, GeoCore-9B demonstrates strong downstream versatility. Beyond standard proxy generative tasks, we show that GeoCore-9B can be effectively adapted for practical EO applications, including highly challenging tasks such as cloud removal and SAR-to-optical cross-modal translation. Extensive evaluations confirm that GeoCore-9B establishes new state-of-the-art performance in both visual fidelity and geographic structural accuracy.

Geo-aware generation at 9B scale

One Flow Matching DiT, trained from scratch on EO data, generates satellite imagery from text, GSD and coordinates.

Ten text prompts in a five-by-two grid, each with images from CRS-Diff, Text2Earth and GeoCore-9B: two tennis courts surrounded by trees, houses on both sides of a highway, houses by the beach, an industrial area with blue workshops, a commercial area, an intersection in a residential area, a curved road through a dense forest, white rocks and brown earth on a mountain, a parking lot next to trees, and a roundabout surrounded by grass.
Text-conditioned generation. Prompts vary while the geospatial metadata stays fixed; CRS-Diff and Text2Earth often show structural artifacts or unnatural textures.
More text prompts Fig. 9 · six prompts
Six more text prompts, each with images from CRS-Diff, Text2Earth and GeoCore-9B: circular farmlands, a sparse residential area on a meadow, a dense commercial area, storage tanks beside a meadow, ice floes on the sea, and a mobile home park.
More text-conditioned generations. Circular farmlands, a sparse residential area, a dense commercial area, storage tanks, ice floes and a mobile home park.

Adapted to cloud removal and SAR-to-optical

GeoCore-9B adapted with LoRA, next to task-specific baselines on the same Sen2-MTC and QXS-SAROPT inputs, six scenes per task. Select a tile to compare it with the ground truth.

Input GeoCore-9B (ours) Ground truth Task-specific baseline

Quantitative results

GeoCore-9B leads all three RSICD metrics after LoRA text-to-image adaptation, with the best PSNR and SSIM for cloud removal and the best FID for SAR-to-optical translation.

Table 1. Comparison with previous text-to-image methods on the RSICD dataset. Bold indicates the best result.
Bold best↓ lower is better · ↑ higher is better
Scroll for more columns
Method IS↑ FID↓ CLIP↑
Attn-GAN11.7195.8120.19
DAE-GAN7.7193.1519.69
StrucGAN5.84––
DF-GAN9.51109.4119.76
Lafite10.7074.1122.52
DALL-E2.59191.9320.13
Txt2Img-MHN5.99102.4420.27
RSDiff7.2266.49–
CRS-Diff18.3950.7220.33
Text2Earth–24.4925.62
GeoCore-9B (Ours)22.1518.8227.15

GeoCore-9B is adapted to RSICD with LoRA; RSICD has no GSD or coordinates, so learned null geospatial embeddings are used in fine-tuning and inference.

Geospatial Semantic Alignment ablation Figs. 5–6, Table 4
Six prompts generated by a model trained without the GSA loss (top row) and with it (bottom row): circular farmlands, a dense commercial area, houses by the beach, an industrial area with blue workshops, a parking lot and a roundabout. Without GSA the layouts are fragmented and the boundaries distorted.
Without and with the GSA loss. Top: μ = 0, with fragmented textures and distorted boundaries; bottom: with GSA, cleaner layouts and sharper object boundaries.
Line plot of FID on a 10K subset of Git-10M against training iterations. With the GSA loss, FID is lower at 50K, 150K and 300K; the run without GSA is also evaluated at 400K.
Faster convergence. FID on a 10K Git-10M subset across training iterations.
Table 4. Matched downstream comparison with and without GSA. All settings other than the pre-training GSA weight are held fixed. HF-SCC uses the corrected, baseline-consistent definition.
Bold best
Scroll for more columns
Task w/ GSA (μ = 0.5) w/o GSA (μ = 0)
RSICD text-to-image
IS22.1519.16
FID18.8228.43
CLIP27.1524.21
QXS-SAROPT translation
FID12.0519.92
LPIPS0.3770.436
HF-SCC0.01630.0098
SSIM0.3700.324
Sen2-MTC cloud removal
PSNR20.80919.553
SSIM0.7990.683
LPIPS0.2560.284

GSA improves every reported metric: FID falls by 9.61 points on RSICD and 7.87 on QXS-SAROPT, and Sen2-MTC PSNR rises by 1.256 dB.

Frozen VAE on EO data Table 3, Fig. 8
Table 3. Frozen-VAE encode–decode fidelity on the pre-training and downstream domains. SAR intensities are replicated from one channel to three channels before VAE encoding.
↓ lower is better · ↑ higher is better
Scroll for more columns
Domain n PSNR↑ SSIM↑ LPIPS↓ HF-SCC↑
Git-10M RGB100,00031.480.931–0.710
QXS-SAROPT optical2,00038.930.9680.0110.766
QXS-SAROPT SAR (1ch → 3ch)2,00028.770.9320.0230.782
Sen2-MTC cloudy68735.480.9430.0130.416
Sen2-MTC cloud-free68735.580.9240.0160.564

HF-SCC from 0.416 to 0.782 on the downstream domains supports the frozen VAE for the evaluated 256 × 256 tasks; the lower Sen2-MTC values expose a domain-dependent limitation.

Six Git-10M satellite images (top row) and their reconstructions by the frozen pretrained VAE (bottom row): farmland, fields along a river, greenhouses, a residential grid, an orchard by a road, and a roundabout next to a large roof.
VAE reconstruction. Top: original Git-10M images; bottom: reconstructions by the frozen pretrained VAE, without EO-specific fine-tuning.
Metadata interventions and feature probes Tables 5–6
Table 5. Paired metadata interventions. “Orig.” and “suppl.” score a shuffled-condition output against its original and newly supplied metadata, respectively. FID values support comparisons only within this protocol.
↓ lower is better · ↑ higher is better
Scroll for more columns
Condition FID↓ GSD-bin acc.↑ 15°-region acc.↑
Full metadata48.320.5550.485
GSD null54.960.2930.460
GSD shuffled49.460.286 (orig.) / 0.376 (suppl.)0.424
Coordinates null68.180.2610.175
Coordinates shuffled51.200.4920.110 (orig.) / 0.366 (suppl.)

1,000 Git-10M samples with fixed caption, seed and sampler. Nulling the coordinates lowers region accuracy from 0.485 to 0.175; with shuffled coordinates, outputs agree more with the supplied regions than with the original ones (0.366 versus 0.110). In a global-feature near-duplicate test against all 10,503,567 pre-training images, none of 500 coordinate-only generations exceeds the threshold.

Table 6. Frozen-feature linear probes on EuroSAT, LoveDA, and BRIGHT. The train/validation sizes are shown in parentheses.
Bold best
Scroll for more columns
Task (train/val) Metric τ = .25 τ = .50 τ = .75 DINOv3-Sat
EuroSAT (12,960/3,240)Top-197.3%97.5%97.3%98.0%
LoveDA (3,000/1,200)mIoU0.3280.3820.2820.451
BRIGHT (2,500/349 pairs)mIoU0.4020.4820.3810.442

One linear layer on frozen DiT features. At τ = 0.50, GeoCore-9B is within 0.5 top-1 points of DINOv3-Sat on EuroSAT and exceeds it on BRIGHT (0.482 versus 0.442); DINOv3-Sat is also the GSA teacher.

Geo-aware conditioning, training-only alignment

GeoCore-9B framework. Training: an EO image is encoded by a frozen VAE encoder and noised along a linear trajectory; a DiT predicts the velocity under the Flow Matching loss. Local text tokens from a frozen text encoder join the image tokens; a global text embedding, the timestep and the metadata (GSD, latitude, longitude) are summed into one conditioning vector for AdaLN modulation. A frozen geospatial feature extractor (teacher) supervises the k-th DiT block through a projection head and the Geospatial Semantic Alignment loss. Inference: an ODE solver and the frozen VAE decoder.
GeoCore-9B framework. A Flow Matching DiT conditioned on text tokens and geospatial metadata; the teacher and projection head are used in training only.
  1. 01A 9B Flow Matching DiT, from scratch. 32 DiT blocks in the latent space of a pretrained VAE, trained on Git-10M.
  2. 02Text and geospatial metadata. T5-XXL tokens join the image tokens; CLIP text, GSD, latitude and longitude embeddings modulate the DiT through AdaLN.
  3. 03Geospatial Semantic Alignment. A frozen DINOv3-Sat teacher supervises block-8 features during training only, with no inference overhead.

BibTeX

@inproceedings{do2026geocore,
  title={GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation},
  author={Do, Jeonghyeok and Kim, Munchurl},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2026}
}