More research

Generalist Foundation Model for SAR-to-EO Image Translation

Jeonghyeok Do Munchurl Kim†

Korea Advanced Institute of Science and Technology (KAIST), South Korea

† Corresponding authorehwjdgur0913@kaist.ac.krmkimee@kaist.ac.kr

arXiv preprint, 2026

Abstract

Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing conditions.

We introduce GeoSET, the first generalist model for SET, built around a single pretrained parent that is adapted to downstream datasets under a common protocol. We curate over 3 million high-quality SAR–EO pairs from a collection of more than 10 million SAR observations, spanning diverse sensors, spatial resolutions, and ground sampling distances. To bridge the modality gap between SAR observations and a pretrained image generator, we develop a speckle-robust SAR encoder and pretrain the conditional generator on this heterogeneous corpus.

The resulting parent supports efficient adaptation across downstream datasets through low-rank adaptation (LoRA), updating only 0.60% of the generator parameters and requiring approximately one hour per dataset. Across six downstream benchmarks, GeoSET achieves state-of-the-art results in FID and DISTS with full fine-tuning or LoRA, demonstrating effective transfer across heterogeneous SAR–EO domains.

Pretrain once, adapt to many

One parent model, pretrained on 3M+ curated SAR–EO pairs, adapted to each benchmark with LoRA or full fine-tuning.

Grid with five benchmark rows (QXS-SAROPT, SAR2Opt, SAR2EO, SpaceNet6, CAP-BSG) and nine columns: input SAR, SD2.1 fine-tuning, BBDM, E3Diff, cBBDM, C-DiffSET, GeoSET with LoRA, GeoSET with full fine-tuning, and ground-truth EO.
Qualitative comparison on SAR-to-EO image translation benchmarks. Columns (g)–(h): GeoSET with LoRA and full fine-tuning; (i): ground-truth EO.
More scenes Fig. 5 · four benchmarks
Grid with four benchmark rows (QXS-SAROPT, SAR2Opt, SAR2EO, SpaceNet6) and nine columns: input SAR, SD2.1 FT, BBDM, E3Diff, cBBDM, C-DiffSET, GeoSET LoRA, GeoSET full FT, and ground-truth EO.
Qualitative comparison on QXS-SAROPT, SAR2Opt, SAR2EO and SpaceNet6. All methods receive the same SAR observation.

Quantitative results

Across full fine-tuning and LoRA, GeoSET achieves the best reported FID on all six benchmarks and the best DISTS on five.

Radar chart of FID and DISTS on QXS-SAROPT, SAR2Opt, SAR2EO, SpaceNet6, CAP-BSG and KOMPSAT, normalized per dataset and metric, for CondDiff, E3Diff, cBBDM, C-DiffSET, GeoSET (LoRA) and GeoSET (full FT). The two GeoSET curves are the outermost on most axes.
Cross-dataset comparison. FID and DISTS per benchmark, normalized as 100 × best/value (outer ring = best). FID and DISTS are the primary metrics.
Table 2. Quantitative comparison on QXS-SAROPT and SAR2Opt. We retrain (evaluate) all competing methods on the same training (test) splits.
Bold bestUnderline second best↓ lower is better · ↑ higher is better
Show
Scroll for more columns
Method Venue QXS-SAROPT SAR2Opt
FID↓ DISTS↓ KID↓ DINO↑ LPIPS↓ SSIM↑ PSNR↑ FID↓ DISTS↓ KID↓ DINO↑ LPIPS↓ SSIM↑ PSNR↑
General image-to-image translation methods
pix2pixCVPR'17174.60.3730.17130.2750.6650.20312.33261.90.3470.21640.2770.6570.19913.39
CycleGANICCV'17115.50.3620.08020.2830.6470.27813.25139.10.3230.03430.3740.6420.18812.68
pix2pixHDCVPR'1885.70.2980.04920.4030.5730.35816.13146.30.2830.06540.4750.5670.26815.95
SPADECVPR'1990.70.2920.06070.3660.5990.32014.53142.50.2650.05180.4470.5970.23414.47
DDPM (SR3)TPAMI'2243.80.3110.01890.4250.6200.35914.04122.50.2950.04370.4970.6100.31313.65
SD2.1 FTCVPR'2219.10.2570.00420.4890.5610.34815.4071.80.2110.00940.6000.5410.29316.24
BBDMCVPR'2376.60.2700.04790.4140.5680.35215.34143.10.2900.06710.4660.5900.27615.29
ControlNetICCV'2350.40.3070.02110.4580.6040.29713.42140.50.3500.04800.4790.6430.21711.73
HI-DiffNeurIPS'23324.30.5390.32690.2150.6920.45717.10319.80.4730.23570.2770.6920.38417.36
ResShiftNeurIPS'23140.20.3340.08720.2950.6070.21714.20141.70.3040.05150.4350.5970.17714.31
StegoGANCVPR'24106.80.3840.07070.2610.6580.25412.96149.80.3320.03960.3620.6520.16212.39
SAR-to-EO image translation (SET) methods
CondDiffGRSL'2388.60.3550.05370.3100.7300.21311.55211.80.4150.13790.3430.6860.24812.48
E3DiffGRSL'2447.80.2780.01670.3790.5300.30216.44104.70.2320.03060.5410.5290.24916.09
cBBDMGRSL'2550.60.2460.02840.4920.5390.37216.02222.30.3770.15210.4130.5710.36117.05
C-DiffSETTCSVT'2619.90.2330.00550.5220.5260.38016.9278.10.2140.01380.6010.5290.31416.81
GeoSET (LoRA)–19.30.2540.00500.5440.5640.32614.6774.60.2000.00660.6260.5350.27815.50
GeoSET (full FT)–16.90.2440.00320.5580.5530.33215.0771.10.1960.00520.6140.5320.27915.66
Pretraining corpus Table 1
Table 1. Stage 2 SAR–EO pretraining corpus. Drop rates are computed relative to the original SAR sample counts and include source-specific pair eligibility selection and quality filtering. Crop-equivalent counts account for native image size and redundancy and determine the source sampling mixture.
Scroll for more columns
Source Orig. samples Kept pairs Drop (%) Crop equivalents Mix (%)
GUSO589,143585,4240.632,341,69638.73
TerraMesh8,194,0481,741,43478.751,741,43428.80
SARLO-8087,87077,93211.311,246,91220.62
SAR-1M1,130,379688,45839.09688,45811.39
3MOS113,074111,4961.4027,8740.46
Total10,114,5143,204,74468.326,046,374100.00
Ablations and analysis Tables 5, 4 and 7
Table 5. Cumulative component comparison on SAR2EO and SpaceNet6.
Bold best↓ lower is better
Scroll for more columns
Configuration SAR2EO SpaceNet6
FID↓ DISTS↓ LPIPS↓ FID↓ DISTS↓ LPIPS↓
SD2.1 Full FT67.70.2700.487124.70.1960.348
FLUX.2 LoRA (Backbone)71.20.2750.481117.40.2010.361
+ Pretraining corpus50.50.2500.470107.90.1940.351
+ SAR encoder39.80.2200.44894.20.1930.347
+ Speckle aug. (GeoSET)29.50.2020.44087.90.1890.338

Each step improves FID, DISTS and LPIPS on both datasets (SAR2EO FID 71.2 → 50.5 → 39.8 → 29.5). The last row is the complete GeoSET (LoRA) configuration, as in Table 3.

Table 4. Stage 1 encoder diagnostics and Stage 3 adaptation efficiency. (a) Reconstruction PSNR on original and perturbed SAR inputs before and after encoder training. EO reports reconstruction with the frozen pretrained codec. (b) Computational costs of full fine-tuning and LoRA.
Scroll for more columns

(a) Stage 1 Encoder Reconstruction

Source Original SAR Perturbed SAR EO
Before After Δ Before After Δ
GUSO27.4729.18+1.7123.5426.96+3.4231.02
TerraMesh20.1132.04+11.9319.1230.48+11.3733.95
SARLO-8017.3919.21+1.8216.7118.49+1.7827.81
SAR-1M29.9531.25+1.3022.0126.99+4.9833.59
3MOS35.1536.07+0.9227.3031.92+4.6134.62
Average26.0129.55+3.5421.7426.97+5.2332.20

(b) Downstream Adaptation Efficiency

GeoSET Full FT LoRA
Trainable parameters3.852B23.10M
Trainable fraction100%0.5997%
Per-dataset storage30.8 GB185 MB
Peak memory96.9 GB49.1 GB
Batch size1616
Throughput67 img/s94 img/s
Wall time/dataset1.34 h0.98 h

(a) Encoder training raises average SAR PSNR by +3.54 dB on original and +5.23 dB on perturbed inputs. (b) LoRA trains 23.10M parameters (0.5997%), stores 185 MB instead of 30.8 GB and takes 0.98 h per dataset.

Reconstructions from the adapted SAR encoder and the frozen EO encoder on six pretraining streams (GUSO, TerraMesh, SARLO-80, SAR-1M, 3MOS-MR, 3MOS-HR). Top row: SAR and EO reconstructions with PSNR values. Bottom row: speckle-perturbed SAR input at L = 4 and its reconstruction, with PSNR values.
SAR and EO reconstruction across pretraining sources. Top: reconstructions of original SAR and EO; bottom: speckle-perturbed SAR (L = 4, left) and its reconstruction (right). Values are PSNR (dB) relative to the original observations; values in parentheses are for the perturbed inputs.
Table 7. SAR conditioning intervention on pretraining sources. True denotes the corresponding SAR observation, shuffled uses a mismatched SAR observation, and null uses an all-zero SAR latent. The null condition is included during training through conditioning dropout.
Bold best↓ lower is better · ↑ higher is better
Scroll for more columns
PSNR↑ LPIPS↓
Source True Shuffled Null True Shuffled Null
GUSO13.3610.2410.580.5480.6780.690
SARLO-8015.8811.9611.650.5140.6780.693
TerraMesh13.6910.9410.800.4910.6700.695

True SAR conditioning is best on PSNR and LPIPS for all three pretraining sources.

Three stages, one parent

GeoSET framework. Top: multi-source SAR-EO pretraining corpus (GUSO, TerraMesh, SARLO-80, SAR-1M, 3MOS) and the downstream SET benchmarks. Stage 1: speckle-robust SAR representation pretraining with a frozen EO decoder. Stage 2: generalist SAR-to-EO pretraining of a DiT with flow matching, conditioned on the SAR latent. Stage 3: downstream adaptation with LoRA or full fine-tuning.
GeoSET framework. Stages 1–2 build one generalist parent; Stage 3 adapts it to each benchmark.
  1. Stage 1Speckle-robust SAR encoder. Reconstructs SAR from speckle-perturbed input through a frozen decoder.
  2. Stage 2Generalist FLUX.2 pretraining. The generator, conditioned on SAR instead of text, is pretrained on over 3M curated SAR–EO pairs.
  3. Stage 3Per-benchmark adaptation. LoRA (0.60% of generator parameters, ≈1 h per dataset) or full fine-tuning.

BibTeX

@article{do2026geoset,
  title={GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation},
  author={Do, Jeonghyeok and Kim, Munchurl},
  journal={arXiv preprint arXiv:XXXX.XXXXX},
  year={2026}
}
Prior work: C-DiffSET (IEEE TCSVT 2026)
@article{do2026cdiffset,
  title={C-diffset: Leveraging latent diffusion for sar-to-eo image translation with confidence-guided reliable object generation},
  author={Do, Jeonghyeok and Lee, Jaehyup and Lee, Seungchul and Kim, Munchurl},
  journal={IEEE Transactions on Circuits and Systems for Video Technology},
  year={2026},
  publisher={IEEE}
}