More research

Leveraging Latent Diffusion for SAR-to-EO Image Translation with Confidence-Guided Reliable Object Generation

Jeonghyeok Do1 Jaehyup Lee2 Seungchul Lee3 Munchurl Kim1,†

1Korea Advanced Institute of Science and Technology (KAIST), South Korea2Kyungpook National University, South Korea3Stellarvision Inc.

† Corresponding authorehwjdgur0913@kaist.ac.krjaehyuplee@knu.ac.krleesc@stellarvision.krmkimee@kaist.ac.kr

IEEE TCSVT 2026

Abstract

Synthetic Aperture Radar (SAR) imagery provides robust environmental and temporal coverage (e.g., during clouds, seasons, day-night cycles), yet its noise and unique structural patterns pose interpretation challenges, especially for non-experts. SAR-to-EO (Electro-Optical) image translation (SET) has emerged to make SAR images more perceptually interpretable. However, traditional approaches trained from scratch on limited SAR-EO datasets are prone to overfitting.

To address these challenges, we introduce Confidence Diffusion for SAR-to-EO Translation, called C-DiffSET, a framework leveraging pretrained Latent Diffusion Model (LDM) extensively trained on natural images, thus enabling effective adaptation to the EO domain. Remarkably, we find that the pretrained VAE encoder aligns SAR and EO images in the same latent space, even with varying noise levels in SAR inputs. To further improve pixel-wise fidelity for SET, we propose a confidence-guided diffusion (C-Diff) loss that mitigates artifacts from temporal discrepancies, such as appearing or disappearing objects, thereby enhancing structural accuracy. C-DiffSET achieves state-of-the-art (SOTA) results on multiple datasets, significantly outperforming the very recent image-to-image translation methods and SET methods with large margins.

Start from pretrained diffusion, not from scratch

A latent diffusion model pretrained on natural images, fine-tuned for SAR-to-EO translation with a confidence-guided loss.

Grid with three rows, SpaceNet6 (full polarization), SAR2Opt and QXS-SAROPT (single polarization), and six columns: input SAR, StegoGAN, BBDM, ControlNet, C-DiffSET and the target EO image.
Qualitative comparison on SpaceNet6, SAR2Opt and QXS-SAROPT. Full-polarization SAR input (SpaceNet6, GSD 0.5 m) and single-polarization input (SAR2Opt and QXS-SAROPT, GSD 1 m), against StegoGAN, BBDM and ControlNet.

Quantitative results

C-DiffSET is best in all five metrics on SAR2Opt, SpaceNet6 and QXS-SAROPT.

Table 1. Quantitative comparison of image-to-image translation methods and SET methods on SAR2Opt and SpaceNet6 datasets.
Bold best↓ lower is better · ↑ higher is better
Show
Scroll for more columns
Method Venue SAR2Opt SpaceNet6
FID↓ LPIPS↓ SCC↑ SSIM↑ PSNR↑ FID↓ LPIPS↓ SCC↑ SSIM↑ PSNR↑
GANs
Pix2PixCVPR 2017196.870.4260.00060.21615.422124.550.2560.01020.52219.357
CycleGANICCV 2017139.720.4250.00220.22414.931114.810.2740.00970.49317.798
SAR-SMTNet†TGRS 2023160.870.4790.00110.21914.661118.960.2940.01030.48317.032
CFCA-SET†TGRS 2023152.270.4300.00090.22315.183164.780.2790.00970.49818.297
StegoGANCVPR 2024144.540.3980.00340.23715.62475.120.2440.01060.51618.958
LDMs
BBDMCVPR 202394.720.4730.00050.23415.13181.860.3020.00190.21717.678
ControlNetICCV 202381.040.4230.00050.21614.461106.590.3920.00270.17814.085
Uni-ControlNetNeurIPS 202380.810.4210.00040.21514.38491.140.3210.00370.18314.333
DGDMECCV 2024156.120.5410.00040.27315.568238.370.4380.00150.25317.124
cBBDM†arXiv 202497.640.3940.00220.28516.59172.770.2430.00790.25419.033
C-DiffSET (Ours)–77.810.3460.00350.28616.61337.440.1420.01510.56721.022

† SET-specific methods without official code, re-implemented from their technical descriptions. All LDM-based methods, including C-DiffSET, are initialised with the same Stable Diffusion v2.1 weights. On SpaceNet6, C-DiffSET lowers FID from 72.77 (cBBDM, the next best) to 37.44.

Ablation: pretrained LDM and C-Diff loss Tables 3 and 8
Table 3. Ablation studies on the SAR2Opt and SpaceNet6 dataset evaluating the impact of pretrained LDM and confidence-guided diffusion (C-Diff) loss.
Bold best↓ lower is better · ↑ higher is better
Scroll for more columns
Pretrained LDM Loss function SAR2Opt SpaceNet6
FID↓ LPIPS↓ SCC↑ SSIM↑ PSNR↑ FID↓ LPIPS↓ SCC↑ SSIM↑ PSNR↑
MSE98.980.390.0010.2615.8260.260.230.0110.4318.16
✓MSE78.140.360.0030.2816.4640.620.160.0140.5220.29
✓C-Diff77.810.340.0040.2916.6137.440.140.0150.5721.02

Starting from the pretrained LDM lowers FID from 98.98 to 78.14 on SAR2Opt and from 60.26 to 40.62 on SpaceNet6 with the same MSE loss; the C-Diff loss then improves every metric on both datasets.

Table 8. Ablation studies on the QXS-SAROPT dataset evaluating the impact of pretrained LDM and confidence-guided diffusion (C-Diff) loss.
Bold best↓ lower is better · ↑ higher is better
Scroll for more columns
Pretrained LDM Loss function QXS-SAROPT
FID↓ LPIPS↓ SCC↑ SSIM↑ PSNR↑
MSE29.040.4070.00060.27914.647
✓MSE19.990.2970.00940.36417.736
✓C-Diff18.150.2930.01080.37218.077

On QXS-SAROPT, starting from the pretrained LDM lowers FID from 29.04 to 19.99 with the same MSE loss; the C-Diff loss then improves every metric.

Computational cost Table 7
Table 7. Comparative analysis of C-DiffSET with other methods by parameters, FLOPs, memory usage, and inference time.
Scroll for more columns
Method Params. (M) FLOPs (G) Memory (MB) Time (s)
GANs
Pix2Pix54.4124.22464.120.06
CycleGAN7.84140.43398.380.08
SAR-SMTNet2.15615.402626.980.22
CFCA-SET26.8098.98431.860.10
StegoGAN13.15227.49461.140.11
LDMs
BBDM949.562122.446147.883.13
ControlNet1312.722231.037567.834.50
Uni-ControlNet1519.182295.208382.004.95
DGDM959.032161.256184.451.47
cBBDM949.582122.496147.933.15
C-DiffSET949.582122.496148.093.27

C-DiffSET has the parameters and FLOPs of cBBDM (949.58 M and 2122.49 G) and takes 3.27 s per 512 × 512 image.

Share the latent, weight by confidence

C-DiffSET framework. Training: the frozen image encoder maps the SAR image X and the EO image Y to latents; noise is added to the EO latent in the forward process; the denoising U-Net, fed the concatenated noisy EO and SAR latents, the timestep and the CLIP embedding of the prompt Electro-Optical Image, predicts the noise and a confidence map, optimised with the C-Diff loss. Inference: starting from Gaussian noise, the U-Net denoises step by step conditioned on the SAR latent, and the image decoder produces the EO prediction.
C-DiffSET framework. Training (left): the U-Net predicts the noise and a confidence map from the noisy EO latent and the SAR latent. Inference (right): iterative denoising from Gaussian noise, then the VAE decoder.
  1. 01One latent space for SAR and EO. The frozen VAE of the pretrained LDM embeds both images; the SAR latent is concatenated channel-wise with the noisy EO latent, keeping pixel-wise correspondence.
  2. 02Fine-tuned, not trained from scratch. The U-Net starts from Stable Diffusion v2.1 weights, with the fixed prompt “electro-optical image” as a stable conditioning signal.
  3. 03Confidence-guided diffusion loss. A predicted confidence map weights the noise error pixel-wise (β-NLL style), so regions where objects appear or disappear between acquisitions are down-weighted.

BibTeX

@article{do2026cdiffset,
  title={C-diffset: Leveraging latent diffusion for sar-to-eo image translation with confidence-guided reliable object generation},
  author={Do, Jeonghyeok and Lee, Jaehyup and Lee, Seungchul and Kim, Munchurl},
  journal={IEEE Transactions on Circuits and Systems for Video Technology},
  year={2026},
  publisher={IEEE}
}