Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation


Jeonghyeok Do1, Seungchul Lee2, Munchurl Kim1*
Corresponding author
1KAIST    2Stellarvision Inc.
arXiv preprint, 2026

Abstract

Synthetic aperture radar images in all weather and at all hours, but its speckle and non-optical geometry make it hard to read, which motivates SAR-to-EO translation. Recent SAR-to-EO methods fine-tune a pretrained latent diffusion model and are therefore bound to that model's autoencoder, whose reconstruction quality is an upper bound no generator built on top of it can exceed.

We present ReFlowSET, which places a conditional flow-matching transformer directly in a frozen high-fidelity autoencoder latent space. A 509M-parameter DiT with an eight-block two-stream SAR/EO front end and sixteen single-stream blocks is trained from scratch to predict the velocity of a linear bridge from noise to the EO latent, conditioned on the SAR latent, with a representation-alignment loss to a frozen vision-transformer teacher that is discarded at inference.

On QXS-SAROPT and SAR2Opt, scored on items identical to fifteen prior methods that we retrained under one protocol and one evaluator, ReFlowSET obtains the best DISTS on both datasets and the best FID and LPIPS on SAR2Opt, and exposes a four-step operating point that samples about 11× faster than its 50-step setting.

QXS-SAROPT

Drag the handle. Top: SAR input → ReFlowSET. Bottom: C-DiffSET → ReFlowSET.

SAR2Opt

Top: SAR input → ReFlowSET. Bottom: C-DiffSET → ReFlowSET.

All methods, same scenes — QXS-SAROPT

Every method was retrained by us and evaluated by one evaluator on the same test items.

All methods, same scenes — SAR2Opt

Quantitative comparison

Bold = best, underline = second best, within each column. All rows were retrained by us and scored by a single evaluator on identical test items (QXS-SAROPT n=3999 at 256×256, SAR2Opt n=627 at 512×512). These numbers are not comparable with the ones printed in the source papers, which use different splits and protocols.

LPIPS convention. We feed x*2−1 to the LPIPS network. Several released evaluators feed [0,1] unscaled and report a systematically lower number — the two differ by about 0.05, so check which one any published LPIPS uses before comparing it with this column.
Two rows are named for a paper whose code we did not run. DDPM (SR3-class) is E3Diff's stage 1, and SD2.1 fine-tune only is C-DiffSET's stage 1 without the confidence channel.

MethodVenueQXS-SAROPTSAR2Opt
FID↓DISTS↓LPIPS↓SSIM↑PSNR↑FID↓DISTS↓LPIPS↓SSIM↑PSNR↑
General image-to-image methods
pix2pixCVPR'17174.60.3730.6650.20312.33261.90.3470.6570.19913.39
CycleGANICCV'17104.40.3760.6530.26212.92143.50.3300.6500.17812.90
pix2pixHDCVPR'1885.70.2980.5730.35816.13146.30.2830.5670.26815.95
SPADECVPR'1990.70.2920.5990.32014.53142.50.2650.5970.23414.47
DDPM (SR3-class)TPAMI'2243.80.3110.6200.35914.04122.50.2950.6100.31313.65
SD2.1 fine-tune onlyCVPR'2219.10.2570.5610.34815.4071.80.2110.5410.29316.24
BBDMCVPR'2376.60.2700.5680.35215.34143.10.2900.5900.27615.29
ControlNetICCV'2350.40.3070.6040.29713.42140.50.3500.6430.21711.73
HI-DiffNeurIPS'23324.30.5390.6920.45717.10319.80.4730.6920.38417.36
ResShiftNeurIPS'23140.20.3340.6070.21714.20141.70.3040.5970.17714.31
StegoGANCVPR'24106.80.3840.6580.25412.96150.10.3470.6550.15812.47
SAR-to-EO methods
Conditional DiffusionGRSL'2388.60.3550.7300.21311.55211.80.4150.6860.24812.48
cBBDMGRSL'2550.60.2460.5390.37216.02222.30.3770.5710.36117.05
E3DiffGRSL'2447.80.2780.5300.30216.44104.70.2320.5290.24916.09
C-DiffSETTCSVT'2619.90.2330.5260.38016.9278.10.2140.5290.31416.81
Ours
ReFlowSET (Ours)19.10.2310.5340.35516.0966.30.1850.5220.28716.06

Identity collapse. For these rows the released implementation, run at its own published protocol, produces an output measurably close to a copy of its SAR input (mean|generated−SAR| / mean|generated−GT| below 1). They are reported in place rather than removed or substituted: what a reader is owed is what the released implementation actually does. CycleGAN on QXS-SAROPT and SAR2Opt; StegoGAN on QXS-SAROPT and SAR2Opt. Every other row, ReFlowSET included, is well clear of that threshold (ReFlowSET QXS-SAROPT 2.47, SAR2Opt 1.91).

Two SAR2Opt cells — C-DiffSET and SD2.1 fine-tune only — had their predictions re-dumped after the only extended-metric pass that scored them, so their DISTS had gone stale. Both were re-measured at full n=627 and the values above are the re-measurement.

BibTeX

@article{do2026reflowset,
  title   = {ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation},
  author  = {Do, Jeonghyeok and Lee, Seungchul and Kim, Munchurl},
  journal = {arXiv preprint arXiv:{{ARXIV_ID}}},
  year    = {2026}
}

Our earlier SAR-to-EO work, which ReFlowSET builds on and compares against:

@article{do2026cdiffset,
  title   = {C-DiffSET: Leveraging Latent Diffusion for SAR-to-EO Image Translation with Confidence-Guided Reliable Object Generation},
  author  = {Do, Jeonghyeok and Lee, Jaehyup and Lee, Seungchul and Kim, Munchurl},
  journal = {IEEE Transactions on Circuits and Systems for Video Technology},
  year    = {2026}
}

Acknowledgements

This work was supported in part by the National Research Foundation of Korea (NRF) grant funded by the Korean government (MSIT) under the Sejong Science Fellowship Program (RS-2026-25484549, “Generative AI-based High-Resolution SAR Image Visualization and Analysis Technology for All-Weather Earth Observation”, 50%) and in part by the NRF grant funded by the MSIT (RS-2025-02222525, “Development of AI-based SAR-to-EO image conversion technology”, 50%).

The frozen autoencoder is the FLUX.2 autoencoder of Black Forest Labs, used unmodified. The released checkpoints ship the Apache-2.0 serialisation from FLUX.2-klein-base-4B; training and evaluation used the FLUX.2-dev serialisation of the same network. Substituting one for the other moves QXS-SAROPT PSNR by less than 0.004 dB in absolute value — two orders of magnitude less than changing the evaluation seed, which moves it by 0.395 dB. The representation-alignment teacher used during training is DINOv3 (Meta Platforms); it is required only for training, is not redistributed here, and is subject to its own licence. We thank the authors of the QXS-SAROPT and SAR2Opt datasets and of every method we retrained for comparison.