More research

Decoupled Semantic Supervision for SAR-EO Foundation Models

Jeonghyeok Do1 Munchurl Kim1,†

1Korea Advanced Institute of Science and Technology (KAIST), South Korea

† Corresponding authorehwjdgur0913@kaist.ac.krmkimee@kaist.ac.kr

arXiv preprint, 2026

Abstract

Synthetic aperture radar (SAR) and electro-optical (EO) imagery provide complementary observations: SAR enables day-and-night, weather-resilient sensing, whereas EO provides rich appearance and fine-grained semantic cues. We introduce SAREO-FM, which avoids forcing a single token stream to serve two distinct roles: modality tokens preserve how each sensor observes the scene through masked reconstruction, while learnable semantic queries capture what the scene contains under guidance from a pretrained vision foundation model (VFM). By jointly encoding these queries with SAR and EO tokens, the queries acquire modality-grounded semantic context, while the modality-token outputs remain the explicit targets of masked reconstruction. This design assigns semantic and reconstruction supervision to separate token streams while preserving their interaction within the shared encoder. Pretrained on the million-scale SAR-1M corpus, SAREO-FM achieves strong unimodal transfer for both SAR-only and EO-only inputs, while delivering substantial gains from joint SAR–EO observations on tasks that benefit from complementary sensing.

Semantics to queries, pixels to tokens

Semantic queries take the VFM supervision, SAR and EO tokens take masked reconstruction, all in one shared encoder.

Two panels. Previous work: masked SAR tokens pass through a SAR encoder whose output tokens are both decoded by a SAR decoder for the pixel loss and aligned with VFM features of the EO image, one stream with two duties (coupled); raw EO is not used as an input. SAREO-FM (ours): masked SAR tokens, masked raw EO tokens and learnable semantic queries enter one SAREO encoder; the SAR and EO output tokens go to separate SAR and EO decoders for the pixel losses, while only the query outputs are aligned with the VFM features (decoupled), for SAR, EO or SAR plus EO input.
Coupled versus decoupled semantic supervision. Top: prior methods align image tokens with EO-derived VFM features, so one token stream serves both alignment and pixel reconstruction. Bottom: SAREO-FM supervises learnable semantic queries and reconstructs SAR and EO from the modality tokens.

Fusion recovers what one sensor misses

Building-damage maps on BRIGHT from pre-event EO and post-event SAR: the fused input recovers destroyed buildings that the single-modality models predict as intact or background.

Input SAREO-FM, SAR+EO (ours) SAREO-FM, one modality Ground truth Baseline

Qualitative, in-corpus examples: BRIGHT also contributes source imagery to SAR-1M, so the scenes illustrate dense transfer rather than a full segmentation benchmark. Paper figure

Label-efficient SAR transfer

With the encoder frozen, SAREO-FM is consistently stronger than SARMAE* on FUSAR-Ship from 2-shot onward and on SAR-ACD from 5-shot onward.

Two line plots of overall accuracy against shots (1, 2, 5, 10, 20, 40) with frozen encoders, for SARMAE* (gray, dashed), SAREO-FM both (teal) and SAREO-FM sem (ochre). FUSAR-Ship, left: the SAREO-FM curves lie above SARMAE* from 2 shots on, with the widest gap at 5 shots. SAR-ACD, right: SARMAE* is higher at 1 and 2 shots; SAREO-FM sem is higher from 5 shots on.
Frozen few-shot SAR transfer. We use fixed encoder features and train only the downstream classifier.
Four SAR-ACD aircraft chips. True A220: SAREO-FM predicts A220, SARMAE predicts ARJ21. True A330: SAREO-FM A330, SARMAE Boeing787. True ARJ21: SAREO-FM ARJ21, SARMAE Boeing737. True Boeing737: SAREO-FM Boeing737, SARMAE Boeing787.
Qualitative comparison on SAR-ACD. Representative test images misclassified by SARMAE* but correctly classified by SAREO-FM.

With only 40 examples per class on the out-of-corpus SAR-ACD, end-to-end fine-tuning with semantic queries raises accuracy from 68.50% to 78.40% (Table 1); the exact frozen few-shot values are in Table 6. SARMAE*: the official SARMAE checkpoint under the same downstream protocol.

Quantitative results

SAREO-FM is best in five of six controlled SAR fine-tuning settings, surpasses its DINOv3 teacher with EO input, and joint SAR–EO input gives the best result on both paired benchmarks.

Table 1. SAR target classification with end-to-end fine-tuning. The upper block reports published OA (%) under the source-specific protocol of each method and is included only as context. The lower block reports our controlled evaluation with identical data splits and downstream pipelines; mean and standard deviation over three seeds are shown. * denotes our re-evaluation of the official SARMAE checkpoint. Bold denotes the best controlled result.
Bold best controlled resultOA (%)
Scroll for more columns
MethodVenueBackboneFUSAR-ShipMSTARSAR-ACD
40-shot30%40-shot30%40-shot30%
Published results, each under its own source protocol (context only)
ResNet-50CVPR'16ResNet-50–58.41–89.94–59.70
Swin TransformerICCV'21Swin-B–60.79–82.97–67.50
BEiTICLR'22ViT-B59.7071.1340.7069.75–79.77
CROMANeurIPS'23ViT-B83.71––––88.99
SAR-JEPAISPRS JPRS'24ViT-B85.80–91.60–75.50–
SARATR-XTIP'25HiViT-B87.70–98.10–76.40–
LoMaRWACV'25ViT-B82.70–77.00–67.40–
SUMMITIJAEOG'25ViT-B81.5071.9163.6098.3968.7084.25
Copernicus-FMICCV'25ViT-B87.61––––92.63
CoDe-MAEarXiv'26HiViT-B89.40–98.70–78.30–
MaRSAAAI'26SwinV2-B77.70–75.50–68.40–
SARMAECVPR'26ViT-B89.3092.9296.7099.61–95.06
SARMAECVPR'26ViT-L90.8692.8097.2498.92–95.63
Controlled evaluation: identical splits and downstream pipelines
Random initialization–ViT-B60.31±1.3982.11±0.5234.71±0.0938.76±8.3533.14±2.6039.19±10.27
SARMAE*CVPR'26ViT-B90.09±0.3692.82±0.2194.54±1.7297.18±0.3268.50±2.6095.01±0.85
SAREO-FM (sem)–ViT-B91.05±0.1693.13±0.1497.57±0.4299.13±0.5478.40±1.2493.34±0.55
SAREO-FM (both)–ViT-B91.41±0.2393.12±0.0696.66±0.3199.36±0.4778.00±1.0092.84±0.68

On the out-of-corpus SAR-ACD, semantic queries raise 40-shot accuracy from 68.50% to 78.40%, a 9.90-point gain; under 30% supervision, SARMAE* remains 1.67 points stronger. MSTAR and FUSAR-Ship are part of the SAR-1M pretraining corpus.

Exact frozen few-shot SAR results Table 6
Table 6. Exact frozen few-shot SAR classification results corresponding to Fig. 3 of the main paper. We freeze the pretrained encoder and optimize only a linear downstream classifier. We report mean OA (%) and sample standard deviation over five matched support draws. sem uses semantic-query features, whereas both concatenates modality-token and semantic-query features. SARMAE* denotes the official checkpoint evaluated with the same downstream protocol. Bold and underline denote the best and second-best values within each dataset and shot count, respectively; shaded rows are ours.
Bold bestUnderline second bestOA (%)
Scroll for more columns
Method / Feature1-shot2-shot5-shot10-shot20-shot40-shot
FUSAR-Ship
SARMAE*54.19±8.6964.95±6.9370.03±3.4878.18±3.2883.44±2.0388.35±0.83
SAREO-FM (both)55.56±7.5064.79±3.8172.84±3.9780.46±1.8185.32±1.7989.81±0.70
SAREO-FM (sem)55.79±9.0768.54±4.3575.74±2.4681.65±2.9185.50±1.7189.56±1.22
SAR-ACD
SARMAE*35.09±7.7739.56±6.5247.12±3.1651.68±4.0860.30±2.8670.11±1.82
SAREO-FM (both)33.81±6.3036.31±6.9144.78±4.5651.87±4.1260.69±2.7571.35±2.01
SAREO-FM (sem)33.64±6.9536.87±5.8348.10±3.0253.65±3.1961.65±3.3571.72±1.74
Masking modes and active losses Tables 7–8
Table 7. Shared-encoder sequence composition. Training counts use the 75% patch-mask ratio; downstream inference uses all available image patches.
Scroll for more columns
ConfigurationSAREOQueriesTotal
Independent/shared masking6464256384
Complete EO dropping640256320
Complete SAR dropping064256320
Unpaired SAR training640256320
SAR-only inference2560256512
EO-only inference0256256512
Joint SAR–EO inference256256256768
Table 8. Student inputs, teacher availability, and active losses. Retained image streams are masked by 75%, and all semantic queries remain active. The loss coefficients are λs = 1, λe = 1, and λalign = 0.5.
Scroll for more columns
Sample / modeProb.Student encoder image streamsTeacher inputActive objective
Paired, independent masks0.35Visible SAR + visible EOClean EOℒpixs + ℒpixe + 0.5ℒalign
Paired, shared mask0.35Visible SAR + visible EOClean EOℒpixs + ℒpixe + 0.5ℒalign
Paired, complete EO dropping0.15Visible SAR onlyClean EO, teacher onlyℒpixs + 0.5ℒalign
Paired, complete SAR dropping0.15Visible EO onlyClean EOℒpixe + 0.5ℒalign
Unpaired SAR–Visible SAR onlyUnavailableℒpixs
Unpaired EO–Not sampled in this work––
Downstream datasets and evaluation scope Table 9
Table 9. Downstream datasets and interpretation. OA denotes overall accuracy; micro-AP denotes micro-averaged average precision. The paired EuroSAT protocol uses geospatially matched SAR observations in addition to standard EO imagery; standard EuroSAT itself is EO-only.
Scroll for more columns
DatasetInputClassesMetricProtocolSAR-1M relationRole in this work
MSTARSAR10OA1–40-shot / 30%In corpusIn-corpus target transfer
FUSAR-ShipSAR10OA1–40-shot / 30%In corpusIn-corpus ship transfer
SAR-ACDSAR5OA1–40-shot / 30%Out of corpusPrimary out-of-corpus SAR test
EuroSAT protocolEO or SAR10OA1/5/10/40-shotNo out-of-corpus claimUnimodal scene transfer
NWPU-RESISC45EO45OAFew-shot / higher supervisionNo out-of-corpus claimEO scene transfer
AIDEO30OAFew-shot / higher supervisionNo out-of-corpus claimEO scene transfer
So2Sat LCZ42SAR, EO, or joint17OA5-shot, city-disjointNo out-of-corpus claimPaired single-label transfer
BigEarthNet-MMSAR, EO, or joint19micro-AP5-shot, official splitNo out-of-corpus claimPaired multi-label transfer
BRIGHTPre-EO + post-SAR4Dense labelsQualitative onlyIn-corpus sourceIn-corpus dense diagnostic
Architecture, components and readouts Tables 10–12
Table 10. Architecture summary.
Scroll for more columns
ComponentSpecification
Input resolution256×256
Patch size / grid16×16 / 16×16
SAR / EO tokens256 per available modality
Shared encoderViT-B/16, 12 blocks, D = 768
Attention / MLP12 heads / ratio 4
Semantic queries256 learnable, spatially indexed
Position encodingShared 2D sine–cosine grid
Stream encodingLearnable SAR / EO / query embeddings
TeacherFrozen DINOv3-7B patch features
Semantic projectorTwo-layer MLP
ReconstructionSeparate SAR/EO decoders
Table 11. Lifecycle of SAREO-FM components.
Scroll for more columns
ComponentPretrainingDownstream
Shared SAREO encoderTrainableUsed
Semantic queriesTrainableUsed
SAR/EO patch embeddingsTrainableUsed as available
SAR/EO reconstruction decodersTrainableDiscarded
Semantic projectorTrainableDiscarded
DINOv3-7B teacherFrozenDiscarded
Table 12. Image-level readout dimensions.
Scroll for more columns
Available inputmodsemboth
SAR or EO7687681,536
SAR + EO1,5367682,304

One encoder for SAR, EO, or both

SAREO-FM overview in three panels. Training: SAR and EO inputs are masked; their visible tokens and the learnable semantic queries enter the SAREO encoder; the SAR and EO output tokens go to a SAR decoder and an EO decoder for the SAR and EO pixel losses, and the query outputs are aligned with the features of a frozen vision foundation model applied to the EO input (alignment loss). Inference (downstream): only SAR, only EO, or joint SAR plus EO tokens, always with the semantic queries, pass through the frozen SAREO encoder to the downstream task. Modality masking strategy: cross-modal interaction with (i) independent and (ii) shared masking, and handling missing modality with (iii) EO masking and (iv) SAR masking.
Overview of SAREO-FM. The semantic-query outputs are aligned with a frozen VFM and never reach the decoders; the modality-token outputs are decoded for masked reconstruction; after pretraining, the encoder alone serves SAR, EO or joint input.
  1. 01Decoupled semantic supervision. 256 learnable semantic queries, one per patch, are aligned patch-wise with frozen DINOv3-7B features and are excluded from the decoders.
  2. 02Modality-specific reconstruction. Separate SAR and EO decoders reconstruct the masked patches (75% of each retained modality) from the modality tokens.
  3. 03Mixed modality masking. Independent, shared, EO-dropped and SAR-dropped inputs (0.35/0.35/0.15/0.15) train one ViT-B/16 encoder for SAR, EO or both.

BibTeX

@article{do2026sareofm,
  title={SAREO-FM: Decoupled Semantic Supervision for SAR-EO Foundation Models},
  author={Do, Jeonghyeok and Kim, Munchurl},
  journal={arXiv preprint arXiv:2610.09317},
  year={2026}
}