Semantics to queries, pixels to tokens
Semantic queries take the VFM supervision, SAR and EO tokens take masked reconstruction, all in one shared encoder.
Fusion recovers what one sensor misses
Building-damage maps on BRIGHT from pre-event EO and post-event SAR: the fused input recovers destroyed buildings that the single-modality models predict as intact or background.
Use the arrows or the ← → keys to change scene · click a tile to compare it with the ground truthSwipe or use the arrows to change scene · tap a tile to compare it with the ground truth
Qualitative, in-corpus examples: BRIGHT also contributes source imagery to SAR-1M, so the scenes illustrate dense transfer rather than a full segmentation benchmark. Paper figure
Label-efficient SAR transfer
With the encoder frozen, SAREO-FM is consistently stronger than SARMAE* on FUSAR-Ship from 2-shot onward and on SAR-ACD from 5-shot onward.
With only 40 examples per class on the out-of-corpus SAR-ACD, end-to-end fine-tuning with semantic queries raises accuracy from 68.50% to 78.40% (Table 1); the exact frozen few-shot values are in Table 6. SARMAE*: the official SARMAE checkpoint under the same downstream protocol.
Quantitative results
SAREO-FM is best in five of six controlled SAR fine-tuning settings, surpasses its DINOv3 teacher with EO input, and joint SAR–EO input gives the best result on both paired benchmarks.
| Method | Venue | Backbone | FUSAR-Ship | MSTAR | SAR-ACD | |||
|---|---|---|---|---|---|---|---|---|
| 40-shot | 30% | 40-shot | 30% | 40-shot | 30% | |||
| Published results, each under its own source protocol (context only) | ||||||||
| ResNet-50 | CVPR'16 | ResNet-50 | – | 58.41 | – | 89.94 | – | 59.70 |
| Swin Transformer | ICCV'21 | Swin-B | – | 60.79 | – | 82.97 | – | 67.50 |
| BEiT | ICLR'22 | ViT-B | 59.70 | 71.13 | 40.70 | 69.75 | – | 79.77 |
| CROMA | NeurIPS'23 | ViT-B | 83.71 | – | – | – | – | 88.99 |
| SAR-JEPA | ISPRS JPRS'24 | ViT-B | 85.80 | – | 91.60 | – | 75.50 | – |
| SARATR-X | TIP'25 | HiViT-B | 87.70 | – | 98.10 | – | 76.40 | – |
| LoMaR | WACV'25 | ViT-B | 82.70 | – | 77.00 | – | 67.40 | – |
| SUMMIT | IJAEOG'25 | ViT-B | 81.50 | 71.91 | 63.60 | 98.39 | 68.70 | 84.25 |
| Copernicus-FM | ICCV'25 | ViT-B | 87.61 | – | – | – | – | 92.63 |
| CoDe-MAE | arXiv'26 | HiViT-B | 89.40 | – | 98.70 | – | 78.30 | – |
| MaRS | AAAI'26 | SwinV2-B | 77.70 | – | 75.50 | – | 68.40 | – |
| SARMAE | CVPR'26 | ViT-B | 89.30 | 92.92 | 96.70 | 99.61 | – | 95.06 |
| SARMAE | CVPR'26 | ViT-L | 90.86 | 92.80 | 97.24 | 98.92 | – | 95.63 |
| Controlled evaluation: identical splits and downstream pipelines | ||||||||
| Random initialization | – | ViT-B | 60.31±1.39 | 82.11±0.52 | 34.71±0.09 | 38.76±8.35 | 33.14±2.60 | 39.19±10.27 |
| SARMAE* | CVPR'26 | ViT-B | 90.09±0.36 | 92.82±0.21 | 94.54±1.72 | 97.18±0.32 | 68.50±2.60 | 95.01±0.85 |
| SAREO-FM (sem) | – | ViT-B | 91.05±0.16 | 93.13±0.14 | 97.57±0.42 | 99.13±0.54 | 78.40±1.24 | 93.34±0.55 |
| SAREO-FM (both) | – | ViT-B | 91.41±0.23 | 93.12±0.06 | 96.66±0.31 | 99.36±0.47 | 78.00±1.00 | 92.84±0.68 |
On the out-of-corpus SAR-ACD, semantic queries raise 40-shot accuracy from 68.50% to 78.40%, a 9.90-point gain; under 30% supervision, SARMAE* remains 1.67 points stronger. MSTAR and FUSAR-Ship are part of the SAR-1M pretraining corpus.
| Method / Feature | SAR | EO | ||||||
|---|---|---|---|---|---|---|---|---|
| 1-shot | 5-shot | 10-shot | 40-shot | 1-shot | 5-shot | 10-shot | 40-shot | |
| DINOv3 (teacher) | 38.17±4.61 | 50.72±2.68 | 56.01±1.81 | 64.35±1.56 | 59.51±3.53 | 75.96±3.15 | 83.34±1.42 | 90.88±0.65 |
| SARMAE* | 37.22±5.71 | 48.64±3.00 | 53.97±1.39 | 64.72±1.12 | 43.56±6.57 | 59.57±3.05 | 66.49±1.97 | 78.36±1.07 |
| CoDe-MAE | – | – | 59.88† | – | – | – | 81.18† | – |
| SAREO-FM (both) | 39.07±5.98 | 53.54±2.21 | 59.36±1.46 | 69.24±0.96 | 64.41±5.05 | 81.09±1.61 | 86.76±1.24 | 92.58±0.75 |
With EO input, SAREO-FM surpasses its frozen DINOv3 teacher by 4.90, 5.13, 3.42 and 1.70 points from 1 to 40 shots; with SAR input it improves over SARMAE* at every shot count, by 1.85 to 5.39 points.
| Input | Method | Feature | Benchmark | |
|---|---|---|---|---|
| So2Sat LCZ42 | BigEarthNet-MM | |||
| SAR+EO | Random initialization | mod | 40.56±2.25 | 54.29±2.30 |
| EO | DINOv3 (teacher) | mod | 54.25±3.87 | 64.73±1.43 |
| EO | SAREO-FM | mod | 58.14±2.90 | 64.42±1.44 |
| SAR | SARMAE* | mod | 28.96±1.68 | 57.44±0.69 |
| SAR | SAREO-FM | mod | 29.59±2.64 | 60.47±0.79 |
| SAR+EO | SAREO-FM | both | 58.79±4.19 | 66.70±0.58 |
On BigEarthNet-MM, joint input reaches 66.70 micro-AP, 2.28 points above the best SAREO-FM unimodal input; on So2Sat LCZ42, adding SAR to EO yields a smaller 0.65-point gain.
| Pretraining configuration | MSTAR (OA↑) | FUSAR (OA↑) | Avg. (OA↑) | Δ |
|---|---|---|---|---|
| Random initialization | 26.26±1.86 | 47.86±4.89 | 37.06 | – |
| Plain MAE | 66.16±3.15 | 76.46±2.52 | 71.31 | +34.25 |
| + DINOv3 supervision | 73.49±2.31 | 79.90±2.49 | 76.70 | +5.39 |
| + Decoupled semantic supervision | 77.11±2.14 | 82.55±2.15 | 79.83 | +3.13 |
| + Raw EO modality (SAREO-FM) | 78.41±2.14 | 84.00±2.88 | 81.21 | +1.38 |
Decoupling semantic prediction from modality reconstruction adds 3.13 points on average; the complete SAREO-FM improves on plain MAE by 9.90 points.
Exact frozen few-shot SAR results Table 6
| Method / Feature | 1-shot | 2-shot | 5-shot | 10-shot | 20-shot | 40-shot |
|---|---|---|---|---|---|---|
| FUSAR-Ship | ||||||
| SARMAE* | 54.19±8.69 | 64.95±6.93 | 70.03±3.48 | 78.18±3.28 | 83.44±2.03 | 88.35±0.83 |
| SAREO-FM (both) | 55.56±7.50 | 64.79±3.81 | 72.84±3.97 | 80.46±1.81 | 85.32±1.79 | 89.81±0.70 |
| SAREO-FM (sem) | 55.79±9.07 | 68.54±4.35 | 75.74±2.46 | 81.65±2.91 | 85.50±1.71 | 89.56±1.22 |
| SAR-ACD | ||||||
| SARMAE* | 35.09±7.77 | 39.56±6.52 | 47.12±3.16 | 51.68±4.08 | 60.30±2.86 | 70.11±1.82 |
| SAREO-FM (both) | 33.81±6.30 | 36.31±6.91 | 44.78±4.56 | 51.87±4.12 | 60.69±2.75 | 71.35±2.01 |
| SAREO-FM (sem) | 33.64±6.95 | 36.87±5.83 | 48.10±3.02 | 53.65±3.19 | 61.65±3.35 | 71.72±1.74 |
Masking modes and active losses Tables 7–8
| Configuration | SAR | EO | Queries | Total |
|---|---|---|---|---|
| Independent/shared masking | 64 | 64 | 256 | 384 |
| Complete EO dropping | 64 | 0 | 256 | 320 |
| Complete SAR dropping | 0 | 64 | 256 | 320 |
| Unpaired SAR training | 64 | 0 | 256 | 320 |
| SAR-only inference | 256 | 0 | 256 | 512 |
| EO-only inference | 0 | 256 | 256 | 512 |
| Joint SAR–EO inference | 256 | 256 | 256 | 768 |
| Sample / mode | Prob. | Student encoder image streams | Teacher input | Active objective |
|---|---|---|---|---|
| Paired, independent masks | 0.35 | Visible SAR + visible EO | Clean EO | ℒpixs + ℒpixe + 0.5ℒalign |
| Paired, shared mask | 0.35 | Visible SAR + visible EO | Clean EO | ℒpixs + ℒpixe + 0.5ℒalign |
| Paired, complete EO dropping | 0.15 | Visible SAR only | Clean EO, teacher only | ℒpixs + 0.5ℒalign |
| Paired, complete SAR dropping | 0.15 | Visible EO only | Clean EO | ℒpixe + 0.5ℒalign |
| Unpaired SAR | – | Visible SAR only | Unavailable | ℒpixs |
| Unpaired EO | – | Not sampled in this work | – | – |
Downstream datasets and evaluation scope Table 9
| Dataset | Input | Classes | Metric | Protocol | SAR-1M relation | Role in this work |
|---|---|---|---|---|---|---|
| MSTAR | SAR | 10 | OA | 1–40-shot / 30% | In corpus | In-corpus target transfer |
| FUSAR-Ship | SAR | 10 | OA | 1–40-shot / 30% | In corpus | In-corpus ship transfer |
| SAR-ACD | SAR | 5 | OA | 1–40-shot / 30% | Out of corpus | Primary out-of-corpus SAR test |
| EuroSAT protocol | EO or SAR | 10 | OA | 1/5/10/40-shot | No out-of-corpus claim | Unimodal scene transfer |
| NWPU-RESISC45 | EO | 45 | OA | Few-shot / higher supervision | No out-of-corpus claim | EO scene transfer |
| AID | EO | 30 | OA | Few-shot / higher supervision | No out-of-corpus claim | EO scene transfer |
| So2Sat LCZ42 | SAR, EO, or joint | 17 | OA | 5-shot, city-disjoint | No out-of-corpus claim | Paired single-label transfer |
| BigEarthNet-MM | SAR, EO, or joint | 19 | micro-AP | 5-shot, official split | No out-of-corpus claim | Paired multi-label transfer |
| BRIGHT | Pre-EO + post-SAR | 4 | Dense labels | Qualitative only | In-corpus source | In-corpus dense diagnostic |
Architecture, components and readouts Tables 10–12
| Component | Specification |
|---|---|
| Input resolution | 256×256 |
| Patch size / grid | 16×16 / 16×16 |
| SAR / EO tokens | 256 per available modality |
| Shared encoder | ViT-B/16, 12 blocks, D = 768 |
| Attention / MLP | 12 heads / ratio 4 |
| Semantic queries | 256 learnable, spatially indexed |
| Position encoding | Shared 2D sine–cosine grid |
| Stream encoding | Learnable SAR / EO / query embeddings |
| Teacher | Frozen DINOv3-7B patch features |
| Semantic projector | Two-layer MLP |
| Reconstruction | Separate SAR/EO decoders |
| Component | Pretraining | Downstream |
|---|---|---|
| Shared SAREO encoder | Trainable | Used |
| Semantic queries | Trainable | Used |
| SAR/EO patch embeddings | Trainable | Used as available |
| SAR/EO reconstruction decoders | Trainable | Discarded |
| Semantic projector | Trainable | Discarded |
| DINOv3-7B teacher | Frozen | Discarded |
| Available input | mod | sem | both |
|---|---|---|---|
| SAR or EO | 768 | 768 | 1,536 |
| SAR + EO | 1,536 | 768 | 2,304 |
One encoder for SAR, EO, or both
- 01Decoupled semantic supervision. 256 learnable semantic queries, one per patch, are aligned patch-wise with frozen DINOv3-7B features and are excluded from the decoders.
- 02Modality-specific reconstruction. Separate SAR and EO decoders reconstruct the masked patches (75% of each retained modality) from the modality tokens.
- 03Mixed modality masking. Independent, shared, EO-dropped and SAR-dropped inputs (0.35/0.35/0.15/0.15) train one ViT-B/16 encoder for SAR, EO or both.
BibTeX
@article{do2026sareofm,
title={SAREO-FM: Decoupled Semantic Supervision for SAR-EO Foundation Models},
author={Do, Jeonghyeok and Kim, Munchurl},
journal={arXiv preprint arXiv:2610.09317},
year={2026}
}