More research

Learning Modality-Consistent Alignment for PAN-Sharpening

Jeonghyeok Do1 Sungpyo Kim1 Geunhyuk Youk1 Jaehyup Lee2,† Munchurl Kim1,†

1Korea Advanced Institute of Science and Technology (KAIST), South Korea2Kyungpook National University, South Korea

† Co-corresponding authorsehwjdgur0913@kaist.ac.krksp04204@kaist.ac.krrmsgurkjg@kaist.ac.krjaehyuplee@knu.ac.krmkimee@kaist.ac.kr

ICCV 2025

Abstract

PAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment—caused by sensor placement, acquisition timing, and resolution disparity—induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present.

To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN’s high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11× faster inference time and 0.63× the memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions.

Misaligned inputs, sharp outputs

Aligns MS texture and PAN structure both ways; beats the latest state of the art in all metrics at 50.11× faster inference.

WV3 scene at full resolution: the input LRMS and PAN images with a red box, their zoomed-in views, and the zoomed-in outputs of S2DBPN, PanDiff, DCPNet, TMDiff, CANConv and PAN-Crafter around parked cars and white-roofed buildings.
Full-resolution comparison on WV3. Input LRMS and PAN with zoomed-in views, five recent methods and PAN-Crafter: minimal artifacts near buildings and cars, where other approaches often yield blurred or distorted results.

Quantitative results

PAN-Crafter is best in 8 of 9 quality metrics on WV3, in all six on GF2, in 4 of 6 on QB and in 4 of 5 on the unseen WV2, with the lowest memory use.

Table 1. Quantitative comparison on the WV3 dataset. Deep learning-based PS methods at full and reduced resolution. Bold indicates the best performance in each metric. The inference time and memory usage are measured on a 256 × 256 × 8 HRMS target at reduced resolution.
Bold best↓ lower is better · ↑ higher is better
Scroll for more columns
MethodVenueFull resolutionReduced resolutionTime↓ (s)Memory↓ (GB)
HQNR↑Ds↓Dλ↓ERGAS↓SCC↑SAM↓Q8↑PSNR↑SSIM↑
PanNetICCV'170.9180.0490.0352.5380.9793.4020.91336.1480.966––
MSDCNNJSTARS'180.9240.0500.0282.4890.9793.3000.91436.3290.967––
FusionNetICCV'210.9200.0530.0292.4280.9813.1880.91636.5690.968––
LAGConvAAAI'220.9150.0550.0332.3800.9813.1530.91636.7320.9700.0043.281
S2DBPNTGRS'230.9460.0300.0252.2450.9853.0190.91737.2160.9720.0052.387
PanDiffTGRS'230.9520.0340.0142.2760.9843.0580.91337.0290.9712.9552.328
DCPNetTGRS'240.9230.0360.0432.3010.9843.0830.91537.0090.9720.1097.213
TMDiffTGRS'240.9240.0590.0182.1510.9862.8850.91537.4770.9739.9979.910
CANConvCVPR'240.9510.0300.0202.1630.9852.9270.91837.4410.9730.4512.713
PAN-Crafter–0.9580.0270.0162.0400.9882.7870.92237.9560.9760.0091.711

PAN-Crafter needs 0.009 s and 1.711 GB: 50.11× faster than CANConv, which relies on k-means clustering for kernel generation, and 328.33× and 1110.78× faster than the diffusion models PanDiff and TMDiff.

Ablation of MARs and CM3A Table 4
Table 4. Ablation studies on CM3A and MARs on the WV3 dataset. The combination of both components achieves the best performance, highlighting their synergistic effect in jointly refining spatial and spectral consistency.
Bold best✓ component used↓ lower is better · ↑ higher is better
Scroll for more columns
CM3AMARsWV3 dataset
HQNR↑ERGAS↓SAM↓PSNR↑Time↓ (s)Memory↓ (GB)
––0.9482.2322.98037.2450.0061.537
✓–0.9492.2122.97037.2850.0071.556
–✓0.9562.1222.87337.6020.0091.701
✓✓0.9582.0402.78737.9560.0091.711

MARs alone raises PSNR from 37.245 to 37.602 dB; CM3A alone gives a marginal gain (37.285 dB), but together with MARs it reaches 37.956 dB and the best HQNR, ERGAS and SAM.

Computational cost Table 10
Table 10. Computational efficiency comparison of deep learning-based PS methods. Inference time (s), memory usage (MB), FLOPs (G) and parameter count (M); the paper lists one column per method.
Scroll for more columns
MethodTime (s)Memory (MB)FLOPs (G)Params. (M)
LAGConv0.0043360.18.430.15
S2DBPN0.0052444.0158.9416.19
PanDiff2.9552383.662.079.52
DCPNet0.1097386.8105.401.414
TMDiff9.99710147.41284.42154.10
CANConv0.4512777.652.210.79
PAN-Crafter0.0091751.979.037.17

PAN-Crafter uses the least memory of the compared methods (1751.9 MB) and is faster than all of them except LAGConv and S2DBPN.

Reconstruct both modalities, align both ways

PAN-Crafter architecture. The up-sampled LRMS image and the PAN image are concatenated and passed through a U-Net of convolutions, ResBlocks and AttnBlocks with down- and up-sampling; a Modality-Adaptive Reconstruction switch selects MS mode, which adds the output to the LRMS image to give the HRMS image, or PAN mode, which adds it to a down- and up-sampled PAN image repeated over the channels to back-reconstruct the PAN image. Insets show the ResBlock (LayerNorm, SiLU, convolution, mode-dependent modulation) and the AttnBlock (LayerNorm, CM3A, select, feed-forward).
PAN-Crafter architecture. An encoder–decoder with ResBlocks and CM3A at multiple scales; the MARs mode switches the output between the HRMS image and the back-reconstructed PAN image.
  1. 01Modality-Adaptive Reconstruction (MARs). One network, two modes: MS mode predicts the HRMS image, PAN mode back-reconstructs the PAN image, whose high-frequency details act as auxiliary self-supervision. Inference uses MS mode.
  2. 02CM3A, in both directions. Local 3 × 3 attention aligns MS texture to PAN structure in MS mode and PAN structure to MS texture in PAN mode; down-sampled input images replace fixed positional embeddings.
  3. 03A multi-scale U-Net. ResBlocks with mode-dependent modulation at every scale; AttnBlocks with CM3A only at the low- and mid-resolution stages, to reduce computational overhead: 0.009 s and 1.711 GB per 256 × 256 × 8 target.

BibTeX

@inproceedings{do2025pancrafter,
  title={PAN-Crafter: Learning modality-consistent alignment for PAN-sharpening},
  author={Do, Jeonghyeok and Kim, Sungpyo and Youk, Geunhyuk and Lee, Jaehyup and Kim, Munchurl},
  booktitle={2025 IEEE/CVF International Conference on Computer Vision (ICCV)},
  pages={4242--4252},
  year={2025},
  organization={IEEE}
}