Misaligned inputs, sharp outputs
Aligns MS texture and PAN structure both ways; beats the latest state of the art in all metrics at 50.11× faster inference.
All methods, same scenes
The zoomed-in crops of the paper’s full-resolution figures at their native resolution: recent methods and PAN-Crafter on the same LRMS and PAN input, including WV2, a satellite unseen in training. Select a tile to compare it with PAN-Crafter.
Use the arrows or the ← → keys to change scene · click a tile to compare it with PAN-CrafterSwipe or use the arrows to change scene · tap a tile to compare it with PAN-Crafter
Whole scenes at full resolution
Complete 512 × 512 full-resolution scenes from each satellite: the input LRMS, CANConv and PAN-Crafter.
Click a tile to open the comparison slider against PAN-CrafterTap a tile to open the comparison slider against PAN-Crafter
Paper figures
Full- and reduced-resolution comparisons with error maps, the unseen WV2 satellite, and the PAN–MS misalignment that CM3A addresses.
Quantitative results
PAN-Crafter is best in 8 of 9 quality metrics on WV3, in all six on GF2, in 4 of 6 on QB and in 4 of 5 on the unseen WV2, with the lowest memory use.
| Method | Venue | Full resolution | Reduced resolution | Time↓ (s) | Memory↓ (GB) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HQNR↑ | Ds↓ | Dλ↓ | ERGAS↓ | SCC↑ | SAM↓ | Q8↑ | PSNR↑ | SSIM↑ | ||||
| PanNet | ICCV'17 | 0.918 | 0.049 | 0.035 | 2.538 | 0.979 | 3.402 | 0.913 | 36.148 | 0.966 | – | – |
| MSDCNN | JSTARS'18 | 0.924 | 0.050 | 0.028 | 2.489 | 0.979 | 3.300 | 0.914 | 36.329 | 0.967 | – | – |
| FusionNet | ICCV'21 | 0.920 | 0.053 | 0.029 | 2.428 | 0.981 | 3.188 | 0.916 | 36.569 | 0.968 | – | – |
| LAGConv | AAAI'22 | 0.915 | 0.055 | 0.033 | 2.380 | 0.981 | 3.153 | 0.916 | 36.732 | 0.970 | 0.004 | 3.281 |
| S2DBPN | TGRS'23 | 0.946 | 0.030 | 0.025 | 2.245 | 0.985 | 3.019 | 0.917 | 37.216 | 0.972 | 0.005 | 2.387 |
| PanDiff | TGRS'23 | 0.952 | 0.034 | 0.014 | 2.276 | 0.984 | 3.058 | 0.913 | 37.029 | 0.971 | 2.955 | 2.328 |
| DCPNet | TGRS'24 | 0.923 | 0.036 | 0.043 | 2.301 | 0.984 | 3.083 | 0.915 | 37.009 | 0.972 | 0.109 | 7.213 |
| TMDiff | TGRS'24 | 0.924 | 0.059 | 0.018 | 2.151 | 0.986 | 2.885 | 0.915 | 37.477 | 0.973 | 9.997 | 9.910 |
| CANConv | CVPR'24 | 0.951 | 0.030 | 0.020 | 2.163 | 0.985 | 2.927 | 0.918 | 37.441 | 0.973 | 0.451 | 2.713 |
| PAN-Crafter | – | 0.958 | 0.027 | 0.016 | 2.040 | 0.988 | 2.787 | 0.922 | 37.956 | 0.976 | 0.009 | 1.711 |
PAN-Crafter needs 0.009 s and 1.711 GB: 50.11× faster than CANConv, which relies on k-means clustering for kernel generation, and 328.33× and 1110.78× faster than the diffusion models PanDiff and TMDiff.
| Method | GF2 | QB | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Full resolution | Reduced resolution | Full resolution | Reduced resolution | |||||||||
| HQNR↑ | Ds↓ | ERGAS↓ | SCC↑ | SAM↓ | PSNR↑ | HQNR↑ | Ds↓ | ERGAS↓ | SCC↑ | SAM↓ | PSNR↑ | |
| PanNet | 0.929 | 0.052 | 1.038 | 0.975 | 1.050 | 39.197 | 0.851 | 0.092 | 4.856 | 0.966 | 5.273 | 35.563 |
| MSDCNN | 0.898 | 0.079 | 0.862 | 0.983 | 0.946 | 40.730 | 0.888 | 0.058 | 4.074 | 0.977 | 4.828 | 37.040 |
| FusionNet | 0.865 | 0.105 | 0.960 | 0.980 | 0.971 | 39.866 | 0.853 | 0.079 | 4.183 | 0.975 | 4.892 | 36.821 |
| LAGConv | 0.895 | 0.078 | 0.816 | 0.985 | 0.886 | 41.147 | 0.892 | 0.035 | 3.845 | 0.980 | 4.682 | 37.565 |
| S2DBPN | 0.935 | 0.046 | 0.686 | 0.990 | 0.772 | 42.686 | 0.908 | 0.036 | 3.956 | 0.980 | 4.849 | 37.314 |
| PanDiff | 0.936 | 0.045 | 0.674 | 0.990 | 0.767 | 42.827 | 0.919 | 0.055 | 3.723 | 0.982 | 4.611 | 37.842 |
| DCPNet | 0.953 | 0.024 | 0.724 | 0.988 | 0.806 | 42.312 | 0.880 | 0.073 | 3.618 | 0.983 | 4.420 | 38.079 |
| TMDiff | 0.942 | 0.030 | 0.754 | 0.988 | 0.764 | 41.896 | 0.901 | 0.068 | 3.804 | 0.981 | 4.627 | 37.642 |
| CANConv | 0.919 | 0.063 | 0.653 | 0.991 | 0.722 | 43.166 | 0.893 | 0.070 | 3.740 | 0.982 | 4.554 | 37.795 |
| PAN-Crafter | 0.964 | 0.017 | 0.522 | 0.994 | 0.596 | 45.076 | 0.920 | 0.039 | 3.570 | 0.984 | 4.426 | 38.195 |
QB is known to be the most challenging dataset; PAN-Crafter is slightly behind in Ds and SAM there but best in HQNR, ERGAS, SCC and PSNR.
| Method | WV2 dataset (unseen satellite dataset) | ||||
|---|---|---|---|---|---|
| HQNR↑ | ERGAS↓ | SCC↑ | SAM↓ | PSNR↑ | |
| PanNet | 0.875 | 5.481 | 0.876 | 7.040 | 27.120 |
| MSDCNN | 0.862 | 4.930 | 0.905 | 5.898 | 27.901 |
| FusionNet | 0.862 | 5.100 | 0.902 | 6.118 | 27.616 |
| LAGConv | 0.902 | 5.133 | 0.885 | 6.094 | 27.525 |
| S2DBPN | 0.813 | 5.703 | 0.915 | 7.063 | 26.748 |
| PanDiff | 0.932 | 4.291 | 0.916 | 5.430 | 28.964 |
| DCPNet | 0.797 | 5.507 | 0.931 | 10.174 | 27.063 |
| TMDiff | 0.874 | 5.157 | 0.875 | 6.087 | 27.473 |
| CANConv | 0.876 | 4.328 | 0.918 | 5.481 | 29.005 |
| PAN-Crafter | 0.942 | 4.169 | 0.924 | 5.078 | 29.276 |
Fully zero-shot: all models are trained on WV3 and tested on WV2 without any fine-tuning.
Ablation of MARs and CM3A Table 4
| CM3A | MARs | WV3 dataset | |||||
|---|---|---|---|---|---|---|---|
| HQNR↑ | ERGAS↓ | SAM↓ | PSNR↑ | Time↓ (s) | Memory↓ (GB) | ||
| – | – | 0.948 | 2.232 | 2.980 | 37.245 | 0.006 | 1.537 |
| ✓ | – | 0.949 | 2.212 | 2.970 | 37.285 | 0.007 | 1.556 |
| – | ✓ | 0.956 | 2.122 | 2.873 | 37.602 | 0.009 | 1.701 |
| ✓ | ✓ | 0.958 | 2.040 | 2.787 | 37.956 | 0.009 | 1.711 |
MARs alone raises PSNR from 37.245 to 37.602 dB; CM3A alone gives a marginal gain (37.285 dB), but together with MARs it reaches 37.956 dB and the best HQNR, ERGAS and SAM.
Computational cost Table 10
| Method | Time (s) | Memory (MB) | FLOPs (G) | Params. (M) |
|---|---|---|---|---|
| LAGConv | 0.004 | 3360.1 | 8.43 | 0.15 |
| S2DBPN | 0.005 | 2444.0 | 158.94 | 16.19 |
| PanDiff | 2.955 | 2383.6 | 62.07 | 9.52 |
| DCPNet | 0.109 | 7386.8 | 105.40 | 1.414 |
| TMDiff | 9.997 | 10147.4 | 1284.42 | 154.10 |
| CANConv | 0.451 | 2777.6 | 52.21 | 0.79 |
| PAN-Crafter | 0.009 | 1751.9 | 79.03 | 7.17 |
PAN-Crafter uses the least memory of the compared methods (1751.9 MB) and is faster than all of them except LAGConv and S2DBPN.
Reconstruct both modalities, align both ways
- 01Modality-Adaptive Reconstruction (MARs). One network, two modes: MS mode predicts the HRMS image, PAN mode back-reconstructs the PAN image, whose high-frequency details act as auxiliary self-supervision. Inference uses MS mode.
- 02CM3A, in both directions. Local 3 × 3 attention aligns MS texture to PAN structure in MS mode and PAN structure to MS texture in PAN mode; down-sampled input images replace fixed positional embeddings.
- 03A multi-scale U-Net. ResBlocks with mode-dependent modulation at every scale; AttnBlocks with CM3A only at the low- and mid-resolution stages, to reduce computational overhead: 0.009 s and 1.711 GB per 256 × 256 × 8 target.
BibTeX
@inproceedings{do2025pancrafter,
title={PAN-Crafter: Learning modality-consistent alignment for PAN-sharpening},
author={Do, Jeonghyeok and Kim, Sungpyo and Youk, Geunhyuk and Lee, Jaehyup and Kim, Munchurl},
booktitle={2025 IEEE/CVF International Conference on Computer Vision (ICCV)},
pages={4242--4252},
year={2025},
organization={IEEE}
}