More research

An Uncertainty-aware Knowledge Distillation Diffusion Framework with Details Enhancement for PAN-Sharpening

Sungpyo Kim1 Jeonghyeok Do1 Jaehyup Lee2,† Munchurl Kim1,†

1Korea Advanced Institute of Science and Technology (KAIST), South Korea2Kyungpook National University, South Korea

† Co-corresponding authorsksp04204@kaist.ac.krehwjdgur0913@kaist.ac.krjaehyuplee@knu.ac.krmkimee@kaist.ac.kr

CVPR 2025

Abstract

Conventional methods for PAN-sharpening often struggle to restore fine details due to limitations in leveraging high-frequency information. Moreover, diffusion-based approaches lack sufficient conditioning to fully utilize Panchromatic (PAN) images and low-resolution multispectral (LRMS) inputs effectively. To address these challenges, we propose an uncertainty-aware knowledge distillation diffusion framework with details enhancement for PAN-sharpening, called U-Know-DiffPAN. The U-Know-DiffPAN incorporates uncertainty-aware knowledge distillation for effective transfer of feature details from our teacher model to a student one. The teacher model in our U-Know-DiffPAN captures frequency details through frequency selective attention, facilitating accurate reverse process learning. By conditioning the encoder on compact vector representations of PAN and LRMS and the decoder on Wavelet transforms, we enable rich frequency utilization. So, the high-capacity teacher model distills frequency-rich features into a lightweight student model aided by an uncertainty map. From this, the teacher model can guide the student model to focus on difficult image regions for PAN-sharpening via the usage of the uncertainty map. Extensive experiments on diverse datasets demonstrate the robustness and superior performance of our U-Know-DiffPAN over very recent state-of-the-art PAN-sharpening methods.

Lighter student, finer details

A diffusion teacher’s uncertainty map guides a lightweight student to difficult regions, such as cars.

Full-resolution WorldView-3 parking lot: LRMS and PAN crops of a group of cars, then PanDiff, TMDiff, CANConv and U-Know-DiffPAN (ours), each with a red box on the cars and an enlarged inset; the U-Know-DiffPAN inset shows the cars most clearly.
Full-resolution WorldView-3. PanDiff, TMDiff, CANConv and U-Know-DiffPAN (the student FSA-S); the red boxes enlarge cars, a highly uncertain region.

Quantitative results

FSA-T or FSA-S is best on every reduced-resolution metric on WV3, QB and GF2; on WV3 and QB, the lightweight student surpasses its teacher on several metrics.

Table 3. Performance comparison of different models on WV3 and QB datasets. Bold and underline highlight the best- and 2nd-best performing models. The results for the full-resolution datasets and detailed results with standard deviations are provided in Table 8.
Bold bestUnderline second best↓ lower is better · ↑ higher is better
Show
Scroll for more columns
Method Venue WV3 (Reduced-Resolution) QB (Reduced-Resolution)
PSNR↑ SSIM↑ SAM↓ ERGAS↓ SCC↑ Q8↑ PSNR↑ SSIM↑ SAM↓ ERGAS↓ SCC↑ Q4↑
Non-diffusion models
PanNetICCV'1736.1480.9663.4022.5380.9790.91335.5630.9395.2734.8560.9660.911
MSDCNNJSTARS'1836.3290.9673.3002.4890.9790.91437.0400.9544.8284.0740.9770.925
FusionNetICCV'2136.5690.9683.1882.4280.9810.91636.8210.9524.8924.1830.9750.923
LAGConvAAAI'2236.7320.9703.1532.3800.9810.91637.5650.9584.6823.8450.9800.930
S2DBPNTGRS'2337.2160.9723.0192.2450.9850.91737.3140.9564.8493.9560.9800.928
DCPNetTGRS'2437.0090.9723.0832.3010.9840.91538.0790.9634.4203.6180.9830.935
CANConvCVPR'2437.4410.9732.9272.1630.9850.91837.7950.9604.5543.7400.9820.935
Diffusion models
PanDiffTGRS'2337.0290.9713.0582.2760.9840.91337.8420.9594.6113.7230.9820.935
TMDiffTGRS'2437.4770.9732.8852.1510.9860.91537.6420.9584.6273.8040.9810.930
FSA-SOurs37.9300.9762.7972.0460.9880.92238.3610.9644.3373.5000.9840.938
FSA-TOurs37.8940.9762.8012.0550.9870.92138.3430.9644.3493.5020.9850.938

FSA-T is the teacher and FSA-S the distilled student, with 25.492M and 9.115M parameters (Table 4).

Table 4. Computational complexity comparison for diffusion-based models. Best values are in bold.
Bold bestlower is better in every column
Scroll for more columns
Method Params. (M) FLOPs (T) Time (s) Memory (GB)
PanDiff12.5560.47119.5223.260
TMDiff153.9395.51767.46110.483
FSA-T25.4921.40225.4955.910
FSA-S9.1150.34612.2872.136

Among the diffusion models, FSA-S has the fewest parameters and FLOPs, the shortest inference time and the lowest memory use.

Full resolution and standard deviations Table 8
Table 8. Additional PAN-sharpening results by our U-Know-DiffPAN and other SOTA methods for the WV3, QB, and GF2 dataset. The best (second best) performance in each block is in bold (underlined).
Bold bestUnderline second best↓ lower is better · ↑ higher is better
Show
Scroll for more columns
Model Reduced-Resolution Full-Resolution
PSNR↑ SSIM↑ SAM↓ ERGAS↓ SCC↑ Q8↑ Dλ↓ Ds↓ HQNR↑
PanNet36.148 ± 1.9580.966 ± 0.0113.402 ± 0.6722.538 ± 0.5970.979 ± 0.0060.913 ± 0.0870.035 ± 0.0140.049 ± 0.0190.918 ± 0.031
MSDCNN36.329 ± 1.7480.967 ± 0.0103.300 ± 0.6542.489 ± 0.6200.979 ± 0.0070.914 ± 0.0870.028 ± 0.0130.050 ± 0.0200.924 ± 0.030
FusionNet36.569 ± 1.6660.968 ± 0.0093.188 ± 0.6282.428 ± 0.6210.981 ± 0.0070.916 ± 0.0870.029 ± 0.0110.053 ± 0.0210.920 ± 0.030
LAGNet36.732 ± 1.7230.970 ± 0.0093.153 ± 0.6082.380 ± 0.6170.981 ± 0.0070.916 ± 0.0870.033 ± 0.0120.055 ± 0.0230.915 ± 0.033
S2DBPN37.216 ± 1.8880.972 ± 0.0093.019 ± 0.5882.245 ± 0.5410.985 ± 0.0050.917 ± 0.0910.025 ± 0.0100.030 ± 0.0100.946 ± 0.018
DCPNet37.009 ± 1.7350.972 ± 0.0083.083 ± 0.5372.301 ± 0.5690.984 ± 0.0050.915 ± 0.0920.043 ± 0.0180.036 ± 0.0120.923 ± 0.027
CANConv37.441 ± 1.7880.973 ± 0.0082.927 ± 0.5362.163 ± 0.4810.985 ± 0.0050.918 ± 0.0820.020 ± 0.0080.030 ± 0.0080.951 ± 0.013
PanDiff37.029 ± 1.7960.971 ± 0.0083.058 ± 0.5672.276 ± 0.5450.984 ± 0.0040.913 ± 0.0840.014 ± 0.0050.034 ± 0.0050.952 ± 0.009
TMDiff37.477 ± 1.9230.973 ± 0.0082.885 ± 0.5492.151 ± 0.4580.986 ± 0.0040.915 ± 0.0860.018 ± 0.0070.059 ± 0.0090.924 ± 0.015
FSA-T37.894 ± 1.8200.976 ± 0.0072.801 ± 0.5172.055 ± 0.4630.987 ± 0.0030.921 ± 0.0830.014 ± 0.0050.032 ± 0.0030.954 ± 0.006
FSA-S37.930 ± 1.8240.976 ± 0.0072.797 ± 0.5262.046 ± 0.4540.988 ± 0.0030.922 ± 0.0830.016 ± 0.0060.029 ± 0.0030.955 ± 0.008
Ablations Tables 5, 6 and 7
Table 5. Comparison of Results with and without FFA and HQFE Blocks in FSA-T. Best values are in bold.
Bold best↓ lower is better · ↑ higher is better
Scroll for more columns
Encoder
FFA
Decoder
HQFE
GF2 (Reduced-Resolution)
SAM↓ ERGAS↓ SCC↑ Q4↑
0.654 ± 0.1120.661 ± 0.0760.992 ± 0.0010.986 ± 0.007
✓0.654 ± 0.1120.636 ± 0.0770.993 ± 0.0020.986 ± 0.007
✓0.617 ± 0.1170.556 ± 0.1040.993 ± 0.0020.987 ± 0.007
✓✓0.603 ± 0.1020.537 ± 0.0770.994 ± 0.0010.988 ± 0.006
Table 6. Impact of our U-know loss function design. Best values are in bold.
Bold best↓ lower is better · ↑ higher is better
Scroll for more columns
Loss func. GF2 (Full-Resolution)
Dλ↓ Ds↓ HQNR↑
ℒ10.026 ± 0.0140.040 ± 0.0170.935 ± 0.020
ℒKD0.025 ± 0.0150.038 ± 0.0160.938 ± 0.021
ℒU-know0.018 ± 0.0110.037 ± 0.0070.944 ± 0.012
Table 7. Comparison of results between DWT and SWT conditioning at the SWTCA block in FSA-T Ψ, with the best values in bold.
Bold best↓ lower is better · ↑ higher is better
Scroll for more columns
Condition GF2 (Reduced-Resolution)
SAM↓ ERGAS↓ SCC↑ Q4↑
DWT0.646 ± 0.1170.567 ± 0.0950.993 ± 0.0020.987 ± 0.007
SWT0.603 ± 0.1020.537 ± 0.0770.994 ± 0.0010.988 ± 0.006
Datasets Table 1
Table 1. Detailed information of Worldview-3, QuickBird, and GaoFen-2 datasets.
Scroll for more columns
Satellite WorldView-3 QuickBird GaoFen-2
Number of Band844
Spatial Resolution (m)PAN0.30.60.8
LRMS1.22.43.2
Radiometric Resolution (bit)111110
Number of (Train / Test) Images9,714 / 2017,139 / 2019,809 / 20
Patch SizePAN64×64×164×64×164×64×1
LRMS16×16×816×16×416×16×4

Trust the teacher where it is certain

U-Know-DiffPAN in two steps. Step 1, teacher pre-training: FSA-T receives the noisy residual, the LRMS and PAN images, the compact vector v and the SWT condition, and outputs the residual and the uncertainty map, trained with the uncertainty-driven diffusion loss. Step 2, student training: the frozen, pretrained FSA-T provides its output, uncertainty map and features; the student FSA-S, conditioned only on LRMS and PAN, is trained with the hard, soft and feature losses.
U-Know-DiffPAN. Step 1 pre-trains the teacher FSA-T, which predicts the output and an uncertainty map; Step 2 distills it into the lightweight student FSA-S.
  1. 01Frequency-selective teacher. FSA-T conditions its encoder on a compact vector of PAN and LRMS (FFA blocks); in its decoder, HQFE blocks refine frequencies with Fourier attention (FTCA) and inject stationary-wavelet components of PAN and LRMS by cross-attention (SWTCA).
  2. 02Uncertainty from the teacher. Trained with an uncertainty-driven diffusion loss, FSA-T predicts the residual together with a pixel-wise uncertainty map θ̂, which is high on object edges.
  3. 03Uncertainty-aware distillation. The student FSA-S (ResBlocks only) learns from the ground truth weighted by τ + θ̂, from the teacher’s output weighted by τ − θ̂, and from the teacher’s intermediate features.

BibTeX

@inproceedings{kim2025uknowdiffpan,
  title={U-Know-DiffPAN: An uncertainty-aware knowledge distillation diffusion framework with details enhancement for PAN-sharpening},
  author={Kim, Sungpyo and Do, Jeonghyeok and Lee, Jaehyup and Kim, Munchurl},
  booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  pages={23069--23079},
  year={2025},
  organization={IEEE}
}