A Novel VLM-Guided Diffusion Model for Remote Sensing Image Super-Resolution

  • Sung, Mingyu; 
  • Gong, Mu-Gyeong; 
  • Ham, Seung-Jae; 
  • Kim, Il-Min; 
  • Yun, Sangseok; 
  • ... Kang, Jae-Mo
Citations

WEB OF SCIENCE

2
Citations

SCOPUS

2

초록

Super-resolution (SR) of remote sensing imagery based on generative AI models is vital for practical applications such as urban planning and disaster assessment. However, current approaches suffer from poor performance tradeoffs among the pivotal, yet competing, objectives: perceptual quality, factual accuracy, and inference speed. To break through this limitation, we propose a novel and high-performing two-stage SR framework for remote sensing imagery based on a generative diffusion model. First, in Stage 1, factually grounded base images are generated by employing a guidance-free diffusion process relying solely on the original low-resolution (LR) images, such that the risk of semantic hallucination can be effectively mitigated. The generated images are refined subsequently in Stage 2 such that high-frequency details for SR quality can be restored via our customized and innovative guidance mechanism with a vision-language model (VLM) and a ControlNet, and a dynamic inference acceleration technique is applied to ensure efficiency. Extensive experimental results confirm that our proposed framework excels in perceptual quality-achieving top contrastive language-image pretraining-based image quality assessment (CLIP-IQA) scores-and in structural integrity while achieving robust performance. In particular, it enables reliable, high-fidelity SR for large-scale, real-world remote sensing pipelines by surpassing the conventional fidelity-hallucination tradeoff at practical inference speed. The source code is available at https://github.com/Bluear7878/Remote-Sensing-Vision-Language-Diffusion-Model

키워드

Remote sensing; Diffusion models; Measurement; Accuracy; Semantics; Image segmentation; Image restoration; Superresolution; Standards; Image synthesis; Denoising diffusion implicit model (DDIM) sampling; diffusion models; lightweight architecture; remote sensing imagery; super-resolution (SR); vision-language model (VLM)
제목
A Novel VLM-Guided Diffusion Model for Remote Sensing Image Super-Resolution
저자
Sung, Mingyu; Gong, Mu-Gyeong; Ham, Seung-Jae; Kim, Il-Min; Yun, Sangseok; Kang, Jae-Mo
DOI
10.1109/LGRS.2025.3608178
발행일
2025-09
유형
Article
저널명
IEEE Geoscience and Remote Sensing Letters
권
22