High-speed and precise virtual try-on with two-stage semantic segmentation and a latent consistency model for optimized diffusion processes

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

3

초록

This work tests the hypothesis that the primary bottleneck for visual quality invirtual try-on (VTON) systems is the precision of input segmentation masks,rather than generative capability. VTON technology empowers users to dressdigital models in desired clothing items virtually. Conventional VTON modelsrely on segmentation models to isolate clothing regions and diffusion modelsto synthesize complete VTON images. This paper introduces high-speed andprecise VTON (HSP-VTON) as a framework that uniquely combines refinedtwo-stage semantic segmentation for enhanced accuracy with a latent consis-tency model to accelerate the diffusion-based image generation process. Thesynergistic integration of these components for VTON addresses critical chal-lenges in both precision and speed. Experimental results on the ATR datasetdemonstrate a 2.8% improvement in mean intersection over union comparedwith existing methods. Furthermore, HSP-VTON achieves superior perfor-mance on the VITON-HD dataset, outperforming state-of-the-art VTONmodels. The latent consistency model also reduces the number of inferencesteps, leading to substantial time savings without compromising image quality.

키워드

deep learning; diffusion; semantic segmentation; virtual try-on
제목
High-speed and precise virtual try-on with two-stage semantic segmentation and a latent consistency model for optimized diffusion processes
저자
Baek, Sangyeop; Lee, Jong Taek
DOI
10.4218/etrij.2024-0592
발행일
2025-10
유형
Article
저널명
ETRI Journal
권
47
호
5
페이지
881 ~ 892