시각언어모델 기반 맥락 추출과 한국어 대규모 언어모델 파인튜닝을 통한 여행 블로그 콘텐츠 자동 생성

VLM-Based Context Extraction and Fine-Tuning of Korean LLMs for Automatic Travel Blog Content Generation
  • 임동훈; 
  • 한승수; 
  • 은의찬; 
  • 김동영; 
  • 최장훈

초록

In this paper, we propose an automated system for generating Korean travel blog content by integrating Vision-Language Model (VLM)–based context extraction, fine-tuned Korean language model, and a text-to-image (T2I) generation model via prompt engineering. Using the Qwen2-VL model, we analyze travel photos to extract visual context and produce emotionally nuanced captions. We then leverage a large-scale crawled corpus of real-world travel blogs to fine-tune HyperCLOVA X, enabling it to create natural, storytelling-oriented blog text. In addition, we employ a travel-specific prompt engineering approach with DALL-E to generate custom postcards and stamp images, providing users with intuitive and creative value for their travel records. A robust system prompt minimizes hallucinations while preserving expressive writing. User surveys indicate that the fine-tuned model is, on average, 60.9% more specialized in travel-related content. These findings demonstrate that the proposed approach can significantly reduce manual effort while producing high-quality travel blog content, suggesting new possibilities for AI-based content creation in the tourism domain.

키워드

시각-언어 모델(VLM); 파인 튜닝; 프롬프트 엔지니어링; 텍스트-이미지(T2I) 생성; 여행 블로그; Vision-Language Model; Fine-Tuning; Prompt Engineering; T2I Generation; Travel Blog
제목
시각언어모델 기반 맥락 추출과 한국어 대규모 언어모델 파인튜닝을 통한 여행 블로그 콘텐츠 자동 생성
제목 (타언어)
VLM-Based Context Extraction and Fine-Tuning of Korean LLMs for Automatic Travel Blog Content Generation
저자
임동훈; 한승수; 은의찬; 김동영; 최장훈
DOI
10.9708/jksci.2025.30.06.077
발행일
2025-06
유형
Y
저널명
한국컴퓨터정보학회논문지
권
30
호
6
페이지
77 ~ 90