근접 정책 최적화 기반 샌드박스 게임의 강화학습 환경 설계와 구현

Design and Implementation of Reinforcement Learning Environments in Sandbox Games with Proximal Policy Optimization

초록

Applying reinforcement learning(RL) in sandbox games without official APIs presents significant challenges due to limited state observability and unstable visual feedback. To address this, we propose a dual-module RL framework integrating multimodal perception and temporal decision-making for non-API environments, demonstrated using complex boss battles in the sandbox game Terraria. Our framework consists of two distinct modules: an external perception module (Worker), utilizing YOLO- based real-time object detection and memory parsing to extract crucial game states (player and boss positions, health points, and distances), and an internal decision-making module (Learner), employing a Convolutional Neural Network(CNN)-Gated Transformer-XL (GTrXL) model combined with Proximal Policy Optimization (PPO) for robust policy learning. Experiments conducted in visually cluttered, temporally extended boss-fight scenarios show that our architecture effectively handles multimodal cognition and long-term sequence prediction tasks, significantly outperforming traditional RL approaches. The proposed dual-module structure not only demonstrates its effectiveness and stability within Terraria but also suggests broad applicability to other complex, non-standard environments where APIs are unavailable. Our results validate the potential for bridging the gap between RL research in controlled benchmarks and practical deployment scenarios.

키워드

Reinforcement Learning; Sandbox Game; Non-API Environment; YOLO; GTrXL; Proximal Policy Optimization
제목
근접 정책 최적화 기반 샌드박스 게임의 강화학습 환경 설계와 구현
제목 (타언어)
Design and Implementation of Reinforcement Learning Environments in Sandbox Games with Proximal Policy Optimization
저자
김일; 조형주
DOI
10.9717/kmms.2025.28.7.893
발행일
2025-07
유형
Y
저널명
멀티미디어학회논문지
권
28
호
7
페이지
893 ~ 912