2차원 술래잡기 게임을 위한 다중 에이전트 강화학습의 보상 정책 비교

Comparing Reward Policies in Multi-Agent Reinforcement Learning for 2D Tag Game

초록

This paper proposes and compares various reward policies for implementing a 2D tag game using multi-agent deep reinforcement learning (MADRL) techniques. Tag is a game where a chaser chases and tags a runner, with the runner aiming to avoid being tagged for as long as possible. The game can be customized with variations such as setting a tag limit, adjusting the environment with obstacles, or altering the number of players on each team. Designing an effective reward policy for an agent in a tag game to maximize survival time can be challenging and involve a substantial amount of experimentation. This paper investigates the impact of various reward policies on agent survival time and presents a methodology for identifying the optimal reward policy. Reinforcement learning experiments were conducted on four distinct stages, each with a unique obstacle layout, where agents were trained to maximize survival time while being chased under predefined rules. Eleven different reward policies were applied to evaluate their effectiveness, each focused on avoiding collisions with walls, obstacles, and chasers. After determining the optimal reward policy that maximizes agent survival time through experiments, a tag game was implemented where human players chase agents trained under this policy.

키워드

강화학습; 학습 파라미터; 술래잡기 게임; 보상 정책; Reinforcement Learning; Learning Hyperparameter; Tag Game; Reward Policy
제목
2차원 술래잡기 게임을 위한 다중 에이전트 강화학습의 보상 정책 비교
제목 (타언어)
Comparing Reward Policies in Multi-Agent Reinforcement Learning for 2D Tag Game
저자
정상훈; 김구진
DOI
10.47116/apjcri.2025.03.28
발행일
2025-03
유형
Y
저널명
아시아태평양융합연구교류논문지
권
11
호
3
페이지
427 ~ 436