재난 상황에서의 생성형 언어 모델에 대한 백도어 공격 시나리오와 실험적 분석

Backdoor Attacks on Generative Language Models in Disaster Situations: Attack Scenarios and Empirical Analysis

초록

With the remarkable advancement of artificial intelligence, AI models have begun to be used across various fields, while simultaneouslyrevealing security vulnerabilities within AI itself. This paper assumes a situation in which a backdoor is inserted into a generative modelused in disaster situations through data poisoning and transfer learning. By means of this backdoor, the model produces normal outputsfor general inputs, yet deliberately yields manipulated outputs for inputs containing a specific trigger, thereby potentially causing socialconfusion. We experimentally demonstrate this possibility and emphasize its associated risks.

키워드

생성형 모델; 재난; 백도어; 데이터 오염; 트리거; Gnenrative model; Disaster; Backdoor; Data poisoning; Trigger
제목
재난 상황에서의 생성형 언어 모델에 대한 백도어 공격 시나리오와 실험적 분석
제목 (타언어)
Backdoor Attacks on Generative Language Models in Disaster Situations: Attack Scenarios and Empirical Analysis
저자
강종영; 이지연
DOI
10.3745/TKIPS.2025.14.7.533
발행일
2025-07
유형
Y
저널명
정보처리학회 논문지
권
14
호
7
페이지
533 ~ 541