상세 보기
다국어 뉴스에서 추출한 식별적 색인어 기반 감염병 위험지수 예측
- 김기후;
- 장종원;
- 장길진
초록
In this paper, we propose a COVID-19 risk level prediction system using Discriminative Term Frequency-Inverse Document Frequency (D-TF-IDF) keyword extraction from multilingual news articles, in addition to the existing natural language-based methods for predicting the increase in COVID-19 confirmed cases. D-TF-IDF extracts keywords by excluding overlaps among key terms appearing in multilingual news, which are then used as inputs for models predicting the increase in COVID-19 confirmed cases and risk levels. The proposed keyword extraction method improves the classification performance metric, micro F1-score, in the COVID-19 confirmed case increase and risk level prediction system. Random Over Sampling, a data augmentation technique, is used to address the imbalance issue in keyword data representing risk levels. Additionally, to enhance classification performance, the system employs the classifier that shows the best performance among Random Forest, Balanced Random Forest, Support Vector Machine, and Gaussian Process Classifiers for predicting COVID-19 risk levels.
키워드
- 제목
- 다국어 뉴스에서 추출한 식별적 색인어 기반 감염병 위험지수 예측
- 제목 (타언어)
- Prediction of Infectious Disease Risk Index using Discriminative Keyword Extraction from Multi-lingual News Articles
- 저자
- 김기후; 장종원; 장길진
- 발행일
- 2025-05
- 유형
- Y
- 저널명
- 전자공학회논문지
- 권
- 62
- 호
- 5
- 페이지
- 35 ~ 45
- 언어
- KOR
- 출판사
- 대한전자공학회
- 발행국가
- 대한민국
- 분량
- 11 페이지
- ISSN
- E 2288-159X
P 2287-5026