상세 보기
Comparative Study of Multiclass Text Classification in Research Proposals Using Pretrained Language Models
- Lee, Eunchan;
- Lee, Changhyeon;
- Ahn, Sangtae
WEB OF SCIENCE
7SCOPUS
11초록
Recently, transformer-based pretrained language models have demonstrated stellar performance in natural language understanding (NLU) tasks. For example, bidirectional encoder representations from transformers (BERT) have achieved outstanding performance through masked self-supervised pretraining and transformer-based modeling. However, the original BERT may only be effective for English-based NLU tasks, whereas its effectiveness for other languages such as Korean is limited. Thus, the applicability of BERT-based language models pretrained in languages other than English to NLU tasks based on those languages must be investigated. In this study, we comparatively evaluated seven BERT-based pretrained language models and their expected applicability to Korean NLU tasks. We used the climate technology dataset, which is a Korean-based large text classification dataset, in research proposals involving 45 classes. We found that the BERT-based model pretrained on the most recent Korean corpus performed the best in terms of Korean-based multiclass text classification. This suggests the necessity of optimal pretraining for specific NLU tasks, particularly those in languages other than English.
키워드
- 제목
- Comparative Study of Multiclass Text Classification in Research Proposals Using Pretrained Language Models
- 저자
- Lee, Eunchan; Lee, Changhyeon; Ahn, Sangtae
- 발행일
- 2022-05
- 유형
- Article
- 저널명
- APPLIED SCIENCES-BASEL
- 권
- 12
- 호
- 9
- 언어
- ENG
- 출판사
- MDPI
- 발행국가
- 스위스
- ISSN
- E 2076-3417