상세 보기
초록
This study explores the productivity and accuracy of neologism detection, focusing on the most recent Korean neologism dataset, Neologisms 2023. It aims to re-evaluate the corpus-based semi-automatic extraction method, which has been the dominant approach in neologism detection, and assess the potential of integrating human expert judgment with Large Language Models (LLMs). Since the introduction of large-scale corpora into neologism research, both Korean and global approaches have widely adopted semi-automatic extraction based on corpus data as the methodological norm. Targeting Korean neologisms, this study conducts two empirical experiments to examine the effectiveness and precision of current detection methods. First, by analyzing the performance of five widely used unknown word extractors using the Neologisms 2023 list and a 78-million-token newspaper corpus, the study assesses the procedural transparency and replicability of corpus-based neologism extraction, while also identifying the strengths and limitations of noun-based unknown word extraction. Second, the study explores the potential for human–machine collaboration in neologism detection by applying Retrieval-Augmented Generation (RAG)-based LLMs to the candidate list previously evaluated by human experts. This approach provides insight into how human linguistic intuition can continue to play a central role—especially in challenging areas such as proper noun and compound word recognition—within a triangulated framework involving corpus data, human judgment, and LLMs.
키워드
- 제목
- 말뭉치, LLMs, 인간 전문가의 협업을 통한 한국어 신어의 탐지 -<신어 2023>의 신어 목록을 중심으로-
- 제목 (타언어)
- A Triangulational Approach to Korean Neologism Detection: Corpora, Large Language Models, and Human Expertise in Neologisms 2023
- 저자
- 남길임; 안진산; 이수진
- 발행일
- 2025-08
- 유형
- Y
- 저널명
- 한국어학
- 권
- 108
- 페이지
- 205 ~ 238
- 언어
- KOR
- 출판사
- 한국어학회
- 발행국가
- 대한민국
- 분량
- 34 페이지
- ISSN
- E 2734-0082
P 1226-9123