대규모 말뭉치에서의 ‘N+N’형 임시어 추출과 그 특성

Large-Scale Extraction of N+N Nonce Words in Korean and Their Characteristics
  • 정예은; 
  • 안진산; 
  • 이찬영

초록

This study constructs a large-scale corpus of 2.5 billion Korean tokens, balanced and refined for linguistic analysis, and employs SoyNLP’s LRNounExtractor_v2 to extract candidate N+N nonce words. The resulting dataset of 8,165 items, filtered by morpho-lexical criteria, has been released as the Korean Nonce Word Database (KNWD). Drawing on this lexical resource, we track nonce-word usage trends across time-series corpora from 2009 to 2023 and classify items into temporary and persistent types (rising, declining, or fluctuating). We then analyze the quantitative properties of these categories and explore socio-cognitive factors that appear to shape their distribution and use.

키워드

lexical resource; temporary types; nonce-word usage trends; persistent types; Korean Nonce Word Database (KNWD); N+N nonce words; 어휘 자원; 일시적 출현형; 임시어 사용 추이; 지속적 출현형; KNWD(Korean Nonce Word Database); ‘N+N’형 임시어
제목
대규모 말뭉치에서의 ‘N+N’형 임시어 추출과 그 특성
제목 (타언어)
Large-Scale Extraction of N+N Nonce Words in Korean and Their Characteristics
저자
정예은; 안진산; 이찬영
DOI
10.20405/kl.2025.11.109.415
발행일
2025-11
유형
Y
저널명
한국어학
권
109
페이지
415 ~ 457