말뭉치언어학의 관점에서 본 <모두의 말뭉치> 구축 현황과 쟁점

Everyone’s Corpus: Corpus Linguistics Issues
  • 안진산; 
  • 남길임; 
  • 강현아; 
  • 이준

초록

This study examined the status of the National Institute of Korean Language’s Everyone’s Corpus and explored key issues in and future directions for building text data in the era of generative artificial intelligence (“AI”). To this end, six analytical criteria were established: purpose of construction, text genre, level and type of processing, multilingual parallelism, and data type. As of March 31, 2025, Everyone’s Corpus includes 82 corpora constituted by 2.7 billion words and various web-based corpus tools. While the corpus is meaningful for its large size and can be used to train and evaluate AI models, over 85% of its content consists of written language, indicating a clear imbalance in text genre representation. Thus, this paper discusses the relevance of traditional balanced corpora such as COCA as benchmarks for generative language models and proposes the construction of inclusive corpora that reflect speaker subjectivity and guarantee linguistic rights for visually and hearing-impaired people.

키워드

모두의 말뭉치; 텍스트 데이터; 생성형 AI; 벤치마크; 주관성; Everyone’s Corpus; text data; generative AI; benchmark; subjectivity
제목
말뭉치언어학의 관점에서 본 <모두의 말뭉치> 구축 현황과 쟁점
제목 (타언어)
Everyone’s Corpus: Corpus Linguistics Issues
저자
안진산; 남길임; 강현아; 이준
DOI
10.22557/HG.2025.6.86.2.489
발행일
2025-06
유형
Y
저널명
한글
권
86
호
2
페이지
489 ~ 525