상세 보기
초록
This study examined the status of the National Institute of Korean Language’s Everyone’s Corpus and explored key issues in and future directions for building text data in the era of generative artificial intelligence (“AI”). To this end, six analytical criteria were established: purpose of construction, text genre, level and type of processing, multilingual parallelism, and data type. As of March 31, 2025, Everyone’s Corpus includes 82 corpora constituted by 2.7 billion words and various web-based corpus tools. While the corpus is meaningful for its large size and can be used to train and evaluate AI models, over 85% of its content consists of written language, indicating a clear imbalance in text genre representation. Thus, this paper discusses the relevance of traditional balanced corpora such as COCA as benchmarks for generative language models and proposes the construction of inclusive corpora that reflect speaker subjectivity and guarantee linguistic rights for visually and hearing-impaired people.
키워드
- 제목
- 말뭉치언어학의 관점에서 본 <모두의 말뭉치> 구축 현황과 쟁점
- 제목 (타언어)
- Everyone’s Corpus: Corpus Linguistics Issues
- 저자
- 안진산; 남길임; 강현아; 이준
- 발행일
- 2025-06
- 유형
- Y
- 저널명
- 한글
- 권
- 86
- 호
- 2
- 페이지
- 489 ~ 525
- 언어
- KOR
- 출판사
- 한글학회
- 발행국가
- 대한민국
- 분량
- 37 페이지
- ISSN
- E 2733-8932
P 1225-0449