역사 자료 형태소 분석 말뭉치 프로그램 개발 및 고도화

Development and Advancement of a Historical Data Morpheme Analysis Corpus Program
  • 장요한; 
  • 옥철영; 
  • 신승용; 
  • 박시온

초록

This paper aims to provide a detailed explanation of the development and enhancement process of the historical data morpheme analysis program ‘UTagger-Hunminjeongeum’. In this paper, we introduce the morpheme analysis algorithm of ‘UTagger- Hunminjeongeum (ver. 0.9)’ and outline the steps taken to improve it. Additionally, we present the structure of the tagging tool ‘UTagger- Hunminjeongeum TCM’, which was independently developed to reduce manual errors and save time. This tool is used to create a small-scale morpheme- analyzed corpus for training ‘UTagger-Hunminjeongeum’. The paper also discusses the enhancements made after the training phase, such as converting Chinese character tagging into Hangul and tagging intonation markers (bangjeom). The program has achieved an accuracy rate of nearly 90% for trained data and over 80% for untrained data, with an overall accuracy rate ranging from 85% to 90%. With continued development and the inclusion of more diverse data, the program is expected to become a versatile and highly accurate morphological analysis tool.

키워드

형태소; 형태소 분석 프로그램; 말뭉치; UTagger-훈민정음; UTagger-훈민정음 TCM; 역사 자료; morpheme; morpheme analysis program; corpus; tagging; historical data
제목
역사 자료 형태소 분석 말뭉치 프로그램 개발 및 고도화
제목 (타언어)
Development and Advancement of a Historical Data Morpheme Analysis Corpus Program
저자
장요한; 옥철영; 신승용; 박시온
DOI
10.29211/soli.2025.54..007
발행일
2025-03
유형
Y
저널명
언어와 정보 사회
권
54
페이지
191 ~ 219