상세 보기
초록
This paper aims to provide a detailed explanation of the development and enhancement process of the historical data morpheme analysis program ‘UTagger-Hunminjeongeum’. In this paper, we introduce the morpheme analysis algorithm of ‘UTagger- Hunminjeongeum (ver. 0.9)’ and outline the steps taken to improve it. Additionally, we present the structure of the tagging tool ‘UTagger- Hunminjeongeum TCM’, which was independently developed to reduce manual errors and save time. This tool is used to create a small-scale morpheme- analyzed corpus for training ‘UTagger-Hunminjeongeum’. The paper also discusses the enhancements made after the training phase, such as converting Chinese character tagging into Hangul and tagging intonation markers (bangjeom). The program has achieved an accuracy rate of nearly 90% for trained data and over 80% for untrained data, with an overall accuracy rate ranging from 85% to 90%. With continued development and the inclusion of more diverse data, the program is expected to become a versatile and highly accurate morphological analysis tool.
키워드
- 제목
- 역사 자료 형태소 분석 말뭉치 프로그램 개발 및 고도화
- 제목 (타언어)
- Development and Advancement of a Historical Data Morpheme Analysis Corpus Program
- 저자
- 장요한; 옥철영; 신승용; 박시온
- 발행일
- 2025-03
- 유형
- Y
- 저널명
- 언어와 정보 사회
- 권
- 54
- 페이지
- 191 ~ 219
- 언어
- KOR
- 출판사
- 서강대학교 언어정보연구소
- 발행국가
- 대한민국
- 분량
- 29 페이지
- ISSN
- E 2713-6817
P 1598-1886