A study on the data fusion using the national health screening data

  • 배세진; 
  • 김달호

초록

The exact combination of data collected from different objects is difficult. Data fusion is the statistical combination process of obtaining an integrated dataset using common variable. We consider three statistical techniques for data fusion: conditional mean matching using linear regression models, gamma regression using nonlinear regression models on two independent datasets, and a distance hot deck nonparametric approach based on the distance of each variable. The National Health Insurance Corporation's National Health Screening Data are used to compare the performance of three models.

키워드

Data fusion; distance hot deck; gamma regression; linear regression model; statistical matching
제목
A study on the data fusion using the national health screening data
저자
배세진; 김달호
DOI
10.7465/jkdi.2021.32.3.695
발행일
2021-05
유형
Y
저널명
한국데이터정보과학회지
권
32
호
3
페이지
695 ~ 703