상세 보기
An investigation on synthetic data generation from original incomplete data
- Seo, Yun-Beom;
- Kim, Young Min
WEB OF SCIENCE
0SCOPUS
0초록
The generation of synthetic data is widely recognized as an effective strategy for protecting sensitive information in original datasets while maintaining their analytical utility. However, when the original data includes missing values, generating synthetic data directly from incomplete datasets can cause considerable bias, resulting in substantial deviations from the underlying data structure. Moreover, existing methods offer limited statistical inference frameworks for handling missing values when imputation is conducted after the synthetic data generation process. To address these challenges, this study proposes handling missing data prior to synthetic data generation. Specifically, we investigate two methodologies for producing fully imputed synthetic datasets and evaluate their performance through extensive simulation studies. We also examine their practical applicability in real-world data analysis settings, highlighting their advantages in preserving both data utility and inferential validity.
키워드
- 제목
- An investigation on synthetic data generation from original incomplete data
- 저자
- Seo, Yun-Beom; Kim, Young Min
- 발행일
- 2025-04-22
- 유형
- Article; Early Access
- 언어
- ENG
- 출판사
- TAYLOR & FRANCIS INC
- 발행국가
- 미국
- ISSN
- E 1532-4141
P 0361-0918