An investigation on synthetic data generation from original incomplete data

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

The generation of synthetic data is widely recognized as an effective strategy for protecting sensitive information in original datasets while maintaining their analytical utility. However, when the original data includes missing values, generating synthetic data directly from incomplete datasets can cause considerable bias, resulting in substantial deviations from the underlying data structure. Moreover, existing methods offer limited statistical inference frameworks for handling missing values when imputation is conducted after the synthetic data generation process. To address these challenges, this study proposes handling missing data prior to synthetic data generation. Specifically, we investigate two methodologies for producing fully imputed synthetic datasets and evaluate their performance through extensive simulation studies. We also examine their practical applicability in real-world data analysis settings, highlighting their advantages in preserving both data utility and inferential validity.

키워드

Incomplete data; Nonparametric Bayesian; Sequential conditional distribution; Synthetic data generation; Two-stage synthetic data generation; MULTIPLE IMPUTATION; MISSING-DATA
제목
An investigation on synthetic data generation from original incomplete data
저자
Seo, Yun-Beom; Kim, Young Min
DOI
10.1080/03610918.2025.2496777
발행일
2025-04-22
유형
Article; Early Access
저널명
Communications in Statistics Part B: Simulation and Computation