상세 보기
이기종 장치의 빅데이터 처리를 위한 분산형 기계 학습 프레임워크
- 노유수;
- 윤주영;
- 서영균
초록
As the need to reliably collect diverse types of data streams from heterogeneous hardware devices and leverage them for machine learning model training increases, building environments that can support these capabilities is becoming increasingly important. With data's growing complexity and scale, traditional machine learning frameworks that focus on processing structured data or single-source inputs are reaching their limits. There is a demand for architectures that can efficiently process and learn from dynamically incoming heterogeneous data streams. In this paper, we propose FALCON, a high-performance, scalable distributed computing platform designed to effectively collect large-scale streaming data and connect it seamlessly to machine learning model training. FALCON utilizes Apache Flume, Kafka, and a schema registry to ensure reliable collection and consistent transmission and management of data from heterogeneous devices, and integrates a PyTorch-based distributed parallel learning architecture to enhance the scalability and efficiency of both model training and real-time analysis. Performance evaluation using real-world data demonstrates that FALCON reduces transfer latency by approximately 3.7× compared to a single-worker baseline, achieves up to 94% reduction in storage footprint versus local loading, and improves data-loading speed by up to 17.4% over JSON serialization via the proposed serialization and compression configurations.
키워드
- 제목
- 이기종 장치의 빅데이터 처리를 위한 분산형 기계 학습 프레임워크
- 제목 (타언어)
- A Distributed Machine Learning Framework for Processing Big Data from Heterogeneous Devices
- 저자
- 노유수; 윤주영; 서영균
- 발행일
- 2025-08
- 유형
- Y
- 저널명
- 데이타베이스연구
- 권
- 41
- 호
- 2
- 페이지
- 74 ~ 87
- 언어
- KOR
- 출판사
- 한국정보과학회
- 발행국가
- 대한민국
- 분량
- 14 페이지
- ISSN
- P 1598-9798