상세 보기
대규모 로그 데이터의 실시간 분석을 위한 아파치 스파크 성능 최적화
- 전승훈;
- 박동철
초록
In modern society, big data serves as a critical resource across various industries, emphasizing the growing importance of distributed data processing platforms capable of handling large-scale data efficiently. Among them, Apache Spark is widely adopted for its high processing performance and scalable architecture in diverse data analysis environments. However, when operated under default configurations, Spark often exhibits performance bottlenecks due to excessive processing times or inefficient resource utilization, especially under varying workload conditions. This study, conducted in collaboration with a mid-sized domestic IT company, investigates the performance optimization of Spark by performing experiments on a commercial-grade Spark server. Specifically, we analyze the impact of CPU core count and memory allocation on performance through a series of controlled experiments. By systematically evaluating various configuration combinations, we identify major performance bottlenecks in Spark and propose optimized resource allocation strategies to address them. All experiments are conducted using real-world, large-scale security monitoring logs provided by the partner company, rather than synthetic data. The optimized configuration achieves, on average, a 9.8-fold improvement in performance compared to the default setting, demonstrating the effectiveness of the proposed approach.
키워드
- 제목
- 대규모 로그 데이터의 실시간 분석을 위한 아파치 스파크 성능 최적화
- 제목 (타언어)
- Performance Optimization of Apache Spark for Real-Time Analysis of Massive Log Data
- 저자
- 전승훈; 박동철
- 발행일
- 2025-08
- 유형
- Y
- 권
- 13
- 호
- 4
- 페이지
- 72 ~ 91