대규모 로그 데이터의 실시간 분석을 위한 아파치 스파크 성능 최적화

Performance Optimization of Apache Spark for Real-Time Analysis of Massive Log Data

초록

In modern society, big data serves as a critical resource across various industries, emphasizing the growing importance of distributed data processing platforms capable of handling large-scale data efficiently. Among them, Apache Spark is widely adopted for its high processing performance and scalable architecture in diverse data analysis environments. However, when operated under default configurations, Spark often exhibits performance bottlenecks due to excessive processing times or inefficient resource utilization, especially under varying workload conditions. This study, conducted in collaboration with a mid-sized domestic IT company, investigates the performance optimization of Spark by performing experiments on a commercial-grade Spark server. Specifically, we analyze the impact of CPU core count and memory allocation on performance through a series of controlled experiments. By systematically evaluating various configuration combinations, we identify major performance bottlenecks in Spark and propose optimized resource allocation strategies to address them. All experiments are conducted using real-world, large-scale security monitoring logs provided by the partner company, rather than synthetic data. The optimized configuration achieves, on average, a 9.8-fold improvement in performance compared to the default setting, demonstrating the effectiveness of the proposed approach.

키워드

Apache SparkPerformance OptimizationResource ManagementBig DataDistributed Systems
제목
대규모 로그 데이터의 실시간 분석을 위한 아파치 스파크 성능 최적화
제목 (타언어)
Performance Optimization of Apache Spark for Real-Time Analysis of Massive Log Data
저자
전승훈박동철
DOI
10.23023/JPT.2025.13.4.072
발행일
2025-08
유형
Y
저널명
Journal of Platform Technology
13
4
페이지
72 ~ 91