상세 보기
Mitigating Resource Contention for Responsive On-device Machine Learning Inferences
- Kim, Minsung;
- Lee, Jihoon;
- Chou, Seongjin;
- Chung, Whisoo;
- Kim, Inwoo;
- ... Kim, Hyosu;
- 외 4명
SCOPUS
0초록
On-device machine learning applications are increasingly deployed in dynamic and open system environments, where resource availability fluctuates unpredictably. This variability, coupled with limited computing resources, poses significant challenges in achieving high responsiveness. Existing on-device machine learning frameworks typically rely on static and coarse-grained resource allocation, leading to performance degradation under resource contention. To address this, we propose FlexOn, a novel framework that combines fine-grained model segmentation and dynamic resource selection to rapidly adapt to highly dynamic runtime conditions and effectively mitigate unpredictable resource contention. A prototype built on LiteRT demonstrates significant improvements in both average and tail latencies of up to 54% and 58%, respectively, across three different embedded platforms under dynamically varying resource availability. To the best of our knowledge, this is the first work that addresses the resource contention in open embedded systems for better machine learning inference responsiveness.
- 제목
- Mitigating Resource Contention for Responsive On-device Machine Learning Inferences
- 저자
- Kim, Minsung; Lee, Jihoon; Chou, Seongjin; Chung, Whisoo; Kim, Inwoo; Kang, Woosung; Kim, Hyosu; Oh, Sangeun; Chwa, Hoon Sung; Lee, Kilho
- 발행일
- 2025
- 유형
- Conference Paper
- 저널명
- IEEE/ACM International Conference on Computer-Aided Design, Digest of Technical Papers, ICCAD