Mitigating Resource Contention for Responsive On-device Machine Learning Inferences

  • Kim, Minsung
  • Lee, Jihoon
  • Chou, Seongjin
  • Chung, Whisoo
  • Kim, Inwoo
  • ... Kim, Hyosu
  • 외 4명
Citations

SCOPUS

0

초록

On-device machine learning applications are increasingly deployed in dynamic and open system environments, where resource availability fluctuates unpredictably. This variability, coupled with limited computing resources, poses significant challenges in achieving high responsiveness. Existing on-device machine learning frameworks typically rely on static and coarse-grained resource allocation, leading to performance degradation under resource contention. To address this, we propose FlexOn, a novel framework that combines fine-grained model segmentation and dynamic resource selection to rapidly adapt to highly dynamic runtime conditions and effectively mitigate unpredictable resource contention. A prototype built on LiteRT demonstrates significant improvements in both average and tail latencies of up to 54% and 58%, respectively, across three different embedded platforms under dynamically varying resource availability. To the best of our knowledge, this is the first work that addresses the resource contention in open embedded systems for better machine learning inference responsiveness.

제목
Mitigating Resource Contention for Responsive On-device Machine Learning Inferences
저자
Kim, MinsungLee, JihoonChou, SeongjinChung, WhisooKim, InwooKang, WoosungKim, HyosuOh, SangeunChwa, Hoon SungLee, Kilho
DOI
10.1109/ICCAD66269.2025.11240707
발행일
2025
유형
Conference Paper
저널명
IEEE/ACM International Conference on Computer-Aided Design, Digest of Technical Papers, ICCAD