상세 보기
SLICE: An Efficient and Tuning-Free Keyframe Sampling Framework for Long-Form Video Understanding
- Han, Sungjin;
- Vu, Thang;
- Kim, Junyeong
WEB OF SCIENCE
0SCOPUS
0초록
Efficiently distilling long videos into semantically significant keyframes is essential for effective Video Question Answering (VideoQA) using Large Multimodal Models (LMMs). However, existing sampling strategies suffer from hyperparameter brittleness, relying heavily on manual hyperparameter tuning or dataset-specific calibration. To address these limitations, we introduce SLICE (Semantic Length-Independent Content Extractor), a parameter-free framework. The core mechanism of SLICE introduces logarithmic scale-adaptive smoothing ( σ=ln(N) ) based on the video length N to suppress visual noise, and logically allocates a limited token budget through Information Density Partitioning. By partitioning videos based on actual Semantic Energy Mass rather than uniform time intervals, SLICE concentrates frames within complex event segments and aggressively omits static backgrounds, thereby securing both temporal diversity and semantic relevance. Experimental results demonstrate that SLICE significantly outperforms hyperparameter-sensitive state-of-the-art (SOTA) methods. Notably, on sparse event localization tasks, SLICE achieves an outstanding visual F1 score of up to 80.2% on HAYSTACK-LVBENCH and dramatically improves the temporal F1 score on HAYSTACK-EGO4D by nearly 3 times compared to previous baselines, all while reducing computational overhead. These findings prove that SLICE is a highly scalable, out-of-the-box solution for real-world LMM deployments.
키워드
- 제목
- SLICE: An Efficient and Tuning-Free Keyframe Sampling Framework for Long-Form Video Understanding
- 저자
- Han, Sungjin; Vu, Thang; Kim, Junyeong
- 발행일
- 2026
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 51576 ~ 51588