SLICE: An Efficient and Tuning-Free Keyframe Sampling Framework for Long-Form Video Understanding

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Efficiently distilling long videos into semantically significant keyframes is essential for effective Video Question Answering (VideoQA) using Large Multimodal Models (LMMs). However, existing sampling strategies suffer from hyperparameter brittleness, relying heavily on manual hyperparameter tuning or dataset-specific calibration. To address these limitations, we introduce SLICE (Semantic Length-Independent Content Extractor), a parameter-free framework. The core mechanism of SLICE introduces logarithmic scale-adaptive smoothing ( σ=ln(N) ) based on the video length N to suppress visual noise, and logically allocates a limited token budget through Information Density Partitioning. By partitioning videos based on actual Semantic Energy Mass rather than uniform time intervals, SLICE concentrates frames within complex event segments and aggressively omits static backgrounds, thereby securing both temporal diversity and semantic relevance. Experimental results demonstrate that SLICE significantly outperforms hyperparameter-sensitive state-of-the-art (SOTA) methods. Notably, on sparse event localization tasks, SLICE achieves an outstanding visual F1 score of up to 80.2% on HAYSTACK-LVBENCH and dramatically improves the temporal F1 score on HAYSTACK-EGO4D by nearly 3 times compared to previous baselines, all while reducing computational overhead. These findings prove that SLICE is a highly scalable, out-of-the-box solution for real-world LMM deployments.

키워드

Keyframe SamplingLarge Multimodal ModelsVideo Question Answering
제목
SLICE: An Efficient and Tuning-Free Keyframe Sampling Framework for Long-Form Video Understanding
저자
Han, SungjinVu, ThangKim, Junyeong
DOI
10.1109/ACCESS.2026.3680314
발행일
2026
유형
Article
저널명
IEEE Access
14
페이지
51576 ~ 51588