Survey on joint compression and fine-tuning of large language models: Methods, toolchains, and open challenges under resource constraints

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Improving the efficiency of large-scale language models without losing task performance remains a central de ployment challenge. This survey investigates this problem by analyzing a review corpus of 165 studies published from the 2010s through 2025 and summarizes representative methods for compressing and fine-tuning large lan guage models. First, the fundamental building blocks of these models and the standard benchmarks used to assess their performance are introduced. Next, the focus shifts to the principal tasks involved in model compression. Methods such as pruning and quantization serve as redundancy reduction mechanisms that remove redundant parameters while maintaining the core capability of the model. In parallel, this survey reviews fine-tuning, a process through which models gain task-specific capabilities through human feedback or through highly targeted and computationally efficient updates. The interaction between compression and fine-tuning is treated as the main focus of the survey. Instead of treating model compression and knowledge transfer as separate objectives, this survey illustrates how these elements can be integrated into a single pipeline, thereby reducing computa tional costs and shortening training time. Finally, this survey provides a list of widely used open-source tools and outlines open challenges for future LLM compression and adaptation research.

키워드

Large language modelsModel compressionFine-tuningJoint compression and fine-tuning
제목
Survey on joint compression and fine-tuning of large language models: Methods, toolchains, and open challenges under resource constraints
저자
Kim, JihyeonLee, HanyongLee, Jaesung
DOI
10.1016/j.neucom.2026.134332
발행일
2026-10
유형
Article
저널명
Neurocomputing
698