Towards refined unlearning: Utilizing the retain knowledge subspace and noise reduction

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

When large language models (LLMs) are adapted to a task, they are often fine-tuned on data that may contain sensitive information. This creates a need for effective unlearning methods to address privacy concerns. Among various unlearning methods, a recent line of work introduces a simple yet effective approach through arithmetic operations on task vectors. Despite their effectiveness, existing task vector-based unlearning methods in LLM still suffer from entangled representation. They also overlook the layer-wise roles of Transformers. In this paper, we propose ERASOR (unlEarning via low-RAnk projection in Subspace with ORthogonal bases), a new unlearning method in LLMs. We found that applying singular value decomposition directly on retain knowledge vectors performs better in unlearning than the previous approach with subspace over the whole model parameters. Additionally, we apply a noise reduction strategy using the explained variance, which preserves retain knowledge subspace in each layer by adjusting its dimensionality according to the layer-specific role. Extensive experiments across diverse unlearning benchmarks and model settings show that ERASOR achieves precise and robust unlearning while maintaining model performance on retained knowledge. Remarkably, it achieves about 63–85% unlearning effectiveness with only a 10% reduction in retained knowledge. The code for ERASOR is available at https://github.com/sun1187/ERASOR.

키워드

LLM unlearningMachine unlearning
제목
Towards refined unlearning: Utilizing the retain knowledge subspace and noise reduction
저자
Kim, EunSunChoi, YeoJunCho, Yoon-Sik
DOI
10.1016/j.knosys.2026.116284
발행일
2026-07
유형
Article
저널명
Knowledge-Based Systems
347