Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning
Abstract
Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations.
To evaluate unlearning effectiveness in multimodal large language models (MLLMs), prior works fine-tune models on fictitious identities, simulating unlearning requests on subsets of these IDs, which are typically uniformly distributed.
However, in realistic scenarios, people from different demographic groups may request to be unlearned at different frequencies, potentially altering the model's internal beliefs for these groups and leading to biased behaviors.
To fill this gap, we propose FAIRGET, the first Visual Question Answering benchmark that evaluates unlearning under unbalanced, realistic, forget requests.
These requests are designed to simulate multiple realistic scenarios, ranging from simple to challenging settings, that lead to biased unlearned models if fairness is not accounted for.
Additionally, we propose FAUN, the first unlearning algorithm for MLLMs that forgets unlearning data while preserving model fairness.
FAUN exploits a bias-aware activation steering mechanism to unlearn identities while accounting for the unbalanced nature of the forget data.
Experiments on FAIRGET and the established FIUBench demonstrate our method's superiority both in unlearning quality and fairness.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요