Membership Inference Attacks for Unseen Classes
Abstract
A key tool in developing safe AI models is \emph{data auditing}, i.e., using statistical tools to determine whether harmful content may have been used in the training data of a black-box model.
Unfortunately, most \emph{membership inference attacks} (MIAs) used to perform this type of auditing themselves assume \emph{access} to examples of harmful content from the same distribution as the query data.
In real-world auditing scenarios, auditors often face legal and ethical restrictions preventing them from accessing a representative set of samples of harmful content to train MIA models effectively.
We abstract and formalize this setting into a new data access model, the ``unseen class'' setting, and show that the state of the art MIAs fail due to the lack of access to the full target distribution.
We show in this setting, \emph{quantile regression attacks} outperform approaches typically considered to be SoTA.
We demonstrate this both empirically and theoretically, showing that quantile regression attacks achieve up to \textbf{11$\times$ the TPR} of shadow model-based approaches in practice, and providing a theoretical model that outlines the generalization properties required for this approach to succeed.
Our work identifies an important failure mode in existing MIAs and provides a cautionary tale for practitioners that aim to directly use existing tools for real-world applications of AI safety.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요