Bigger Is Safer: Provable Robustness in In-Context Learning Scales with Capacity
Abstract
In-context learning (ICL) allows large language models to adapt to new tasks from a few examples without updating their parameters.
Existing theories explain ICL by assuming the test task distribution matches pretraining -- an assumption that breaks down under adversarial distribution shifts.
We introduce a distributionally robust meta-learning framework that provides worst-case guarantees for ICL under Wasserstein-based distribution shifts.
Focusing on linear self-attention Transformers, we derive a non-asymptotic bound connecting adversarial perturbation strength ($\rho$), model capacity ($m$), and the number of in-context examples ($N$).
The analysis reveals that the maximum safe perturbation radius scales as $\rho_{\max} \propto \sqrt{m}$, while maintaining performance under adversarial shift requires additional in-context examples with $N_\rho - N_0 \propto \rho^2$.
Experiments on synthetic tasks confirm these scaling laws, and experiments on 21 real pretrained models (0.1B--7B parameters, 5 families) provide qualitative evidence consistent with the theory's predictions, while revealing that ICL capability is a prerequisite for robustness.
These findings advance the theoretical understanding of ICL under adversarial conditions and formalize the sense in which larger models are safer under distributional shift.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요