Perturbation is All You Need for Extrapolating Language Models
Abstract
This paper develops a statistical theory of extrapolation for large language models, by reinterpreting them through pre-post-additive noise models.
In contrast to the standard autoregressive next-token prediction based on an exact prefix, we introduce a perturbation-based procedure that first transforms the prefix into a semantic neighbour and then conditions on this perturbed variant for next-token prediction.
This yields a hierarchical model with a pre-post-additive noise structure.
Within this framework, we develop a rigorous theory of extrapolability, namely, the capacity of a model class to make reliable predictions for token sequences that lie outside the empirical support of the training corpus, by establishing five properties of the proposed procedure: adaptivity, contractivity, robustness, extrapolability, and double robustness.
We evaluate the finite sample performance of the proposed procedure using both synthetic and real world language data.
Results show that the proposed method consistently improves out-of-support prediction while maintaining competitive in-support performance, demonstrating that perturbation offers a practical route to language modelling.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요