Response drift across frontier large language models
Abstract
All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet the magnitude and structure of this drift remain uncharacterised by systematic human evaluation.
Here we report a fully crossed evaluation in which 47 geographically diverse participants each assessed all 62 multidomain questions across ten frontier LLMs under blinded conditions, yielding 29,140 independent assessments.
Every model drifts, but drift magnitude varies substantially: eight models converge on a statistically indistinguishable ceiling (78-81% deviation), while two achieve lower deviation (47-49%).
Drift profiles differ across six domains and 62 questions, with pairwise correlations among ceiling models exceeding r = 0.85.
Automated similarity metrics explain less than 2% of variance in human judgements.
These findings reveal that response drift is universal across frontier LLMs, domain- and question-dependent in structure, and accessible only through human-centred evaluation.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요