SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
Abstract
Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is identified as a key driver of Loss of Control (LoC) risk.
In this work, we introduce SysAdmin, a benchmark that positions frontier language models as autonomous system administrators in a high-fidelity Linux sandbox to measure power-seeking propensity across five dimensions: self-preservation, increasing autonomy, resource acquisition, environment modification, and strategic concealment.
We evaluated seven frontier models across four experimental conditions in a total of 2800 tasks.
After bias correction using human-annotated calibration data, corrected power-seeking estimates ranged from 0 to about 5 percent per model.
We also conducted a positive control with explicit power-seeking prompts that achieved 100% detection, validating measurement sensitivity.
Our findings indicate current frontier models exhibit minimal spontaneous power-seeking in naturalistic system administration contexts, though model-specific failure modes suggest evaluations must test diverse misalignment patterns.
Nevertheless, we discovered other more pronounced failure modes (than power-seeking) such as specification gaming and resistance to goal modification.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요