GuardianAgentBench: Where Agents Fail and How to Guard Them
Abstract
As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical.
We present GuardianAgentBench (GABench), a benchmark of 580 scenarios across six domains evaluated on three production-ready frameworks: LangChain, LlamaIndex, and Vectara.
The benchmark incorporates rigorous multi-stage validation and five adversarial attack modes.
Experiments with six state-of-the-art models reveal that even the strongest configuration achieves only 74.8% overall accuracy and expose two distinct failure regimes: stronger models under-call required tools, while weaker models mis-select and over-call tools.
Performance degrades monotonically with both tool-set size and sequential turn depth, with long-horizon planning proving the steeper bottleneck.
Our guardrail implementation consistently outperforms system-prompt-based defenses across all models, recovering 19.9% of failures at a false positive rate of just 0.5%.
These results demonstrate that execution-time structural intervention improves safety without disrupting correct agent behavior.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요