미디어 커버리지1건1개 미디어
학술
기타

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

arXiv CS.AI
CC BY
이 매체는 공공·자유 라이선스로 본문을 직접 표시합니다.

Abstract

Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential.

We present NEXUS (Neural EXecution Utility and Safety), a structured-plan safety monitor that applies a formal intervention policy to select among four actions: allow, block, request confirmation, or request revision.

NEXUS combines deterministic safety rules, argument-level inspection, and a calibrated logistic-regression risk score for graded escalation.

On a 128-instance synthetic benchmark, NEXUS achieves an F1 score of 0.949 and a 4-class intervention accuracy of 0.6406, outperforming rule-only intervention selection by 27.3 percentage points.

It also improves over rule-only on R-Judge (F1 = 0.861 vs.

0.849), matches rule-only on AgentHarm due to threat-model limits, and achieves 0% ASR at 99% control allow on IPI.

On the rule-blind NEXUS-Stress benchmark, NEXUS reaches an F1 score of 0.881, highlighting the difficulty of fine-grained intervention routing.

With 0.205 ms median latency, NEXUS adds under 0.1% overhead to typical agent loops.

Code, benchmarks, and the calibrated risk scorer are publicly released.

전문 보기

이 뉴스, 어떠셨어요?

탭 한 번으로 반응 · 로그인 불필요

관련 뉴스

관련 뉴스 제보는 로그인 후 가능합니다.