Explainability Framework for Policy-Aware Autonomous Agents
Abstract
In the field of Artificial Intelligence, an agent is a system which is able to autonomously make decisions in order to reach a desired goal.
As these systems grow more prevalent in our day-to-day lives, there has been an increased need to add explainability features which can provide an account for an agent's behavior.
We therefore propose a framework that outlines how to produce comprehensible explanations for policy-aware agents, or agents which have rule-enforcing policies incorporated in their decision-making framework.
This framework is designed using insights from the social sciences on how to produce good explanations.
It is implemented in the Answer Set Programming language while using Python to assist with information extraction and natural-language translation.
Because these agents incur penalties when violating policies, we are able to leverage these penalties to detect undesirable events in scenarios that are counterfactual to the agents' original actions.
This lends itself to creating contrastive explanations (e.g., "the agent performed this action because, had it not, undesirable event X would have occurred."), which formulate the core component for our explainability framework.
The framework is evaluated using a survey wherein human participants provide feedback on our program-generated explanations.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요