Explainable Reinforcement Learning

Authors

  • Elena Hansen Author

DOI:

https://doi.org/10.5281/zenodo.19614772

Keywords:

explainable reinforcement learning; XRL; interpretable AI; counterfactual explanations; policy distillation; saliency maps; causal RL; human-in-the-loop

Abstract

Reinforcement learning (RL) has achieved remarkable results in game-playing, robotic control, and sequential decision-making -- yet the policies learned by deep RL agents remain largely opaque, limiting deployment in domains where decisions must be auditable, contestable, or legally justified. Explainable reinforcement learning (XRL) addresses this by generating human-interpretable accounts of why an agent took a given action or followed a given policy. This study systematically evaluates six XRL approaches -- saliency maps, attention visualisation, SHAP-based state attribution, policy distillation into decision trees, counterfactual explanation generation, and causal influence diagrams - across four RL benchmarks: Atari game-playing, MuJoCo continuous control, a medical treatment dosing simulation, and an autonomous driving safety scenario. Evaluation spans three dimensions: faithfulness to the underlying policy (fidelity metrics), human comprehensibility (user study, n=84 domain experts), and computational overhead. Counterfactual explanations achieve the highest human comprehensibility ratings (4.1/5) and the best fidelity on continuous control tasks (Spearman rho = 0.82). Policy distillation into decision trees sacrifices 6.4-12.8% task performance but achieves the highest auditor satisfaction ratings (4.3/5) for compliance-critical domains. Saliency maps are the most computationally efficient but score lowest on expert comprehensibility (2.6/5). A domain-appropriate XRL selection framework is proposed, distinguishing post-hoc local explanation, global policy explanation, and inherently interpretable policy design.

Downloads

Published

2026-08-19

How to Cite

Explainable Reinforcement Learning. (2026). Bio-QI  Journal, 3(2), 41-48. https://doi.org/10.5281/zenodo.19614772

Most read articles by the same author(s)