Explainable Reinforcement Learning
DOI:
https://doi.org/10.5281/zenodo.19614772Keywords:
explainable reinforcement learning; XRL; interpretable AI; counterfactual explanations; policy distillation; saliency maps; causal RL; human-in-the-loopAbstract
Reinforcement learning (RL) has achieved remarkable results in game-playing, robotic control, and sequential decision-making -- yet the policies learned by deep RL agents remain largely opaque, limiting deployment in domains where decisions must be auditable, contestable, or legally justified. Explainable reinforcement learning (XRL) addresses this by generating human-interpretable accounts of why an agent took a given action or followed a given policy. This study systematically evaluates six XRL approaches -- saliency maps, attention visualisation, SHAP-based state attribution, policy distillation into decision trees, counterfactual explanation generation, and causal influence diagrams - across four RL benchmarks: Atari game-playing, MuJoCo continuous control, a medical treatment dosing simulation, and an autonomous driving safety scenario. Evaluation spans three dimensions: faithfulness to the underlying policy (fidelity metrics), human comprehensibility (user study, n=84 domain experts), and computational overhead. Counterfactual explanations achieve the highest human comprehensibility ratings (4.1/5) and the best fidelity on continuous control tasks (Spearman rho = 0.82). Policy distillation into decision trees sacrifices 6.4-12.8% task performance but achieves the highest auditor satisfaction ratings (4.3/5) for compliance-critical domains. Saliency maps are the most computationally efficient but score lowest on expert comprehensibility (2.6/5). A domain-appropriate XRL selection framework is proposed, distinguishing post-hoc local explanation, global policy explanation, and inherently interpretable policy design.Downloads
Published
2026-08-19
Issue
Section
Articles
How to Cite
Explainable Reinforcement Learning. (2026). Bio-QI Journal, 3(2), 41-48. https://doi.org/10.5281/zenodo.19614772

