Explainable Deep Reinforcement Learning for Microgrid Resilient Energy Management
Nejati Amiri, Mohammad Hossein (2026) Explainable Deep Reinforcement Learning for Microgrid Resilient Energy Management. Doctoral thesis, Birmingham City University.
Preview |
Text
Mohammad Hossein Nejati Amiri PhD Thesis_Final Version_Final Award July 2026.pdf - Accepted Version Download (40MB) |
Abstract
Microgrids enable the integration of distributed renewable generation and can support decarbonisation, particularly in rural or weakly connected areas where main grid access is limited or unreliable. This thesis focuses on a standalone rural microgrid in a cyclone-prone coastal region, where weather-driven High-Impact, Low-Probability (HILP) events are the main resilience concern. In this context, HILP events can substantially reduce renewable generation, interrupt normal operation, and increase the risk of unserved priority loads. The studied system relies on photovoltaic generation, wind generation, and a Battery Energy Storage System (BESS), with the battery treated as the only energy storage and fault-ride-through resource during disruptive periods. The central challenge is therefore to keep critical loads energised during weather-related HILP events while limiting battery ageing and providing explanations that operators can trust. This thesis examines whether an interpretable, data-driven controller can match a tuned Model Predictive Control (MPC) benchmark with lower computational cost and improved battery preservation within this defined standalone microgrid context.
A cyclone-prone coastal site in India is adopted as the case study, with asset sizing de-rived in HOMER Pro from historical demand and weather records that include the Cyclone Laila period to ensure realistic constraints. The benchmark is a multi-objective MPC formulation implemented as a mixed-integer linear programme in Pyomo. It balances resilience, power imbalance, switching effort, and life-cycle degradation through adaptive State Of Charge (SOC) limits and optional grid trading. The proposed alternative is a Deep Reinforcement Learning (DRL) policy based on Proximal Policy Optimisation (PPO), trained with a four-stage curriculum that progressively increases renewable perturbations and sce-nario difficulty, and interpreted post hoc with SHapley Additive exPlanations (SHAP) and Local Interpretable Model-Agnostic Explanations (LIME). Robustness is assessed through Monte Carlo experiments under smoothly correlated Perlin noise uncertainty.
Under deterministic conditions, the learned policy achieves a Resilience Index (RI) of 0.996 versus 0.999 for MPC and extends battery life Expected Year (EY) from 14.44 to 15.90 years by smoothing charge-discharge behaviour. The rate of change and the acceleration of SOC fall by 12–19% and 26–38%, respectively. Under uncertainty, resilience remains within ∼ 0.3 percentage of MPC (RI 0.996 versus 0.999) and EY improves by ∼ 5% (15.88 versus 15.06 years). Across 4,000 runs, performance concentrates tightly with mean RI 0.992 and EY 15.95 years. Inference is 5,000× lighter per step than MPC, and lifetime compute is almost 3× lower, indicating suitability for edge deployment. Explanations agree with power system intuition: charge under surplus, discharge under deficit, and use medium-priority loads as the primary source of flexibility.
Actions (login required)
![]() |
View Item |

Tools
Tools