Reinforcement Learning for Industrial Automation

Authors

  • Erik Dubois Author

DOI:

https://doi.org/10.5281/zenodo.19614649

Keywords:

reinforcement learning; industrial automation; process control; safe RL; offline RL; job-shop scheduling; energy optimisation; sim-to-real transfer

Abstract

Industrial automation -- encompassing manufacturing process control, supply chain scheduling, energy management, and quality assurance -- presents a compelling yet underexplored domain for reinforcement learning, where sequential decision-making under uncertainty could yield substantial efficiency gains over hand-tuned heuristic controllers. This study presents a controlled evaluation of five RL approaches -- model-free PPO, model-based Dreamer, safe constrained RL (CPO), offline RL from historical logs (CQL), and hybrid RL-PID control -- across four industrial automation tasks: chemical reactor temperature control (OpenAI Safety Gym + custom CSTR simulator), job-shop scheduling (OR-Library benchmarks with stochastic disruptions), HVAC energy optimisation (EnergyPlus building simulation), and visual quality inspection (defect detection with active camera positioning). A total of 1,920 experiments were conducted in high-fidelity simulators validated against real industrial data. The hybrid RL-PID controller achieved the best balance of performance and safety on chemical process control: 12.4 +- 1.8% energy reduction versus the baseline PID with zero safety constraint violations across 10,000 episodes, compared to 18.2% energy reduction but 2.4% violation rate for unconstrained PPO. Offline RL from historical logs achieved 84.6% of online RL performance without any simulator interaction, enabling deployment where building an accurate simulator is infeasible. Model-based Dreamer required 8x fewer environment interactions than PPO to reach equivalent performance on scheduling, making it suitable for expensive-to-simulate industrial processes. Safe constrained RL (CPO) reduced safety violations by 94.2% relative to unconstrained PPO while sacrificing only 4.8% of reward. The sim-to-real gap on a physical CSTR pilot plant was 8.6% for the hybrid controller versus 22.4% for unconstrained PPO, confirming that safety constraints improve real-world transfer. A practical industrial RL deployment framework mapping process characteristics, safety requirements, and data availability to recommended approaches is proposed.

Downloads

Published

2026-08-19

How to Cite

Reinforcement Learning for Industrial Automation. (2026). Bio-QI  Journal, 2(1), 28-36. https://doi.org/10.5281/zenodo.19614649