Research on Optimal Power Grid Scheduling Based on Transfer Reinforcement Learning
Main Article Content
Abstract
To enhance power grid adaptability amid rising renewable energy integration, this paper proposes M3-PPO, a meta-reinforcement learning algorithm that enables efficient the strategy transfer and rapid adaptation across tasks with varying energy mixes. Built upon a base framework (M-PPO) that integrates PPO and MAML, M3-PPO introduces two key innovations to overcome MAML’s training instability: a Mamba-based context encoder for richer task representation in the inner loop, and a global-local momentum update mechanism for smoother meta-parameter optimization in the outer loop. Experiments on the Grid2Op platform demonstrate that M3-PPO significantly outperforms baseline algorithms in generalization and scheduling efficiency, achieving robust performance even when simulating complex energy environments. The approach is particularly suitable for integration with antenna-enabled smart grid monitoring, wireless data acquisition, and edge-computing platforms, providing an engineering-oriented solution for adaptive, real-time, and robust power grid scheduling in modern renewable-rich energy systems.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
S. Sridhar and M. Govindarasu, “Model-Based Attack Detection and Mitigation for Automatic Generation Control,” IEEE Transactions on Smart Grid, vol. 5, no. 2, pp. 580-591, 2014, doi: 10.1109/TSG.2014.2298195.
Y. Li, R. Huang, and L. Ma, “False Data Injection Attack and Defense Method on Load Frequency Control,” IEEE Internet of Things Journal, vol. 8, no. 4, pp. 2910-2919, 2021, doi: 10.1109/JIOT.2020.3021429.
A. O. Erick and K. A. Folly, “Reinforcement Learning Approaches to Power Management in Grid-tied Microgrids: A Review,” Proceedings of the 2020 Clemson University Power Systems Conference (PSC); 10-13 March 2020; Clemson, SC, USA, doi: 10.1109/PSC50246.2020.9131138.
L. Lei, Y. Tan, G. Dahlenburg, W. Xiang, and K. Zheng, “Dynamic Energy Dispatch Based on Deep Reinforcement Learning in IoT-Driven Smart Isolated Microgrids,” IEEE Internet of Things Journal, vol. 8, no. 10, pp. 7938-7953, 2021, doi: 10.1109/JIOT.2020.3042007.
X. Zhou, J. Wang, X. Wang, and S. Chen, “Deep Reinforcement Learning for Microgrid Operation Optimization: A Review,” Proceedings of the 2023 8th Asia Conference on Power and Electrical Engineering (ACPEE); 14-16 April 2023; Tianjin, China, doi: 10.1109/ACPEE56931.2023.10135713.
E. Samadi, A. Badri, and R. Ebrahimpour, “Decentralized multi-agent based energy management of microgrid using reinforcement learning,” International Journal of Electrical Power & Energy Systems, vol. 122, Art. no. 106211, 2020, doi: 10.1016/j.ijepes.2020.106211.
J. Duan, “Deep-Reinforcement-Learning-Based Autonomous Voltage Control for Power Grid Operations,” IEEE Transactions on Power Systems, vol. 35, no. 1, pp. 814-817, 2020, doi: 10.1109/TPWRS.2019.2941134.
R. Huang, “Learning and Fast Adaptation for Grid Emergency Control via Deep Meta Reinforcement Learning,” IEEE Transactions on Power Systems, vol. 37, no. 6, pp. 4168-4178, 2022, doi: 10.1109/TPWRS.2022.3155117.
L. Xiong, “Meta-Reinforcement Learning-Based Transferable Scheduling Strategy for Energy Management,” IEEE Trans. Circuits Syst. I, vol. 70, no. 4, pp. 1685-1695, 2023, doi: 10.1109/TCSI.2023.3240702.
D. Jakobeit, M. Schenke, and O. Wallscheid, “Meta-Reinforcement-Learning-Based Current Control of Permanent Magnet Synchronous Motor Drives for a Wide Range of Power Classes,” IEEE Transactions on Power Electronics, vol. 38, no. 7, pp. 8062-8074, 2023, doi: 10.1109/TPEL.2023.3256424.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” 2014. arXiv preprint. arXiv:1412.6980, doi: 10.48550/arXiv.1412.6980.
A. Gu and T. Dao, “Mamba: Linear-Time Sequence Modeling with Selective State Spaces,” 2024. arXiv: arXiv:2312.00752, doi: 10.48550/arXiv.2312.00752.
C. Chen, M. Cui, F. Li, S. Yin, and X. Wang, “Model-Free Emergency Frequency Control Based on Reinforcement Learning,” IEEE Transactions on Industrial Informatics, vol. 17, no. 4, pp. 2336-2346, 2020, doi: 10.1109/TII.2020.3001095.
Z. Yan and Y. Xu, “A Multi-Agent Deep Reinforcement Learning Method for Cooperative Load Frequency Control of a Multi-Area Power System,” IEEE Transactions on Power Systems, vol. 35, no. 6, pp. 4599-4608, 2020, doi: 10.1109/TPWRS.2020.2999890.
X. Hu, Y. Zhang, H. Xia, W. Wei, Q. Dai, and J. Li, “Toward Fair Power Grid Control: A Hierarchical Multiobjective Reinforcement Learning Approach,” IEEE Internet of Things Journal, vol. 11, no. 4, pp. 6582-6595, 2024, doi: 10.1109/JIOT.2023.3314522.
A. Marot, I. Guyon, B. Donnot, G. Dulac-Arnold, P. Panciatici, and M. Awad, “L2RPN: Learning to Run a Power Network in a Sustainable World NeurIPS2020 challenge design,” Réseau de Transport d’Électricité. Paris, France: White Paper; 2020.
J. R. Vázquez-Canteli, J. Kämpf, G. Henze, and Z. Nagy, “CityLearn v1.0: An OpenAI Gym Environment for Demand Response with Deep Reinforcement Learning,” Proceedings of the 6th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation, BuildSys ’19; 13-14 November 2019; New York, NY, USA. New York, NY, USA: Association for Computing Machinery; 2019, doi: 10.1145/3360322.3360998.
K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman, “Quantifying Generalization in Reinforcement Learning,” Proceedings of the 36th International Conference on Machine Learning, PMLR, 202, [Online]. Available: https://proceedings.mlr.press/v97/cobbe19a.html.
T. Wu and J. Wang, “Artificial intelligence for operation and control: The case of microgrids,” The Electricity Journal, vol. 34, no. 1, Art. no. 106890, 2021, doi: 10.1016/j.tej.2020.106890.
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel, “RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning,” 2016. arXiv: arXiv:1611.02779, doi: 10.48550/arXiv.1611.02779.
C. Finn, P. Abbeel, and S. Levine, “Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks,” Proceedings of the 34th International Conference on Machine Learning, PMLR, 2017, [Online]. Available: https://proceedings.mlr.press/v70/finn17a.html.
Á. Belmonte-Baeza, J. Lee, G. Valsecchi, and M. Hutter, “Meta Reinforcement Learning for Optimal Design of Legged Robots,” IEEE Robotics and Automation Letters, vol. 7, no. 4, pp. 12134-12141, 2022, doi: 10.1109/LRA.2022.3211785.
F. Ye, P. Wang, C.-Y. Chan, and J. Zhang, “Meta Reinforcement Learning-Based Lane Change Strategy for Autonomous Vehicles,” Proceedings of the 2021 IEEE Intelligent Vehicles Symposium (IV); Nagoya, Japan; 1 November 2021. New York, NY, USA: IEEE; 2021, doi: 10.1109/IV48863.2021.9575379.
Q. Liu, “A few-shot disease diagnosis decision making model based on meta-learning for general practice,” Artificial Intelligence in Medicine, vol. 147, Art. no. 102718, 2024, doi: 10.1016/j.artmed.2023.102718.
H. Zheng, Z. Yang, W. Liu, J. Liang, and Y. Li, “Improving deep neural networks using softplus units,” Proceedings of the 2015 International Joint Conference on Neural Networks (IJCNN); Killarney; 12-17 July 2015; New York, NY, USA: IEEE; 2015, doi: 10.1109/IJCNN.2015.7280459.
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal Policy Optimization Algorithms,” 2017. arXiv: arXiv:1707.06347, doi: 10.48550/arXiv.1707.06347.
H. Fu, “Towards Effective Context for Meta-Reinforcement Learning: An Approach based on Contrastive Learning,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 8, pp. 7457-7465, 2021, doi: 10.1609/aaai.v35i8.16914.
L. Kirsch, J. Harrison, J. Sohl-Dickstein, and L. Metz, “General-Purpose In-Context Learning by Meta-Learning Transformers,” 2024. arXiv: arXiv:2212.04458, doi: 10.48550/arXiv.2212.04458.