INTELLIGENT SCHEDULING OF PV–STORAGE–CHARGING INTEGRATED STATIONS VIA GTRXL-PPO WITH CURRICULUM LEARNING

Authors

  • SongCheng Lu (Corresponding Author) School of Electrical Engineering and Automation, Hefei University of Technology, Xuancheng 242000, Hefei, China.

Keywords:

Long range temporal dependency decision making, PV–storage–charging integration, Curriculum learning, Gated Transformer XL, Proximal Policy Optimization

Abstract

To address the long-horizon sequential decision-making task, characterized by complex temporal dependencies, non-stationary dynamics, and high stochasticity in distribution-level PV–storage–charging systems, this paper develops a deep reinforcement learning framework that combines Gated Transformer‑XL (GTrXL) with Proximal Policy Optimization (PPO). Cross-segment memory captures long‑range temporal dependencies and time‑aware encodings reinforce intraday periodicity. Training adopts a five‑stage curriculum with adaptive KL control and auxiliary multi‑task heads to improve sample efficiency. A group‑normalized, potential‑based reward unifies economic performance, grid friendliness, and storage health. In simulation, the agent learns a structured six-phase daily policy and achieves a 94.4% charging completion rate and 92.8% PV utilization, reduces average daily electricity purchase cost by 15%, and keeps grid peak power within a 60 kW soft limit. Across five seeds, returns improve by 78.4% over a feed-forward PPO baseline and by 23.6% over a vanilla GTrXL-PPO, demonstrating the benefits of long-memory RL for coordinated PV–storage–charging operation. The framework enforces feasibility via continuous action mapping with ramp-rate/jerk constraints and supports millisecond-level inference. Uncertainty-aware shaping improves robustness; gains are statistically significant across five seeds via paired tests.

References

[1] Rehman W U, Bo R, Mehdipourpich H, et al. Sizing battery energy storage and PV system in an extreme fast charging station considering uncertainties and battery degradation. Applied Energy, 2022, 313: 118745.

[2] Zhao Z, Lee C K M, Yan X, et al. Reinforcement learning for electric vehicle charging scheduling: a systematic review. Transportation Research Part E: Logistics and Transportation Review, 2024, 190: 103698.

[3] Liao X, Cao N, Li M, et al. Research on short-term load forecasting using XGBoost based on similar days. 2019 International Conference on Intelligent Transportation, Big Data & Smart City (ICITBS). Changsha, China, 2019: 675-678.

[4] Zhu K, Jiang Y, Wang K, et al. Day-ahead campus load interval forecast based on similar day and kernel function estimation. 2020 International Conference on Intelligent Computing, Automation and Systems (ICICAS). 2020: 145-148.

[5] Sutton R S, Barto A G. Reinforcement Learning: An Introduction. 2nd ed. Cambridge, MA: MIT Press, 2018.

[6] Mnih V, Kavukcuoglu K, Silver D, et al. Human-level control through deep reinforcement learning. Nature, 2015, 518(7540): 529-533.

[7] Schulman J, Wolski F, Dhariwal P, et al. Proximal policy optimization algorithms. arXiv, 2017. DOI: 10.48550/arXiv.1707.06347.

[8] Alonso M, Amaris H, Martin D, et al. Proximal policy optimization for energy management of electric vehicles and PV storage units. Energies, 2023, 16(15): 5689.

[9] Upadhyay S, Ahmed I, Mihet-Popa L. Energy management system for an industrial microgrid using optimization-algorithms-based reinforcement learning technique. Energies, 2024, 17(16): 3898.

[10] Heendeniya C B, Nespoli L. A stochastic deep reinforcement learning agent for grid-friendly electric vehicle charging management. Energy Informatics, 2022, 5(S1): 28.

[11] Li H, Dai X, Goldrick S, et al. Reinforcement learning for EV fleet smart charging with on-site renewable energy sources. Energies, 2024, 17(21): 5442.

[12] Hermans B A L M, Walker S, Ludlage J H A, et al. Model predictive control of vehicle charging stations in grid-connected microgrids: an implementation study. Applied Energy, 2024, 368: 123210.

[13] Pascanu R, Mikolov T, Bengio Y. On the difficulty of training recurrent neural networks. Proceedings of the 30th International Conference on Machine Learning (ICML 2013). PMLR 28(3), 2013: 1310-1318.

[14] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS 2017). 2017: 5998-6008.

[15] Dai Z, Yang Z, Yang Y, et al. Transformer-XL: attentive language models beyond a fixed-length context. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL 2019). 2019: 2978-2988.

[16] Parisotto E, Song H F, Rae J W, et al. Stabilizing transformers for reinforcement learning. Proceedings of the 37th International Conference on Machine Learning (ICML 2020). PMLR 119, 2020: 7487-7498.

[17] Li H, Wan Z, He H. Constrained EV charging scheduling based on safe deep reinforcement learning. IEEE Transactions on Smart Grid, 2020, 11(3): 2427–2439.

[18] Li W, Luo H, Lin Z, et al. A survey on transformers in reinforcement learning. arXiv, 2023. DOI: 10.48550/arXiv.2301.03044.

Downloads

Published

2026-07-31

How to Cite

SongCheng Lu. Intelligent Scheduling Of Pv–Storage–Charging Integrated Stations Via Gtrxl-Ppo With Curriculum Learning. Journal of Computer Science and Electrical Engineering. 2026, 8(6): 9-21. DOI: https://doi.org/10.61784/jcsee3148.