Constructing the Optimal Bidding Path for VPPS Participating in the Spot Market Using the DDPG Reinforcement Learning Algorithm
Main Article Content
Abstract
The increasing penetration of distributed renewable energy sources has intensified the need for intelligent bidding strategies in virtual power plants (VPPs), where reliable communication and real-time information exchange are essential for coordinated energy management. This study proposes an optimal bidding path construction framework based on the Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm for VPP participation in electricity spot markets. A Markov decision process is established to characterize dynamic market interactions, and customized state-space optimization, constrained action-space design, and a multi-objective reward function are integrated into the Actor–Critic architecture to jointly maximize economic returns while satisfying operational constraints. The framework further incorporates communication-aware resource coordination mechanisms that leverage edge computing and low-latency information exchange to enhance decision consistency under uncertain renewable generation and volatile market conditions. Experimental evaluation demonstrates that the improved DDPG algorithm increases average daily revenue by 39.1% compared with conventional DDPG, accelerates convergence by approximately 15%, reduces revenue volatility by 12%, and maintains the constraint violation rate at 1.2%. In addition to intelligent energy scheduling, the proposed methodology provides valuable insights into communicationenabled power systems, distributed electromagnetic information networks, and wireless coordination infrastructures requiring adaptive decision-making and reliable multi-node information
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
References
G. Wang, J. Deng, and D. Duan, “Data-driven H2/H∞control for full-car active suspension systems via stochastic reinforcement learning algorithm,” Proceedings of the Institution of Mechanical Engineers, Part C: Journal of Mechanical Engineering Science, vol. 240, no. 5, pp. 1303-1321, 2026, doi: 10.1177/09544062251406280.
H. Guo, Z. Chai, and Y. Li, “QER-LPD3QN: A quantum-Inspired Sequence-Aware deep reinforcement learning algorithm for path planning,” Expert Systems with Applications, vol. 313, Art. no. 131575, 2026, doi: 10.1016/J.ESWA.2026.131575.
M. Bongiovi, “Deep Reinforcement Learning algorithms learn important classes of repeated games optimally—Theoretical and empirical analysis,” Franklin Open, vol. 14, pp. 100503-100503, 2026, doi: 10.1016/J.FRAOPE.2026.100503.
Z. Wang, J. Song, Y. Liu, and J. Zhao, “Reinforcement learning algorithm for reusable resource allocation with unknown rental time distribution,” European Journal of Operational Research, vol. 331, no. 1, pp. 186-199, 2026, doi: 10.1016/J.EJOR.2025.09.012.
R. Wang, Y. Shen, D. Wang, and W. Li, “A cognitive internet of things resource allocation method based on multiagent reinforcement learning algorithm,” Scientific reports, vol. 16, pp. 7756, 2026, doi: 10.1038/S41598-026-36380-X.
H. Yu, C. H. Yang, and Y. Huang, “A multi-objective goal-oriented reinforcement learning algorithm for dynamic multi-objective sequential decision making,” Autonomous Agents and Multi-Agent Systems, vol. 40, no. 1, pp. 5, 2026, doi: 10.1007/S10458-026-09735-X.
D. Belomestny, I. Levin, A. Naumov, and S. Samsonov, “UVIP: Model-Free Approach to Evaluate Reinforcement Learning Algorithms,” Journal of Optimization Theory and Applications, vol. 208, no. 3, pp. 89, 2026, doi: 10.1007/S10957-025-02903-1.
S. Cammarota, M. Ferrante, A. Carosi, R. M. D’Angelillo, and N. Toschi, “Beam angle optimization for radiotherapy using LLMs via reinforcement-learning inspired iterative refinement,” Medical physics, vol. 53, no. 2, Art. no. e70258, 2026, doi: 10.1002/MP.70258.
Y. Yang, T. Wang, Y. Fu, J. Huang, and D. Zhou, “Portfolio management based on value distribution reinforcement learning algorithm,” Frontiers in Artificial Intelligence, vol. 8, pp. 1709493-1709493, 2026, doi: 10.3389/FRAI.2025.1709493.
H. Huang, M. Li, Y. Sun, J. Zhang, and X. Lin, “Multi-agent Co-optimized battery aging-aware control strategy for a CVT-hybrid electric vehicle by using PPO and DQN reinforcement learning algorithm,” Energy, vol. 344, pp. 139971-139971, 2026, doi: 10.1016/J.ENERGY.2026.139971.
S. Khan, A. A. Khan, R. Mahendran K, M. Fazil, A. U. Rehman, W. Jiang, et al., “C2DEEP-OT: Utilizing Multi-Agent Deep Reinforcement Learning Algorithm and Optimized Attentive Transformer Network for Cervical Cancer Detection,” Information Sciences, vol. 738, pp. 123047-123047, 2026, doi: 10.1016/J.INS.2025.123047.
J. L. Mpoporo, A. P. Owolawi, and C. Tu, “Deep Reinforcement Learning Algorithms for Intrusion Detection: A Bibliometric Analysis and Systematic Review,” Applied Sciences, vol. 16, no. 2, pp. 1048, 2026, doi: 10.3390/APP16021048.
S. Zong, J. Chen, Y. Hu, and J. Li, “An Iterative Reinforcement Learning Algorithm for Speed Drop Compensation in Rolling Mills,” Algorithms, vol. 19, no. 1, pp. 84-84, 2026, doi: 10.3390/A19010084.
C. Pan, Z. Zhang, S. Wen, M. Zhu, Z. Han, and Y. Chen, “Efficient turbine placement optimization in large-scale offshore wind farms: A space-constrained deep reinforcement learning algorithm,” Energy, vol. 344, Art. no. 139929, 2026, doi: 10.1016/J.ENERGY.2026.139929.
Q. Xu, Z. Zhang, J. Li, and X. Qi, “Hierarchical multi-agent reinforcement learning algorithm for multi-UAV roundup strategy,” Applied Mathematical Modelling, vol. 155, pp. 116728-116728, 2026, doi: 10.1016/J.APM.2025.116728.
U. Khekare and R. Vedaraj IS, “Optimized multi agent reinforcement learning algorithms with hybrid BiLSTM for cost efficient EV charging scheduling,” Frontiers in Artificial Intelligence, vol. 8, Art. no. 1700664, 2026, doi: 10.3389/FRAI.2025.1700664.
A. A. Amer, S. Bayhan, H. Rub A, A. Massoud, and M. Ehsani, “A review of deep reinforcement learning algorithms for grid services in grid-interactive efficient buildings,” Energy Reports, vol. 15, Art. no. 108900, 2026, doi: 10.1016/J.EGYR.2025.12.037.
M. Subramaniyan, X. Jin, S. Nagaraja, A. Wallqvist, and J. Reifman, “A reinforcement learning algorithm to optimize resource utilization in combat casualty care,” Scientific Reports, vol. 15, no. 1, Art. no. 44534, 2025, doi: 10.1038/S41598-025-28021-6.
A. Priya, R. Tiwari, P. Agrawal, and S. Kumar, “Reward shaping of deep reinforcement learning algorithm for autonomous navigation in a structured environment,” Intelligent Service Robotics, vol. 19, no. 1, pp. 8, 2025, doi: 10.1007/S11370-025-00673-3.
O. Mortabit, M. Ahachad, and I. S. Kaitouni, “Implementing deep reinforcement learning algorithms for optimal building VRF performance considering static and adaptive thermal comfort models,” Applied Thermal Engineering, vol. 287, Art. no. 129354, 2026, doi: 10.1016/J.APPLTHERMALENG.2025.129354.
M. Almatared, M. Abuhussain, Z. Andleeb, and F. M. Bashir, “Adaptive reinforcement learning algorithm for realtime energy optimization in building digital twins with heterogeneous IoT sensor networks,” Automation in Construction, vol. 182, Art. no. 106714, 2026, doi: 10.1016/J.AUTCON.2025.106714.
K. Peng, K. Yue, P. Xiao, and V. C. Leung, “Security-aware computation offloading in internet of vehicles: a multiagent reinforcement learning algorithm with attention mechanism,” Journal of Cloud Computing, vol. 15, no. 1, pp. 10, 2025, doi: 10.1186/S13677-025-00821-1.
Y. Tong, B. Xie, Z. Zhao, Z. Chen, Z. Lu, Z. Niu, et al., “Improved SAC reinforcement learning algorithm: Active control of drive wheel torque to improve energy utilization and reduce power consumption in electric tractors,” Computers and Electronics in Agriculture, vol. 241, Art. no. 111258, 2026, doi: 10.1016/J.COMPAG.2025.111258.
S. Moazzami, A. Mirzaei, M. Aminian, R. Karimi, and N. Mikaeilvand, “A hybrid fuzzy logic and deep reinforcement learning algorithm for adaptive task scheduling and resource allocation in heterogeneous Fog– Cloud environments,” Sustainable Computing: Informatics and Systems, vol. 49, Art. no. 101260, 2026, doi: 10.1016/J.SUSCOM.2025.101260.
S. A. M. Bobi, I. Rodriguez, B. J. F. López, A. Muñoz, J. Anguera, D. Gonzalez-Calvo, et al., “TD3 Reinforcement Learning Algorithm Used for Health Condition Monitoring of a Cooling Water Pump,” Computers, vol. 14, no. 12, pp. 540, 2025, doi: 10.3390/COMPUTERS14120540.
J. H. Asl, V. A. Le, V. B. Minh, and M. R. Elara, “Model-free inverse reinforcement learning algorithms for continuoustime and discrete-time zero-sum games,” Neurocomputing, vol. 665, Art. no. 132101, 2026, doi: 10.1016/J.NEUCOM.2025.132101.
H. Zhang, J. Fu, Y. Zhang, and H. Du, “Multi-Agent Reinforcement Learning Algorithm Based on Local Observation Imitation Learning,” IET Control Theory & Applications, vol. 19, no. 1, Art. no. e70097, 2025, doi: 10.1049/CTH2.70097.
J. Alponse, C. Yaashuwanth, and K. Prathibanandhi, “A Novel Scalable Trust-Aware Deep Reinforcement Learning Algorithm for Energy-Efficient and Secure Routing in Software-Defined Wireless Sensor Networks for IoT,” Measurement Science Review, vol. 25, no. 6, pp. 358-365, 2025, doi: 10.2478/MSR-2025-0039.
J. Tang, “Deep-reinforcement-learning–guided resource allocation and task offloading for 6G edge intelligence,” Computer Communications, vol. 245, Art. no. 108364, 2026, doi: 10.1016/J.COMCOM.2025.108364.
C. Ma, F. Gao, L. Ji, C. Zhang, L. Long, and J. Zhang, “Toward personalized risk-sensitive decision-making: A novel risk preference adaptive distributional reinforcement learning algorithm for stock trading,” Applied Soft Computing, vol. 186, no. PD, Art. no. 114269, 2026, doi: 10.1016/J.ASOC.2025.114269.