PDF(2375 KB)
Multi-Agent Reinforcement Learning-Driven P2P Energy Trading Strategy for Community Prosumers
CHEN Guili, CHEN Danhong, WEN Hongwu, ZHENG Yongwei, MAI Liang, FENG Xiashan, ZHONG Qishen, HU Yiming, ZHENG Jiehui
Electric Power Construction ›› 2026, Vol. 47 ›› Issue (8) : 14-25.
PDF(2375 KB)
PDF(2375 KB)
Multi-Agent Reinforcement Learning-Driven P2P Energy Trading Strategy for Community Prosumers
[Objective] To address the variations in resource allocation and energy consumption behavior among prosumers in peer-to-peer (P2P) community energy trading scenarios, as well as the limited adaptability of traditional model-based methods in uncertain environments, this paper proposes a multi-agent reinforcement learning method that features both scalability and privacy protection capabilities. [Methods] First, three representative types of heterogeneous prosumer models are constructed. Second, a community energy trading model based on the mid-market rate pricing mechanism is established, and a flexibility incentive mechanism is introduced. Finally, the energy trading decision-making problem of prosumers is formulated as a partially observable Markov decision process, and a soft actor-critic algorithm based on dynamic mean-field (DMF-SAC) approximation is proposed to solve the energy management strategies of prosumers. [Results] Simulation results demonstrate that the proposed method outperforms baseline methods in terms of convergence performance, computational overhead, and operating costs. It also effectively improves the local consumption of distributed energy and enhances peak-shaving and valley-filling capabilities. [Conclusions] The proposed method effectively improves the efficiency and economic benefits of collaborative optimization for heterogeneous prosumers while balancing privacy protection and system scalability, which holds significant value for energy trading and management in community-based markets.
peer-to-peer (P2P) energy trading / energy management / multi-agent reinforcement learning / privacy preservation / dynamic mean-field soft actor-critic(DMF-SAC) algorithm
| [1] |
李晖, 刘栋, 姚丹阳. 面向碳达峰碳中和目标的我国电力系统发展研判[J]. 中国电机工程学报, 2021, 41(18): 6245-6259.
|
| [2] |
陈郑平, 李文忠, 陈飞雄, 等. 分布式资源助力新型电力系统灵活性提升研究综述[J]. 电力工程技术, 2025, 44(2): 145-159.
|
| [3] |
|
| [4] |
鲁卓欣, 徐潇源, 严正, 等. 不确定性环境下数据驱动的电力系统优化调度方法综述[J]. 电力系统自动化, 2020, 44(21): 172-183.
|
| [5] |
宋铎洋, 薛田良, 李艺瀑, 等. 考虑风光不确定性的虚拟电厂合作博弈调度及收益分配策略[J]. 电力工程技术, 2025, 44(1): 193-206.
|
| [6] |
陈景文, 单茜, 刘耀先, 等. 面向电力市场的用户侧电力电量预测综述[J]. 电网与清洁能源, 2024, 40(2): 10-20.
|
| [7] |
周军, 李佳旺, 马鸿君, 等. 考虑点对点电能共享的智能楼宇群分布式优化调度[J]. 电力自动化设备, 2021, 41(10): 113-121.
|
| [8] |
孙宇军, 朱子旭, 赵兴勇, 等. 基于多智能体深度强化学习的微电网群能量管理与交易联合优化策略[J]. 电力建设, 2022, 43(11): 112-120.
|
| [9] |
范宏, 盛哲祺, 张树卿. 分布式灵活性资源的动态聚合调控与协同效益分配研究综述[J]. 浙江电力, 2025, 44(9): 30-45.
|
| [10] |
|
| [11] |
李扬, 马文捷, 卜凡金, 等. 多智能体深度强化学习驱动的跨园区能源交互优化调度[J]. 电力建设, 2024, 45(5): 59-70.
为协调多园区综合能源系统各个园区之间的能量交互,多能源子系统之间的能源转换,实现综合能源系统整体优化调度,提出一种利用多智能体深度强化学习算法学习不同园区的负荷特征,并在此基础上进行决策的综合调度模型。该模型将多园区综合能源系统的调度问题转化为马尔科夫决策过程,并利用深度强化学习算法进行求解,避免了对多园区、多能源子系统之间复杂的能量耦合关系进行建模。仿真结果表明,所提方法可以很好地捕捉到不同园区的负荷特性,并利用其中的互补特性协调不同园区之间进行合理的能量交互,可以实现弃风率由16.3%降低至0,并可以使总运行成本降低5 445.6元,具有良好的经济效益和环保效益。
In order to coordinate energy interactions among various communities and energy conversions among multi-energy subsystems within the multi-community integrated energy system under uncertain conditions, and achieve overall optimization and scheduling of the comprehensive energy system, this paper proposes a comprehensive scheduling model that utilizes a multi-agent deep reinforcement learning algorithm to learn load characteristics of different communities and make decisions based on this knowledge. In this model, the scheduling problem of the integrated energy system is transformed into a Markov decision process and solved using a data-driven deep reinforcement learning algorithm, which avoids the need for modeling complex energy coupling relationships between multi-communities and multi-energy subsystems. The simulation results show that the proposed method effectively captures the load characteristics of different communities and utilizes their complementary features to coordinate reasonable energy interactions among them. This leads to a reduction in wind curtailment rate from 16.3% to 0% and lowers the overall operating cost by 5445.6 Yuan, demonstrating significant economic and environmental benefits. |
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
李美成, 梅文明, 张凌康, 等. 基于可再生能源不确定性的多能源微网调度优化模型研究[J]. 电网技术, 2019, 43(4): 1260-1270.
|
| [22] |
赵鹏杰, 吴俊勇, 林凯骏, 等. 基于一致性算法的多微电网点对点分布式能量交易策略[J]. 电网技术, 2023, 47(1): 205-218.
|
| [23] |
|
| [24] |
|
| [25] |
龚迪阳, 唐雅洁, 高为举, 等. 基于Q学习和PCA的分布式光伏集群优化调度[J]. 浙江电力, 2025, 44(11): 72-82.
|
| [26] |
李忠凡, 陈曦, 黄海涛. 基于多智能体分层强化学习的多园区综合能源系统优化运行[J]. 浙江电力, 2025, 44(9): 46-57.
|
| [27] |
李钟平, 向月. 深度强化学习驱动的风储系统参与能量-调频市场竞价策略[J]. 电力工程技术, 2025, 44(3): 30-42.
|
| [28] |
耿天旭, 梁俊宇, 龚新勇, 等. 基于深度强化学习的配电网多主体协同电压控制方法[J]. 电网与清洁能源, 2024, 40(9): 74-80, 91.
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
梁泽庭, 郑杰辉, 方家琨, 等. 基于多智能体强化学习的差异化产销者参与社区能源交易方法[J]. 电网技术, 2025, 49(5): 1826-1836.
|
| [34] |
|
| [35] |
|
| [36] |
U.S. Department of Transportation. 2017 National household travel survey[EB/OL]. Washington, DC: U.S. Department of Transportation, 2017 [2026-03-31]. https://nhts.ornl.gov.
|
| [37] |
|
| [38] |
The coordination of distributed energy resources (DERs) within virtual power plants (VPPs) is expected to generate significant economic benefits and enhance the operational stability of modern power systems. However, achieving massive coordination of heterogeneous and uncertain DERs remains a challenge in current research. To address this issue, this article proposes a novel bi-level optimization approach based on mean-field reinforcement learning (MFRL) to enable the coordination of massive DERs in VPPs. The problem is decomposed into multiple subproblems: the upper-level subproblem models power dispatch among integrated energy systems (IESs) in response to coordinated demand, while a series of lower-level subproblems determine the operational schemes of DERs within individual IESs. Considering the large decision space, an MFRL algorithm with fast Shapley credit allocation is developed to efficiently solve the upper-level optimization. Meanwhile, the lower-level subproblems are formulated as small-scale mixed-integer linear programming (MILP) problems, addressing the difficulties caused by IES heterogeneity in applying mean-field approximation. Simulation results show that the proposed approach significantly improves convergence speed and reduces the global cost of VPP operation, especially in massive-scale scenarios. In test scenarios ranging from 10 to 500 agents, the proposed bi-level optimization approach improves the objective by 4.8%-26.6%, compared to the advanced baseline method.
|
| [39] |
This paper develops a multi-timescale coordinated operation method for microgrids based on modern deep reinforcement learning. Considering the complementary characteristics of different storage devices, the proposed approach achieves multi-timescale coordination of battery and supercapacitor by introducing a hierarchical two-stage dispatch model. The first stage makes an initial decision irrespective of the uncertainties using the hourly predicted data to minimize the operational cost. For the second stage, it aims to generate corrective actions for the first-stage decisions to compensate for real-time renewable generation fluctuations. The first stage is formulated as a non-convex deterministic optimization problem, while the second stage is modeled as a Markov decision process solved by an entropy-regularized deep reinforcement learning method, i.e., the Soft Actor-Critic. The Soft Actor-Critic method can efficiently address the exploration–exploitation dilemma and suppress variations. This improves the robustness of decisions. Simulation results demonstrate that different types of energy storage devices can be used at two stages to achieve the multi-timescale coordinated operation. This proves the effectiveness of the proposed method.\n
|
| [40] |
|
| [41] |
Open Power System Data. Data package time series[EB/OL]. [2026-03-31]. https://data.open-power-system-data.org/time_series/2020-10-06.
|
利益冲突声明(Conflict of Interests) 所有作者声明不存在利益冲突。
作者贡献声明(Authors' Contributions) 陈桂力负责研究设计、模型构建与论文撰写;陈丹红负责案例选择与数据收集;文宏武参与研究方案制定;郑勇伟负责文献调研与整理;麦亮参与论文撰写与模型构建;冯霞山参与数据可视化和图表制作;钟麒深负责制定实验并分析数据;胡一鸣负责论文撰写与修改工作;郑杰辉负责论文审核工作。所有作者均阅读并同意了论文终稿内容。
/
| 〈 |
|
〉 |