维普中文期刊产品整合服务

An Adaptive Strategy via Reinforcement Learning for the Prisoner's Dilemma Game

查看全文 作  者:Lei [1]Xue;Changyin [1,5]Sun;Donald [2,5]Wunsch;Yingjiang [3]Zhou;Fang [4]Yu 高影响力作者 机构地区:[1]Key Laboratory of Measurement and Control of Complex Systems of Engineering, Ministry of Education, School of Automation, Southeast University, Nanjing 210096, China;[2]Department of Electrical and Computer Engineering, Missouri University of Science and Technology, Rolla, MO 65409, USA;[3]College of Automation, Nanjing University of Posts and Telecommunications, Nanjing 210023, China;[4]Institute of Logistics Science and Engineering, Shanghai Maritime University, Shanghai 210096, China;[5]IEEE高影响力机构 出  处:《IEEE/CAA Journal of Automatica Sinica》索引2018年第5卷第1期,共10页高影响力期刊 基  金:supported by the National Natural Science Foundation(NNSF)of China(61603196,61503079,61520106009,61533008);the Natural Science Foundation of Jiangsu Province of China(BK20150851);China Postdoctoral Science Foundation(2015M581842);Jiangsu Postdoctoral Science Foundation(1601259C);Nanjing University of Posts and Telecommunications Science Foundation(NUPTSF)(NY215011);Priority Academic Program Development of Jiangsu Higher Education Institutions,the open fund of Key Laboratory of Measurement and Control of Complex Systems of Engineering,Ministry of Education(MCCSE2015B02);the Research Innovation Program for College Graduates of Jiangsu Province(CXLX1309) 摘  要:The iterated prisoner's dilemma(IPD) is an ideal model for analyzing interactions between agents in complex networks. It has attracted wide interest in the development of novel strategies since the success of tit-for-tat in Axelrod's tournament. This paper studies a new adaptive strategy of IPD in different complex networks, where agents can learn and adapt their strategies through reinforcement learning method. A temporal difference learning method is applied for designing the adaptive strategy to optimize the decision making process of the agents. Previous studies indicated that mutual cooperation is hard to emerge in the IPD. Therefore, three examples which based on square lattice network and scale-free network are provided to show two features of the adaptive strategy. First, the mutual cooperation can be achieved by the group with adaptive agents under scale-free network, and once evolution has converged mutual cooperation, it is unlikely to shift. Secondly, the adaptive strategy can earn a better payoff compared with other strategies in the square network. The analytical properties are discussed for verifying evolutionary stability of the adaptive strategy. 关 键 词:Complex network prisoner’s dilemma reinforcement learning temporal differences learning
相关文献

参考文献(30)

引证文献(9)

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费