维普中文期刊产品整合服务

A policy gradient algorithm integrating long and short-term rewards for soft continuum arm control

查看全文 作  者:DONG [1]Xiang;ZHANG [1]Jing;CHENG [2]Long;XU [3]WenJun;SU [3]Hang;MEI [3]Tao 高影响力作者 机构地区:[1]School of Electrical Engineering and Automation,Anhui University,Hefei 230601,China;[2]State Key Laboratory for Control and Management of Complex Systems,Institute of Automation,Chinese Academy of Sciences,Beijing 100190,China;[3]Robotics Research Center,Peng Cheng Laboratory,Shenzhen 518055,China高影响力机构 出  处:《Science China(Technological Sciences)》索引2022年第65卷第10期,共11页高影响力期刊 基  金:partially supported by the National Key Research and Development Project Monitoring and Prevention of Major Natural Disasters Special Program (Grant No. 2020YFC1512202);the Anhui University Cooperative Innovation Project (Grant No. GXXT-2019-003) 摘  要:The soft continuum arm has extensive application in industrial production and human life due to its superior safety and flexibility. Reinforcement learning is a powerful technique for solving soft arm continuous control problems, which can learn an effective control policy with an unknown system model. However, it is often affected by the high sample complexity and requires huge amounts of data to train, which limits its effectiveness in soft arm control. An improved policy gradient method, policy gradient integrating long and short-term rewards denoted as PGLS, is proposed in this paper to overcome this issue. The shortterm rewards provide more dynamic-aware exploration directions for policy learning and improve the exploration efficiency of the algorithm. PGLS can be integrated into current policy gradient algorithms, such as deep deterministic policy gradient(DDPG). The overall control framework is realized and demonstrated in a dynamics simulation environment. Simulation results show that this approach can effectively control the soft arm to reach and track the targets. Compared with DDPG and other model-free reinforcement learning algorithms, the proposed PGLS algorithm has a great improvement in convergence speed and performance. In addition, a fluid-driven soft manipulator is designed and fabricated in this paper, which can verify the proposed PGLS algorithm in real experiments in the future. 关 键 词:soft arm control Cosserat rod deep reinforcement learning policy gradient algorithm high sample complexity
相关文献

参考文献(35)

引证文献(3)

耦合文献(70)

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费