维普中文期刊产品整合服务
共被期刊论文引用了2次 您的检索式:您选中1篇文献正在查看引证文献汇总
    题名 作者 年代 出处 被引量
1Navigation for autonomous vehicles via fast-stable and smooth reinforcement learning显示文摘This paper investigates the navigation problem of autonomous vehicles based on reinforcement learning(RL)with both stability and smoothness guarantees.By introducing a data-based Lyapunov function,the stability criterion in mean cost is obtained,where the Lyapunov function has a property of fast descending.Then,an off-policy RL algorithm is proposed to train safe policies,in which a more strict constraint is exerted in the framework of model-free RL to ensure the fast convergence of policy generation,in contrast with the existing RL merely with stability guarantee.In addition,by simultaneously introducing constraints on action increments and action distribution variations,the difference between the adjacent actions is effectively alleviated to ensure the smoothness of the obtained policy,instead of only seeking the similarity of the distributions of adjacent actions as commonly done in the past literature.A navigation task of a ground differentially driven mobile vehicle in simulations is adopted to demonstrate the superiority of the proposed algorithm on the fast stability and smoothness.ZHANG RuiXian YANG JiaNan LIANG Ye LU ShengAo DONG YiFei YANG BaoQing ZHANG LiXian 2024Science China(Technological Sciences)2024,67,2:0
2Robust reinforcement learning with UUB guarantee for safe motion control of autonomous robots显示文摘This paper addresses the issue of safety in reinforcement learning(RL)with disturbances and its application in the safety-constrained motion control of autonomous robots.To tackle this problem,a robust Lyapunov value function(rLVF)is proposed.The rLVF is obtained by introducing a data-based LVF under the worst-case disturbance of the observed state.Using the rLVF,a uniformly ultimate boundedness criterion is established.This criterion is desired to ensure that the cost function,which serves as a safety criterion,ultimately converges to a range via the policy to be designed.Moreover,to mitigate the drastic variation of the rLVF caused by differences in states,a smoothing regularization of the rLVF is introduced.To train policies with safety guarantees under the worst disturbances of the observed states,an off-policy robust RL algorithm is proposed.The proposed algorithm is applied to motion control tasks of an autonomous vehicle and a cartpole,which involve external disturbances and variations of the model parameters,respectively.The experimental results demonstrate the effectiveness of the theoretical findings and the advantages of the proposed algorithm in terms of robustness and safety.ZHANG RuiXian HAN YiNing SU Man LIN ZeFeng LI HaoWei ZHANG LiXian 2024Science China(Technological Sciences)2024,67,1:0
返回顶部 每页显示:
共1页 首页 上一页 第1页 下一页 末页 /1 跳转

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费