维普中文期刊产品整合服务
2篇 您的检索式:作者名="Fengbin Tu"
    题名 作者 年代 出处 被引量
1SWG:an architecture for sparse weight gradient computation显示文摘On-device training for deep neural networks(DNN)has become a trend due to various user preferences and scenarios.The DNN training process consists of three phases,feedforward(FF),backpropagation(BP),and weight gradient(WG)update.WG takes about one-third of the computation in the whole training process.Current training accelerators usually ignore the special computation property of WG and process it in a way similar to FF/BP.Besides,the extensive data sparsity existing in WG,which brings opportunities to save computation,is not well explored.Nevertheless,exploiting the optimization opportunities would meet three underutilization problems,which are caused by(1)the mismatch between WG data dimensions and hardware parallelism,(2)the full sparsity,i.e.,the sparsity of feature map(Fmap),error map(Emap),and gradient,and(3)the workload imbalance resulting from irregular sparsity.In this paper,we propose a specific architecture for sparse weight gradient(SWG)computation.The architecture is designed based on hierarchical unrolling and sparsity-aware(HUSA)dataflow to exploit the optimization opportunities of the special computation property and full data sparsity.In HUSA dataflow,the data dimensions are unrolled hierarchically on the hardware architecture.A valid-data trace(VDT)mechanism is embedded in the dataflow to avoid the underutilization caused by the two-sided input sparsity.The gradient is unrolled in PE to alleviate the underutilization induced by output sparsity while maintaining the data reuse opportunities.Besides,we design an intra-and inter-column balancer(IIBLC)to dynamically tackle the workload imbalance problem resulting from the irregular sparsity.Experimental results show that with HUSA dataflow exploiting the full sparsity,SWG achieves a speedup of 12.23×over state-of-the-art gradient computation architecture,TrainWare.SWG helps to improve the energy efficiency of the state-of-the-art training accelerator LNPU from 7.56 to 10.58 TOPS/W.Weiwei WU Fengbin TU Xiangyu LI Shaojun WEI Shouyi YIN 2024Science China(Information Sciences)2024,67,2:0
2Towards efficient generative AI and beyond-AI computing:New trends on ISSCC 2024 machine learning accelerators显示文摘Compared to the last decade when the convolution neu-ral network(CNN)dominated the research field,machine learn-ing(ML)algorithms have reached a pivotal moment called the generative artificial intelligence(AI)era.With the emer-gence of large-scale foundation models[1],such as large multi-modal model(LMM)GPT-4[2]and text-to-image generative model DALL·E[3].Bohan Yang Jia Chen Fengbin Tu 2024Journal of Semiconductors2024,45,4:0
返回顶部 每页显示:
共1页 首页 上一页 第1页 下一页 末页 /1 跳转

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费