维普中文期刊产品整合服务

Design of high parallel CNN accelerator based on FPGA for AIoT

查看全文 作  者:Lin [1,2]Zhijian;Gao [1]Xuewei;Chen [1]Xiaopei;Zhu [1]Zhipeng;Du [1]Xiaoyong;Chen [2]Pingping 高影响力作者 机构地区:[1]School of Advanced Manufacturing,Fuzhou University,Quanzhou 362251,China;[2]College of Physics and Information Engineering,Fuzhou University,Fuzhou 350108,China高影响力机构 出  处:《The Journal of China Universities of Posts and Telecommunications》索引2022年第29卷第5期,共10页高影响力期刊 基  金:supported by the National Natural Science Foundation of China(61871132,62171135)。 摘  要:To tackle the challenge of applying convolutional neural network(CNN)in field-programmable gate array(FPGA)due to its computational complexity,a high-performance CNN hardware accelerator based on Verilog hardware description language was designed,which utilizes a pipeline architecture with three parallel dimensions including input channels,output channels,and convolution kernels.Firstly,two multiply-and-accumulate(MAC)operations were packed into one digital signal processing(DSP)block of FPGA to double the computation rate of the CNN accelerator.Secondly,strategies of feature map block partitioning and special memory arrangement were proposed to optimize the total amount of off-chip access memory and reduce the pressure on FPGA bandwidth.Finally,an efficient computational array combining multiplicative-additive tree and Winograd fast convolution algorithm was designed to balance hardware resource consumption and computational performance.The high parallel CNN accelerator was deployed in ZU3 EG of Alinx,using the YOLOv3-tiny algorithm as the test object.The average computing performance of the CNN accelerator is 127.5 giga operations per second(GOPS).The experimental results show that the hardware architecture effectively improves the computational power of CNN and provides better performance compared with other existing schemes in terms of power consumption and the efficiency of DSPs and block random access memory(BRAMs). 关 键 词:artificial intelligence of things(AIoT) convolutional neural network(CNN)accelerator Winograd convolution field-programmable gate array(FPGA)
相关文献

参考文献(18)

耦合文献(25)

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费