维普中文期刊产品整合服务

Deep Model Compression for Mobile Platforms:A Survey

查看全文 作  者:Kaiming [1]Nan;Sicong [1]Liu;Junzhao [2]Du;Hui [1]Liu 高影响力作者 机构地区:[1]School of Computer Science and Technology,Xidian University,Xi'an 710071,China;[2]School of Software and Institute of Software Engineering,Xidian University,Xi'an 710071,China高影响力机构 出  处:《Tsinghua Science and Technology》索引2019年第24卷第6期,共17页高影响力期刊 基  金:supported by the National Key Research and Development Program of China (No. 2018YFB1003605);Foundations of CARCH (No. CARCH201704);the National Natural Science Foundation of China (No. 61472312);Foundations of Shaanxi Province and Xi’an Science;Technology Plan (Nos. B018230008 and BD34017020001);the Foundations of Xidian University (No. JBZ171002) 摘  要:Despite the rapid development of mobile and embedded hardware, directly executing computationexpensive and storage-intensive deep learning algorithms on these devices’ local side remains constrained for sensory data analysis. In this paper, we first summarize the layer compression techniques for the state-of-theart deep learning model from three categories: weight factorization and pruning, convolution decomposition, and special layer architecture designing. For each category of layer compression techniques, we quantify their storage and computation tunable by layer compression techniques and discuss their practical challenges and possible improvements. Then, we implement Android projects using TensorFlow Mobile to test these 10 compression methods and compare their practical performances in terms of accuracy, parameter size, intermediate feature size,computation, processing latency, and energy consumption. To further discuss their advantages and bottlenecks,we test their performance over four standard recognition tasks on six resource-constrained Android smartphones.Finally, we survey two types of run-time Neural Network(NN) compression techniques which are orthogonal with the layer compression techniques, run-time resource management and cost optimization with special NN architecture,which are orthogonal with the layer compression techniques. 关 键 词:DEEP learning MODEL compression run-time RESOURCE management COST optimization
相关文献

参考文献(44)

引证文献(8)

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费