维普中文期刊产品整合服务

MSF-Net: A Multilevel Spatiotemporal Feature Fusion Network Combines Attention for Action Recognition

查看全文 作  者:Mengmeng [1]Yan;Chuang [1,2]Zhang;Jinqi [1]Chu;Haichao [1]Zhang;Tao [1]Ge;Suting [1]Chen 高影响力作者 机构地区:[1]School of Electronic and Information Engineering,Nanjing University of Information Science and Technology,Nanjing,210044,China;[2]Jiangsu Key Laboratory of Meteorological Observation and Information Processing,Nanjing,210044,China高影响力机构 出  处:《Computer Systems Science & Engineering》索引2023年第47卷第11期,共17页高影响力期刊 基  金:supported by the General Program of the National Natural Science Foundation of China (62272234);the Enterprise Cooperation Project (2022h160);the Priority Academic Program Development of Jiangsu Higher Education Institutions Project. 摘  要:An action recognition network that combines multi-level spatiotemporal feature fusion with an attention mechanism is proposed as a solution to the issues of single spatiotemporal feature scale extraction,information redundancy,and insufficient extraction of frequency domain information in channels in 3D convolutional neural networks.Firstly,based on 3D CNN,this paper designs a new multilevel spatiotemporal feature fusion(MSF)structure,which is embedded in the network model,mainly through multilevel spatiotemporal feature separation,splicing and fusion,to achieve the fusion of spatial perceptual fields and short-medium-long time series information at different scales with reduced network parameters;In the second step,a multi-frequency channel and spatiotemporal attention module(FSAM)is introduced to assign different frequency features and spatiotemporal features in the channels are assigned corresponding weights to reduce the information redundancy of the feature maps.Finally,we embed the proposed method into the R3D model,which replaced the 2D convolutional filters in the 2D Resnet with 3D convolutional filters and conduct extensive experimental validation on the small and medium-sized dataset UCF101 and the largesized dataset Kinetics-400.The findings revealed that our model increased the recognition accuracy on both datasets.Results on the UCF101 dataset,in particular,demonstrate that our model outperforms R3D in terms of a maximum recognition accuracy improvement of 7.2%while using 34.2%fewer parameters.The MSF and FSAM are migrated to another traditional 3D action recognition model named C3D for application testing.The test results based on UCF101 show that the recognition accuracy is improved by 8.9%,proving the strong generalization ability and universality of the method in this paper. 关 键 词:3D convolutional neural network action recognition MSF FSAM
相关文献

参考文献(40)

网站首页 | 关于我们 | 联系我们 | 产品服务 | 客服中心 | 广告服务 | 版权声明 | 网站联盟 | 友情链接 | 售卡网点

版权所有© 渝B2-20050021-1 渝公网安备 50019002500403号 违法和不良信息举报中心

互联网出版许可证 新出网证(渝)字10号 全国400电话 - 免长途话费