|
|
|
题名
|
作者
|
年代
|
出处
|
被引量
|
| 1 | Hierarchical representation of on-chip context to reduce reconfiguration time and implementation area for coarse-grained reconfigurable architecture显示文摘In reconfigurable system,fast reconfiguration and small size of configuration contexts are strongly required to enhance the processing performance and reduce the implementation overhead.In this paper,a hierarchical representation of contexts for CGRA called HCC is proposed to satisfy the above requirements.In HCC,the contexts are constructed in a hierarchical fashion to thoroughly eliminate the repetitive portions of the contexts,not only reducing the overall contexts storage size,but also alleviating the contexts transportation overhead.The fast context-indexing mechanism is proposed in HCC to achieve high configuration speed,since the hierarchically organized contexts can be located and accessed conveniently.HCC has been verified in a reconfigurable processor called REMUS HP.Owing to HCC,when implementing H.264 decoding on REMUS HP,76.67%of the overall contexts are reduced compared with the traditional non-hierarchical one;and the configuration speed is averagely 23×increased compared with the latest reported optimized configuration mechanism on Virtex-4 FX60.REMUS HP is implemented on a 48.9 mm2silicon with TSMC 65 nm technology.Simulation shows that 1920×1088@30 fps could be achieved for H.264 high-profile decoding when exploiting a 200 MHz working frequency.Compared with the high performance version of XPP,the performance is 181%boosted. | WANG YanSheng LIU LeiBo YIN ShouYi ZHU Min CAO Peng YANG Jun WEI ShaoJun | 2013 | Science China(Information Sciences)2013,56,11: | 7 |
| 2 | An efficient VLSI architecture of speeded-up robust feature extraction for high resolution and high frame rate video显示文摘This paper proposes a VLSI architecture of the optimized Speeded-Up Robust Feature (SURF) algorithm. The SURF algorithm which is widely used in computer vision applications, locates interest points (IPoints) and extracts feature descriptors based on the surrounding gradients. In the proposed approach, the SURF algorithm is modified to make it more suitable for hardware implementation with little accuracy lost compared to OpenSURF, an open source software implementation based on OpenCV. The resource cost and throughput gain are balanced. Word Length Reduction (WLR) is adopted to compress the integral image and reduce the occupied memory resources. The orientations and the feature descriptors of the IPoints are calculated in a more efficient way, which significantly reduces memory accesses. Moreover, the operations are pipelined both in and among the proposed hardware modules. This VLSI architecture has been validated on FPGA (Xilinx Virtex-4 XC4VLX80), which is able to detect IPoints and extract SURF feature descriptors in a VGA (640 × 480) 64 fps video input with a 96 MHz working frequency while dissipating 1.278 W. This throughput is more than double of the ones reported in the latest literatures. | ZHANG WeiLong LIU LeiBo YIN ShouYi ZHOU RenYan CAI ShanShan WEI ShaoJun | 2013 | Science China(Information Sciences)2013,56,7: | 6 |
| 3 | Vertical Distribution Characteristics of PM2.5 Observed by a Mobile Vehicle Lidar in Tianjin, China in 2016显示文摘We present mobile vehicle lidar observations in Tianjin, China during the spring, summer, and winter of 2016. Mobile observations were carried out along the city border road of Tianjin to obtain the vertical distribution characteristics of PM_(2.5). Hygroscopic growth was not considered since relative humidity was less than 60% during the observation experiments. PM_(2.5) profile was obtained with the linear regression equation between the particle extinction coefficient and PM_(2.5) mass concentration. In spring, the vertical distribution of PM_(2.5) exhibited a hierarchical structure. In addition to a layer of particles that gathered near the ground, a portion of particles floated at 0.6–2.5-km height. In summer and winter, the fine particles basically gathered below 1 km near the ground. In spring and summer, the concentration of fine particles in the south was higher than that in the north because of the influence of south wind. In winter, the distribution of fine particles was opposite to that measured during spring and summer. High concentrations of PM_(2.5) were observed in the rural areas of North Tianjin with a maximum of 350 μg m–3 on 13 December2016. It is shown that industrial and ship emissions in spring and summer and coal combustion in winter were the major sources of fine particles that polluted Tianjin. The results provide insights into the mechanisms of haze formation and the effects of meteorological conditions during haze–fog pollution episodes in the Tianjin area. | Lihui LYU Yunsheng DONG Tianshu ZHANG Cheng LIU Wenqing LIU Zhouqing XIE Yan XIANG Yi ZHANG Zhenyi CHEN Guangqiang FAN Leibo ZHANG Yang LIU Yuchen SHI Xiaowen SHU | 2018 | Journal of Meteorological Research2018,32,1: | 5 |
| 4 | Row-based configuration mechanism for a 2-D processing element array in coarse-grained reconfigurable architecture显示文摘Using the coarser operand grain and simplified interconnection patterns, CGRA(coarse grained reconfigurable architectures) has been proven to be energy efficient in several specific domains. As we know, the speed at which the contexts are applied to a PEA(processing element array) directly determines the performance of CGRA. In this paper, the design space in CGRA is further developed from the configuration granularity perspective by one middle-grained configuration granularity—the row-based configuration mechanism(RCM).The most prominent feature of the RCM is that a large DFG(data flow graph) can be mapped onto a small array in once reconfiguration, which is carried out on a row-by-row basis. Compared with an ordinary DFGpartitioning solution, the reconfiguration time and the data transfer time are well reduced. Furthermore, the proposed RCM offers much more efficient storage for the contexts. Compared with the DFG partitioning solution,the performance is boosted from 2.6% to 57.8%, while the area penalty is only 4.79% and the power penalty is only 7.22%. The RCM has been used in one reconfigurable processor called REMUS HPA(reconfigurable multi-media system, high performance version advanced). REMUS HPA has been implemented on a 50.5 mm2 silicon with TSMC 65 nm technology. Simulation shows that 1920×1088@37 fps can be achieved for H.264high-profile decoding when exploiting a 200 MHz working frequency. Compared with the high performance version of XPP(one commercial reconfigurable processor), the performance is 247% boosted. | LIU LeiBo WANG YanSheng YIN ShouYi ZHU Min WANG Xing WEI ShaoJun | 2014 | Science China(Information Sciences)2014,57,10: | 3 |
| 5 | Dynamically reconfigurable architecture for symmetric ciphers显示文摘In this paper, a very large scale integration(VLSI) architecture for a reconfigurable cryptographic processor is presented. Several optimization methods have been introduced into the design process. The interconnection tree between rows(ICTR) method reduces the interconnection complexity and results in a small area overhead. The hierarchical context organization(HCO) scheme reduces the total context size and increases the dynamic configuration speed. Most symmetric ciphers, including AES, DES, SHACAL-1, SMS4, and ZUC, can be implemented using the proposed architecture. Experimental results show that the proposed architecture has obvious advantages over current state-of-the-art architectures reported in the literature in terms of performance,area efficiency(throughput/area) and energy efficiency(throughput/power). | Bo WANG Leibo LIU | 2016 | Science China(Information Sciences)2016,59,4: | 2 |
| 6 | ReSSIM:a mixed-level simulator for dynamic coarse-grained reconfigurable processor显示文摘This paper proposes a mixed-level simulator for dynamic coarse-grained reconfigurable processor(CGRP),called ReSSIM(reconfigurable system simulation implementation mechanism),and the corresponding simulation tool-chain,including task compiler,profiler and debugger.A generic modeling methodology supporting convenient extension of on-chip modules is also proposed.In order to explore the details of the interested modules while maintaining reasonable simulation speed,RCA(reconfigurable computing array),the key reconfigurable device in ReSSIM,is modeled on cycle-accurate level,while the other modules are modeled on transaction level.The typical parameters of RCA are scalable and adjustable,which helps the architects to explore the massive details of the reconfigurable device.Experiment shows that simulation speedup achieved ranges from 9.26× to 18.39× compared with VCS(Synopsys verilog compiler simulator) when running three computingintensive kernel tasks of H.264 decoding algorithm-IDCT(inverse discrete cosine transform),deblocking and MC-chroma(motion compensation).Simulation speed for a set of real applications,such as MPEG4,G.729 and EFR,is 35× slower than the corresponding native executions(i.e.measured from the real chip).And the relative simulation errors are 11% less than the measured IPC(instructions per cycle) of the real chip. | LIU LeiBo JIA Wen YIN ShouYi WANG Dong SUN GuanYi TANG Eugene WEI ShaoJun | 2013 | Science China(Information Sciences)2013,56,6: | 2 |
| 7 | Hierarchical Optimization of Multi-objective Embedded System Design Using Pareto-frontier Algebra显示文摘 | HAN Muhua LIU Leibo WEI Shaojun | 2008 | Chinese Journal of Electronics2008,17,3: | 2 |
| 8 | Olfactory and gustatory dysfunctions of COVID-19 patients in China:A multicenter study显示文摘Introduction:With the spread of the epidemic worldwide,an increasing number of doctors abroad have observed the following atypical symptoms of coronavirus disease 2019(COVID-19):olfactory or taste disorders.Therefore,clarifying the incidence and clinical characteristics of olfactory and taste disorders in Chinese COVID-19 patients is of great significance and urgency.Materials and Methods:A retrospective study was conducted,which included 229 severe acute respiratory syndrome coronavirus 2 confirmed patients,through face-to-face interviews and telephone follow-up.Following the completion of questionnaires,the patients participating in the study,were categorized according to the degree of olfactory and taste disorders experienced,and the proportion of each clinical type of patient with olfactory and taste disorders and the time when symptoms appeared were recorded.Results:Among the 229 patients,31(13.54%)had olfactory dysfunction,and 44(19.21%)had gustatory dysfunction.For the patients with olfactory dysfunction,6(19.35%)developed severe disease and became critically ill.Olfactory dysfunction appeared before the other symptoms in 21.43%of cases.The proportion of females with olfactory and gustatory dysfunction was higher than that of males(P<0.001).Conclusions:The incidence of olfactory and gustatory dysfunction was much lower than that reported abroad;the prognosis of patients with olfactory dysfunction is relatively favorable;olfactory and gustatory dysfunction can be used as a sign for early screening;females are more prone to olfactory and gustatory dysfunction. | Jianhui Li Yi Sun Enqiang Qin Hu Yuan Mingbo Liu Wenqi Yi Zhu Chen Chengcheng Huang Fengjie Zhou Ruiyao Chen Leibo Zhang Ning Yu Qiong Liu Xuejun Zhou Jingjing He Boyu Li Fusheng Wang Changliang Yang Shiming Yang | 2022 | World Journal of Otorhinolaryngology-Head and Neck Surgery2022,8,4: | 2 |
| 9 | MicroRNA-101 inhibits human hepatocellular carcinoma progression through EZH2 downregulation and increased cytostatic drug sensitivity显示文摘 | Leibo Xu Susanne Beckebaum Speranta Iacob Gang Wu Gernot M. Kaiser Arnold Radtke Chao Liu Iyad Kabar Hartmut H. Schmidt Xiaoyong Zhang Mengji Lu Vito R. Cicinnati | 2013 | Journal of Hepatology2013,,: | 2 |
| 10 | Architecture, challenges and applications of dynamic reconfigurable computing显示文摘As a computing paradigm that combines temporal and spatial computations,dynamic reconfigurable computing provides superiorities of flexibility,energy efficiency and area efficiency,attracting interest from both academia and industry.However,dynamic reconfigurable computing is not yet mature because of several unsolved problems.This work introduces the concept,architecture,and compilation techniques of dynamic reconfigurable computing.It also discusses the existing major challenges and points out its potential applications. | Yanan Lu Leibo Liu Jianfeng Zhu Shouyi Yin Shaojun Wei | 2020 | Journal of Semiconductors2020,41,2: | 2 |
| 11 | A VLS1 architecture of JPEG2000 encoder显示文摘 | LIU Leibo CHEN Ning MENG Hongying | 2004 | lEEE Journal of Solid-State Circuits2004,39,11: | 1 |
| 12 | A VLSI Architecture for the Node of Wireless Image Sensor Network显示文摘 | ZHOU Renyan LIU Leibo YIN Shouyi LUO Ao CHEN Xinkai WEI Shaojun | 2011 | Chinese Journal of Electronics2011,20,4: | 1 |
| 13 | A cycle-accurate simulator for a re- configurable multi-media system 显示文摘 | ZHU Min LIu Leibo YIN Shouyi YIN Chongyong WEI Shaojun | 2010 | The Institute of Electronics Information and Communication En- gineers Transactions on Information and Systems2010,93,12: | 1 |
| 14 | Reliability-aware mapping for various No C topologies and routing algorithms under performance constraints显示文摘The flexibility of manycore systems to extensive applications is achieved by reconfiguring the interconnections between processing elements(PEs) and the function of PEs. The efficiency of the system is crucially determined by the mapping technique of applications. In this paper, a highly flexible reliability-aware application mapping approach is proposed for manycore network-on-chip(No C) systems. A reliability cost model(RCM) is first presented to measure the reliability cost for a mapping pattern. This model uses the binary number 0/1 to model the reliability cost of each communication path. The overall reliability cost of a mapping pattern is evaluated by taking the cost of each path as a discrete random variable. Based on RCM,a mapping method called reliability cost ratio based branch and bound(RCRBB) is used. With this method,the best mapping among all the possible patterns is found efficiently by discarding those nonoptimal candidate mappings at early stages. The proposed application mapping approach with reliability awareness is applicable to various No C topologies and routing algorithms, while other state-of-the-art approaches on the same topic are only limited to a specific topology and routing algorithm. Even for the same No C architecture, the proposed approach shows significant superiority in many aspects. Experiments show that RCRBB achieves up to 9.07%reliability enhancement on average. Also, it outperforms other approaches in throughput and latency with a relatively low run time. | WU Chen DENG ChenChen LIU LeiBo YIN ShouYi HAN Jie WEI ShaoJun | 2015 | Science China(Information Sciences)2015,58,8: | 1 |
| 15 | Phthalate esters (PAEs) in indoor PM 10 /PM 2.5 and human exposure to PAEs via inhalation of indoor air in Tianjin, China显示文摘 | Leibo Zhang Fumei Wang Yaqin Ji Jiao Jiao Dekun Zou Lingling Liu Chunyan Shan Zhipeng Bai Zengrong Sun | 2013 | Atmospheric Environment2013,,: | 1 |
| 16 | A fast face detection architecture for auto-focus in smart-phones and digital cameras显示文摘Auto-focus is very important for capturing sharp human face centered images in digital and smart phone cameras. With the development of image sensor technology, these cameras support more and more highresolution images to be processed. Currently it is difficult to support fast auto-focus at low power consumption on high-resolution images. This work proposes an efficient architecture for an Ada Boost-based face-priority auto-focus. The architecture supports block-based integral image computation to improve the processing speed on high-resolution images; meanwhile, it is reconfigurable so that it enables the sub-window adaptive cascade classification, which greatly improves the processing speed and reduces power consumption. Experimental results show that 96% detection rate in average and 58 fps(frame per second) detection speed are achieved for the1080p(1920×1080) images. Compared with the state-of-the-art work, the detection speed is greatly improved and power consumption is largely reduced. | Peng OUYANG Shouyi YIN Chenchen DENG Leibo LIU Shaojun WEI | 2016 | Science China Earth Sciences2016,59,12: | 1 |
| 17 | Implementation of AVS Jizhun decoder with HW/SW partitioning on a coarse-grained reconfigurable multimedia system显示文摘In this paper,a TPP(Task-based Parallelization and Pipelining)scheme is proposed to implement AVS(Audio Video coding Standard)video decoding algorithm on REMUS(REconfigurable MUltimedia System),which is a coarse-grained reconfigurable multimedia system.An AVS decoder has been implemented with the consideration of HW/SW optimized partitioning.Several parallel techniques,such as MB(Macro-Block)-based parallel and block-based parallel techniques,and several pipeline techniques,such as MB level pipeline and block level pipeline techniques are adopted by hardware implementation,for performance improvement of the AVS decoder.Also,most computation-intensive tasks in AVS video standards,such as MC(Motion Compensation),IP(Intra Prediction),IDCT(Inverse Discrete Cosine Transform),REC(REConstruct)and DF(Deblocking Filter),are performed in the two RPUs(Reconfigurable Processing Units),which are the major computing engines of REMUS.Owing to the proposed scheme,the decoder introduced here can support AVS JP(Jizhun Profile)1920×1088@39fps streams when exploiting a 200 MHz working frequency. | LIU LeiBo CHEN YingJie YIN ShouYi ZHOU Li YUAN Hang WEI ShaoJun | 2014 | Science China(Information Sciences)2014,57,8: | 1 |
| 18 | Implementation of multi-standard video decoder on a heterogeneous coarse-grained reconfigurable processor显示文摘This paper proposes a task-based hybrid parallel and hybrid pipeline(THPHP)scheme to implement multi-standard video algorithms,including MPEG-2,H.264,and audio video coding standard(AVS),on a heterogeneous coarse-grained reconfigurable processor,called the reconfigurable multimedia system(REMUS).The proposed schemes greatly improve decoding performance and satisfy the real-time requirements of various high-definition(HD)video decoding standards.In THPHP,we propose both a task-based hybrid parallel scheme,in which macro-block(MB)-level,block-level,and sub-block-level decoding tasks are parallelized to improve data processing throughput,and a hybrid pipeline scheme,in which slice-level,MB-level,block-level and sub-block-level computations are pipelined to improve efficiency.Computation-intensive tasks,such as motion compensation,intra prediction,inverse discrete cosine transform,reconstruction,and deblocking filter,are implemented on two reconfigurable processing units,which are the core computing engines of REMUS.Thanks to the proposed schemes,the implementations can achieve H.264 high profile(HP)1920×1080@30 fps streams,AVS Jizhun profile(JP)1920×1080@39 fps streams,and MPEG-2 main profile(MP)1920×1080@41 fps streams when working at 200 MHz frequency.Compared with XPP-III(a commercial reconfigurable processor),when implementing H.264 HD decoding,the performance and energy efficiency on REMUS are improved by1.81×and 14.3×,respectively. | LIU LeiBo CHEN YingJie WANG Dong YIN ShouYi WANG Xing WANG Long LEI Hao CAO Peng WEI ShaoJun | 2014 | Science China(Information Sciences)2014,57,8: | 1 |
| 19 | A cycle-accurate simulator for a reconfigurable multi-media显示文摘 | Zhu Min Liu Leibo Yin Shouyi | 2010 | IEICE Trans on Information and Systems2010,93,12: | 1 |
| 20 | Optimization of speeded-up robust feature algorithm for hardware implementation显示文摘Speeded-Up Robust Feature(SURF)is a widely-used robust local gradient feature detection and description algorithm.The algorithm itself can be implemented easily on general-purpose processors.However,the software implementation of SURF cannot achieve a performance high enough to meet the practical real-time requirements.And what is more,the huge data storage and the floating point operation of SURF algorithm make it hard and onerous to design and verify corresponding hardware implementation.This paper customized a SURF algorithm for hardware implementation,which combined several optimization methods in previous literature and three approaches(named Word Length Reduction(WLR),Low Bits Abandon(LBA),and Sampling Radius Reduction(SRR)).The computation operations of the simplified and optimized SURF(P-SURF)were reduced by 50%compared with the original SURF.At the same time,the Recall and Precision of the SURF feature descriptor are only dropped by 0.31 on average in the typical testing set,which are within an acceptable accuracy range.P-SURF has been implemented on hardware using TSMC 65 nm process,and the architecture of the whole system mainly contains four modules,including Integral Image Generator,IPoint Detector,IPoint Orientation Assigner,and IPoint Feature Vector Extractor.The chip size is 3.4×4 mm2.The power usage is less than 220mW according to the Synopsys Prime time while extracting IPoints in a video input of VGA(640×480)172 fps operating at 200 MHz.The performance is better than the results reported in literature. | CAI ShanShan LIU LeiBo YIN ShouYi ZHOU RenYan ZHANG WeiLong WEI ShaoJun | 2014 | Science China(Information Sciences)2014,57,4: | 1 |