|
|
|
题名
|
作者
|
年代
|
出处
|
被引量
|
| 1 | Characterizing and annotating the genome using RNA-seq data显示文摘Bioinformatics methods for various RNA-seq data analyses are in fast evolution with the improvement of sequencing technologies. However, many challenges still exist in how to efficiently process the RNA-seq data to obtain accurate and comprehensive results. Here we reviewed the strategies for improving diverse transcriptomic studies and the annotation of genetic variants based on RNA-seq data. Mapping RNA-seq reads to the genome and transcriptome represent two distinct methods for quantifying the expression of genes/transcripts. Besides the known genes annotated in current databases, many novel genes/transcripts(especially those long noncoding RNAs) still can be identified on the reference genome using RNA-seq. Moreover, owing to the incompleteness of current reference genomes, some novel genes are missing from them. Genome-guided and de novo transcriptome reconstruction are two effective and complementary strategies for identifying those novel genes/transcripts on or beyond the reference genome. In addition, integrating the genes of distinct databases to conduct transcriptomics and genetics studies can improve the results of corresponding analyses. | Geng Chen Tieliu Shi Leming Shi | 2017 | Science China(Life Sciences)2017,60,2: | 22 |
| 2 | Class I histone deacetylases are major histone decrotonylases: evidence for critical and broad function of histone crotonylation in transcription显示文摘为 histone crotonylation 的酶和读者蛋白质上的最近的研究在抄写支持 histone crotonylation 的功能。然而,为 histone decrotonylation (HDCR ) 负责的酶仍然保持糟糕定义。而且,如果 histone crotonylation 从生理地重要、机能上地不同或对 histone acetylation 冗余,它尚待坚定。这里我们一级 histone deacetylases (HDAC ) 而非 sirtuin 家庭 deacetylases (SIRT ) 是主要 histone decrotonylases 的现在的证据,和那 histone crotonylation 象在哺乳动物的房间的 histone acetylation 一样动态。尤其是,我们与损害 HDAC 产生了新奇 HDAC1 和 HDAC3 异种但是未经触动的 HDCR 活动。用这些异种,我们证明在哺乳动物的房间的选择 HDCR 与宽广 transcriptional 压抑和 crotonylation 然而并非 acetylation 读者蛋白质的减少的倡导者协会相关。而且,我们证明那 histone crotonylation 被充实在并且为老鼠的自强要求了胚胎的干细胞。 | Wei Wei Xiaoguang Liu Jiwei Chen Shennan Gao Lu Lu Huifang Zhang Guangjin Ding Zhiqiang Wang Zhongzhou Chen Tieliu Shi Jiwen Li Jianjun Yu Jiemin Wong | 2017 | Cell Research2017,27,7: | 18 |
| 3 | Overview of available methods for diverse RNA-Seq data analyses显示文摘RNA-Seq technology is becoming widely used in various transcriptomics studies;however,analyzing and interpreting the RNA-Seq data face serious challenges.With the development of high-throughput sequencing technologies,the sequencing cost is dropping dramatically with the sequencing output increasing sharply.However,the sequencing reads are still short in length and contain various sequencing errors.Moreover,the intricate transcriptome is always more complicated than we expect.These challenges proffer the urgent need of efficient bioinformatics algorithms to effectively handle the large amount of transcriptome sequencing data and carry out diverse related studies.This review summarizes a number of frequently-used applications of transcriptome sequencing and their related analyzing strategies,including short read mapping,exon-exon splice junction detection,gene or isoform expression quantification,differential expression analysis and transcriptome reconstruction. | CHEN Geng WANG Charles SHI TieLiu | 2011 | Science China(Life Sciences)2011,54,12: | 16 |
| 4 | The Challenge and promise of rare disease diagnosis in China显示文摘Rare diseases are chronic and serious,featuring early onset at birth or in childhood,rapid deterioration and high mortality rate,which creates a burden on society and public health systems.Of the known rare diseases,80 percent are genetic in origin,and half of those affected worldwide are children.In China,the rare disease patients are over 10 million,and70 percent of the patients are children(Song et al.,2012;Liu et al.,2010). | Xin Ni Tieliu Shi | 2017 | Science China(Life Sciences)2017,60,7: | 4 |
| 5 | De novo transcriptome assembly of RNA-Seq reads with different strategies显示文摘De novo transcriptome assembly is an important approach in RNA-Seq data analysis and it can help us to reconstruct the transcriptome and investigate gene expression profiles without reference genome sequences.We carried out transcriptome assemblies with two RNA-Seq datasets generated from human brain and cell line,respectively.We then determined an efficient way to yield an optimal overall assembly using three different strategies.We first assembled brain and cell line transcriptome using a single k-mer length.Next we tested a range of values of k-mer length and coverage cutoff in assembling.Lastly,we combined the assembled contigs from a range of k values to generate a final assembly.By comparing these assembly results,we found that using only one k-mer value for assembly is not enough to generate good assembly results,but combining the contigs from different k-mer values could yield longer contigs and greatly improve the overall assembly. | CHEN Geng YIN KangPing WANG Charles SHI TieLiu | 2011 | Science China(Life Sciences)2011,54,12: | 4 |
| 6 | Next-generation sequencing technologies for personalized medicine:promising but challenging显示文摘In the past several years,next-generation sequencing(NGS) technologies have greatly revolutionized our approaches to explore and depict the characteristics and functions of the genomes for various species.The NGS technologies have been broadly used in diverse fields including genomics(genome sequencing and exome sequencing) [1,2],transcriptomics(RNA-Seq) [3,4] and epigenomics(ChIP-Seq, | CHEN Geng SHI TieLiu | 2013 | Science China(Life Sciences)2013,56,2: | 4 |
| 7 | Prediction and systematic study of protein-protein interaction networks of Leptospira interrogans显示文摘Leptospira interrogans serovar Lai is a pathogenic bacterium that causes a spirochetal zoonosis in humans and some animals. With its complete genome sequence available, it is possible to analyze protein-protein interactions from a whole- genome standpoint. Here we combine four recently developed computational approaches (gene fusion method, gene neighbor method, phylogenetic profiles method, and operon method) to predict protein-pro- tein interaction networks of Leptospira interrogans strain Lai. Through comprehensive analysis on in- teractions among proteins of motility and chemotaxis system, signal transduction, lipopolysaccaride bio- synthesis and a series of proteins related to adhesion and invasion, we provided information for further studying on its pathogenic mechanism. In addition, we also assigned 203 previously uncharacterized proteins with possible functions based on the known functions of its interacting partners. This work is helpful for further investigating L. interrogans strain Lai. | SUN Jingchun XU Jinlin CAO Jianping LIU Qi GUO Xiaokui SHI Tieliu LI Yixue | 2006 | Chinese Science Bulletin2006,51,11: | 3 |
| 8 | Lack of correlation between aristolochic acid exposure and hepatocellular carcinoma显示文摘Besides upper tract urothelial cell carcinoma(UTUCs),a recent study published in Science Translational Medicine has indicated that liver cancer may be associated with the exposure of aristolochic acids and similar derivatives(collectively,AA).However,according to our research,this study needs more number of samples for further verification which should be sampled from a wider range of people. | Xiangjun Ji Guoshuang Feng Geng Chen Tieliu Shi | 2018 | Science China(Life Sciences)2018,61,6: | 3 |
| 9 | Towards efficiency in rare disease research: what is distinctive and important?显示文摘Characterized by their low prevalence, rare diseases are often chronically debilitating or life threatening. Despite their low prevalence, the aggregate number of individuals suffering from a rare disease is estimated to be nearly 400 million worldwide.Over the past decades, efforts from researchers, clinicians, and pharmaceutical industries have been focused on both the diagnosis and therapy of rare diseases. However, because of the lack of data and medical records for individual rare diseases and the high cost of orphan drug development, only limited progress has been achieved. In recent years, the rapid development of next-generation sequencing(NGS)-based technologies, as well as the popularity of precision medicine has facilitated a better understanding of rare diseases and their molecular etiology. As a result, molecular subclassification can be identified within each disease more clearly, significantly improving diagnostic accuracy. However, providing appropriate care for patients with rare diseases is still an enormous challenge. In this review, we provide a brief introduction to the challenges of rare disease research and make suggestions on where and how our efforts should be focused. | Jinmeng Jia Tieliu Shi | 2017 | Science China(Life Sciences)2017,60,7: | 3 |
| 10 | Proteomics provides individualized options of precision medicine for patients with gastric cancer显示文摘While precision medicine driven by genome sequencing has revolutionized cancer care,such as lung cancer,its impact on gastric cancer(GC)has been minimal.GC patients are routinely treated with chemotherapy,but only a fraction of them receive the clinical benefit.There is an urgent need to develop biomarkers or algorithms to select chemo-sensitive patients or apply targeted therapy.Here,we carried out retrospective analyses of 1,020 formalin-fixed,paraffin-embedded GC surgical resection samples from 5 hospitals and developed a mass spectrometry-based workflow for proteomic subtyping of GC.We identified two proteomic subtypes:the chemo-sensitive group(CSG)and the chemo-insensitive group(CIG)in the discovery set.The 5-year overall survival of CSG was significantly improved in patients who had received adjuvant chemotherapy after surgery compared with those who received surgery only(64.2%vs.49.6%;Cox P-value=0.002),whereas no such improvement was observed in CIG(50.0%vs.58.6%;Cox P-value=0.495).We validated these results in an independent validation set.Further,differential proteome analysis uncovered 9 FDA-approved drugs that may be applicable for targeted therapy of GC.A prospective study is warranted to test these findings for future GC patient care. | Wenwen Huang Dongdong Zhan Yazhuo Li Nairen Zheng Xin Wei Bin Bai Kecheng Zhang Mingwei Liu Xuefei Zhao Xiaotian Ni Xia Xia Jinwen Shi Cheng Zhang Zhihao Lu Jiafu Ji Juan Wang Shiqi Wang Gang Ji Jipeng Li Yongzhan Nie Wenquan Liang Xiaosong Wu Jianxin Cui Yongsheng Meng Feilin Cao Tieliu Shi Weimin Zhu Yi Wang Lin Chen Qingchuan Zhao Hongwei Wang Lin Shen Jun Qin | 2021 | Science China(Life Sciences)2021,64,8: | 3 |
| 11 | HCCNet: an integrated network database of hepatocellular carcinoma显示文摘 | Bing He Xiaojie Qiu Peng Li Lishan Wang Qi Lv Tieliu Shi | 2010 | Cell Research2010,20,6: | 2 |
| 12 | Comparative analysis of whole-genome sequences of Streptococcus suis显示文摘The outbreak of Streptococcus suis re-cently in some districts of Sichuan Province in China has caused over 30 deaths and over 200 infections in human beings. In order to study the pathogenicity mechanism and to prevent the bacteria from spreading and infecting human beings and swine, we have annotated and analyzed the genomes of two strains, Streptococcus suis P1/7 and 89-1591 re-spective1y. The whole length of P1/7 is 2.007 Mb, and has 1969 ORFs. In contrast, the partial genome sequence of 89-1591 is 1.98 Mb in length and exists in 177 contigs with 1918 ORFs. Analysis shows that the average lengths of CDSs in two genomes are very close, and the numbers of the homolog ORFs are 1306 between those two strains. Most of the tox-icity factors of the two strains are homologeous, but there are still some significant differences between those two strains. For example, among the 11 genes (cps2A―cps2K) encoding for the capsules in P1/7, 4 (cps2A, 2B, 2I, 2J) are not detected in strain 89-1591. At the same time, the genes encoding EF and Haemolysin in P1/7 are also not found in strain 89-1591. Besides, the genes related to DNA replica-tion, repair and recombination differ from each other significantly and there also exist certain differences among the surface proteins. Those characteristics indicate that those two strains have evolved their ownspecific functions to adapt to the different environ-ments and that the pathogenesis of the two strains is different. We have accumulated comprehensive ge-nomics information for future systematic studies of S. sui. Our results are helpful for disease prevention, vaccine development, as well as drug design for S. suis. | WEI Wu DING Guohui WANG Xiaojing SUN dingchun TU Kang HAO Pei WANG Chuan CAO Zhiwei SHI Tieliu LI Yixue | 2006 | Chinese Science Bulletin2006,51,10: | 2 |
| 13 | Novel loci and potential mechanisms of major depressive disorder,bipolar disorder,and schizophrenia显示文摘Different psychiatric disorders share genetic relationships and pleiotropic loci to certain extent.We integrated and analyzed datasets related to major depressive disorder(MDD),bipolar disorder(BIP),and schizophrenia(SCZ)from the Psychiatric Genomics Consortium using multitrait analysis of genome-wide association analysis(MTAG).MTAG significantly increased the effective sample size from 99,773 to 119,754 for MDD,from 909,061 to 1,450,972 for BIP,and from 856,677 to 940,613 for SCZ.We discovered 7,32,and 43 novel lead single nucleotide polymorphisms(SNPs)and 1,6,and 3 novel causal SNPs for MDD,BIP,and SCZ,respectively,after fine-mapping.We identified rs8039305 in the FURIN gene as a novel pleiotropic locus across the three disorders.We performed marker analysis of genomic annotation(MAGMA)and Hi-C-coupled MAGMA(H-MAGMA)based gene-set analysis and identified 101 genes associated with the three disorders,which were enriched in the regulation of postsynaptic membranes,postsynaptic membrane dopaminergic synapses,and Notch signaling pathway.Next,we performed Mendelian randomization analysis using different tools and detected a causal effect of BIP on SCZ.Overall,we demonstrated the usage of combined genome-wide association studies summary statistics for exploring potential novel mechanisms of the three psychiatric disorders,providing an alternative approach to integrate publicly available summary data. | He Wang Zhenghui Yi Tieliu Shi | 2022 | Science China(Life Sciences)2022,65,1: | 2 |
| 14 | Analysis and application of large-scale protein-protein interaction data sets显示文摘Protein-protein interactions play key roles in cells. Lots of experimental approaches and in silico methods have been developed to identify and predict large-scale pro- tein-protein interactions. However, compared with the tradi- tionally experimental results, the high-throughput pro- tein-protein interaction data often contain the false positives in high probability. In order to fully utilize the large-scale data, it is necessary to develop bioinformatic methods for systematically evaluating those data in order to further im- prove the data reliability and mine biological information. This review summarizes the methodologies of analysis and application of high-throughput protein-protein interaction data, including the evaluation methods, the relationship be- tween protein-protein interaction data and other protein biological information, and their applications in biological study. In addition, this paper also suggests some interesting topics on mining high-throughput protein-protein interaction data. | SUN Jingchun XU Jinlin LI Yixue SHI Tieliu | 2005 | Chinese Science Bulletin2005,50,20: | 2 |
| 15 | Comparative study of de novo assembly and genome-guided assembly strategies for transcriptome reconstruction based on RNA-Seq显示文摘Transcriptome reconstruction is an important application of RNA-Seq,providing critical information for further analysis of transcriptome.Although RNA-Seq offers the potential to identify the whole picture of transcriptome,it still presents special challenges.To handle these difficulties and reconstruct transcriptome as completely as possible,current computational approaches mainly employ two strategies:de novo assembly and genome-guided assembly.In order to find the similarities and differences between them,we firstly chose five representative assemblers belonging to the two classes respectively,and then investigated and compared their algorithm features in theory and real performances in practice.We found that all the methods can be reduced to graph reduction problems,yet they have different conceptual and practical implementations,thus each assembly method has its specific advantages and disadvantages,performing worse than others in certain aspects while outperforming others in anther aspects at the same time.Finally we merged assemblies of the five assemblers and obtained a much better assembly.Additionally we evaluated an assembler using genome-guided de novo assembly approach,and achieved good performance.Based on these results,we suggest that to obtain a comprehensive set of recovered transcripts,it is better to use a combination of de novo assembly and genome-guided assembly. | LU BingXin ZENG ZhenBing SHI TieLiu | 2013 | Science China(Life Sciences)2013,56,2: | 2 |
| 16 | Identifying and annotating human bifunctional RNAs reveals their versatile functions显示文摘Bifunctional RNAs that possess both protein-coding and noncoding functional properties were less explored and poorly understood. Here we systematically explored the characteristics and functions of such human bifunctional RNAs by integrating tandem mass spectrometry and RNA-seq data. We first constructed a pipeline to identify and annotate bifunctional RNAs,leading to the characterization of 132 high-confidence bifunctional RNAs. Our analyses indicate that bifunctional RNAs may be involved in human embryonic development and can be functional in diverse tissues. Moreover, bifunctional RNAs could interact with multiple miRNAs and RNA-binding proteins to exert their corresponding roles. Bifunctional RNAs may also function as competing endogenous RNAs to regulate the expression of many genes by competing for common targeting miRNAs. Finally,somatic mutations of diverse carcinomas may generate harmful effect on corresponding bifunctional RNAs. Collectively,our study not only provides the pipeline for identifying and annotating bifunctional RNAs but also reveals their important gene-regulatory functions. | Geng Chen Juan Yang Jiwei Chen Yunjie Song Ruifang Cao Tieliu Shi Leming Shi | 2016 | Science China(Life Sciences)2016,59,10: | 1 |
| 17 | Predicting rRNA-, RNA-, and DNA-binding proteins from primary structure with support vector machines显示文摘 | Xiaojing Yu Jianping Cao Yudong Cai Tieliu Shi Yixue Li | 2005 | Journal of Theoretical Biology2005,,2: | 1 |
| 18 | Acidic domains differentially read histone H3 lysine 4 methylation status and are widely present in chromatin-associated proteins显示文摘Histone methylation is believed to provide binding sites for specific reader proteins, which translate histone code into biological function. Here we show that a family of acidic domain-containing proteins including nucleophosmin (NPM1), pp32, SET/TAF1β, nucleolin (NCL) and upstream binding factor (UBF) are novel H3K4me2-binding proteins. These proteins exhibit a unique pattern of interaction with methylated H3K4, as their binding is stimulated by H3K4me2 and inhibited by H3K4me1 and H3K4me3. These proteins contain one or more acidic domains consisting mainly of aspartic and/or glutamic residues that are necessary for preferential binding of H3K4me2. Furthermore, we demonstrate that the acidic domain with sufficient length alone is capable of binding H3K4me2 in vitro and in vivo. NPM1, NCL and UBF require their acidic domains for association with and transcriptional activation of rDNA genes. Interestingly, by defining acidic domain as a sequence with at least 20 acidic residues in 50 continuous amino acids, we identified 655 acidic domain-containing protein coding genes in the human genome and Gene Ontology (GO) analysis showed that many of the acidic domain proteins have chromatin-related functions. Our data suggest that acidic domain is a novel histone binding motif that can differentially read the status of H3K4 methylation and is broadly present in chromatin-associated proteins. | Meng Wu Wei Wei Jiwei Chen Rong Cong Tieliu Shi Jiwen Li Jiemin Wong James X.Du | 2017 | Science China(Life Sciences)2017,60,2: | 1 |
| 19 | Significant variations in alternative splicing patterns and expression profiles between human-mouse orthologs in early embryos显示文摘Human and mouse orthologs are expected to have similar biological functions; however, many discrepancies have also been reported. We systematically compared human and mouse orthologs in terms of alternative splicing patterns and expression profiles. Human-mouse orthologs are divergent in alternative splicing, as human orthologs could generally encode more isoforms than their mouse orthologs. In early embryos, exon skipping is far more common with human orthologs, whereas constitutive exons are more prevalent with mouse orthologs. This may correlate with divergence in expression of splicing regulators. Orthologous expression similarities are different in distinct embryonic stages, with the highest in morula. Expression differences for orthologous transcription factor genes could play an important role in orthologous expression discordance. We further detected largely orthologous divergence in differential expression between distinct embryonic stages. Collectively, our study uncovers significant orthologous divergence from multiple aspects, which may result in functional differences and dynamics between human-mouse orthologs during embryonic development. | Geng Chen Jiwei Chen Jianmin Yang Long Chen Xiongfei Qu Caiping Shi Baitang Ning Leming Shi Weida Tong Yongxiang Zhao Meixia Zhang Tieliu Shi | 2017 | Science China(Life Sciences)2017,60,2: | 1 |
| 20 | Sequencing XMET genes to promote genotype-guided risk assessment and precision medicine显示文摘High-throughput next generation sequencing (NGS) is a shotgun approach applied in a parallel fashion by which the genome is fragmented and sequenced through small pieces and then analyzed either by aligning to a known reference genome or by de novo assembly without reference genome.This technology has led researchers to conduct an explosion of sequencing related projects in multidisciplinary fields of science.However,due to the limitations of sequencing-based chemistry,length of sequencing reads and the complexity of genes,it is difficult to determine the sequences of some portions of the human genome,leaving gaps in genomic data that frustrate further analysis.Particularly,some complex genes are difficult to be accurately sequenced or mapped because they contain high GC-content and/or low complexity regions,and complicated pseudogenes,such as the genes encoding xenobiotic metabolizing enzymes and transporters (XMETs).The genetic variants in XMET genes are critical to predicate interindividual variability in drug efficacy,drug safety and susceptibility to environmental toxicity.We summarized and discussed challenges,wet-lab methods,and bioinformatics algorithms in sequencing 'complex' XMET genes,which may provide insightful information in the application of NGS technology for implementation in toxicogenomics and pharmacogenomics. | Yaqiong Jin Geng Chen Wenming Xiao Huixiao Hong Joshua Xu Yongli Guo Wenzhong Xiao Tieliu Shi Leming Shi Weida Tong Baitang Ning | 2019 | Science China(Life Sciences)2019,62,7: | 1 |