Search PubMed⌕ Search

Biomedical subjects

Yixue Li

Publications and source records attributed to Yixue Li.

23 records · Page 2Linked to original sources

Nucleocapsid protein of SARS coronavirus tightly binds to human cyclophilin A.

Severe acute respiratory syndrome coronavirus (SARS-CoV) is responsible for SARS infection. Nucleocapsid protein (NP) of SARS-CoV (SARS_NP) functions in enveloping the entire genomic RNA and interacts with viron structural proteins, thus playing important roles in the process of virus particle assembly and release. Protein-protein interaction analysis using bioinformatics tools indicated that SARS_NP may bind to human cyclophilin A (hCypA), and surface plasmon resonance (SPR) technology revealed this binding with the equilibrium dissociation constant ranging from 6 to 160nM. The probable binding sites of these two proteins were detected by modeling the three-dimensional structure of the SARS_NP-hCypA complex, from which the important interaction residue pairs between the proteins were deduced. Mutagenesis experiments were carried out for validating the binding model, whose correctness was assessed by the observed effects on the binding affinities between the proteins. The reliability of the binding sites derived by the molecular modeling was confirmed by the fact that the computationally predicted values of the relative free energies of the binding for SARS_NP (or hCypA) mutants to the wild-type hCypA (or SARS_NP) are in good agreement with the data determined by SPR. Such presently observed SARS_NP-hCypA interaction model might provide a new hint for facilitating the understanding of another possible SARS-CoV infection pathway against human cell.

Amino Acid Sequence↗

Identification of alternatively spliced mRNA variants related to cancers by genome-wide ESTs alignment.

Several databases have been published to predict alternative splicing of mRNAs by analysing the exon linkage relationship by alignment of expressed sequence tags (ESTs) to the genome sequence; however, little effort has been made to investigate the relationship between cancers and alternative splicing. We developed a program, Alternative Splicing Assembler (ASA), to look for splicing variants of human gene transcripts by genome-wide ESTs alignment. Using ASA, we constructed the biosino alternative splicing database (BASD), which predicted splicing variants for reference sequences from the reference sequence database (RefSeq) and presented them in both graph and text formats. EST clusters that differ from the reference sequences in at least one splicing site were counted as splicing variants. Of 4322 genes screened, 3490 (81%) were observed with at least one alternative splicing variants. To discover the variants associated with cancers, tissue sources of EST sequences were extracted from the UniLib database and ESTs from the same tissue type were counted. These were regarded as the indicators for gene expression level. Using Fisher's exact test, alternative splicing variants, of which EST counts were significantly different between cancer tissues and their counterpart normal tissues, were identified. It was predicted that 2149 variants, or 383 variants after Bonferroni correction, of 26 812 variants were likely tumor-associated. By reverse transcription-PCR, 11 of 13 novel alternative splicing variants and eight of nine variants' tissue specificity were confirmed in hepatocellular carcinoma and in lung cancer. The possible involvement of alternative splicing in cancer is discussed.

Alternative Splicing↗

A brief review of computational gene prediction methods.

With the development of genome sequencing for many organisms, more and more raw sequences need to be annotated. Gene prediction by computational methods for finding the location of protein coding regions is one of the essential issues in bioinformatics. Two classes of methods are generally adopted: similarity based searches and ab initio prediction. Here, we review the development of gene prediction methods, summarize the measures for evaluating predictor quality, highlight open problems in this area, and discuss future research directions.

Computational Biology↗

The HUPO PSI's molecular interaction format--a community standard for the representation of protein interaction data.

A major goal of proteomics is the complete description of the protein interaction network underlying cell physiology. A large number of small scale and, more recently, large-scale experiments have contributed to expanding our understanding of the nature of the interaction network. However, the necessary data integration across experiments is currently hampered by the fragmentation of publicly available protein interaction data, which exists in different formats in databases, on authors' websites or sometimes only in print publications. Here, we propose a community standard data model for the representation and exchange of protein interaction data. This data model has been jointly developed by members of the Proteomics Standards Initiative (PSI), a work group of the Human Proteome Organization (HUPO), and is supported by major protein interaction data providers, in particular the Biomolecular Interaction Network Database (BIND), Cellzome (Heidelberg, Germany), the Database of Interacting Proteins (DIP), Dana Farber Cancer Institute (Boston, MA, USA), the Human Protein Reference Database (HPRD), Hybrigenics (Paris, France), the European Bioinformatics Institute's (EMBL-EBI, Hinxton, UK) IntAct, the Molecular Interactions (MINT, Rome, Italy) database, the Protein-Protein Interaction Database (PPID, Edinburgh, UK) and the Search Tool for the Retrieval of Interacting Genes/Proteins (STRING, EMBL, Heidelberg, Germany).

Database Management Systems↗

Identification of beta-barrel membrane proteins based on amino acid composition properties and predicted secondary structure.

Unlike all-helices membrane proteins, beta-barrel membrane proteins can not be successfully discriminated from other proteins, especially from all-beta soluble proteins. This paper performs an analysis on the amino acid composition in membrane parts of 12 beta-barrel membrane proteins versus beta-strands of 79 all-beta soluble proteins. The average and variance of the amino acid composition in these two classes are calculated. Amino acids such as Gly, Asn, Val that are most likely associated with classification are selected based on Fishers discriminant ratio. A linear classifier built with these selected amino acids composition in observed beta-strands achieves 100% classification accuracy for 12 membrane proteins and 79 soluble proteins in a four-fold cross-validation experiment. Since at present the accuracy of secondary structure prediction is quite high, a promising method to identify beta-barrel membrane proteins is presented based on the linear classifier coupled with predicted secondary structure. Applied to 241 beta-barrel membrane proteins and 3855 soluble proteins with various structures, the method achieves 85.48% (206/241) sensitivity and 92.53% specificity (3567/3855).

Amino Acids↗