Search PubMed⌕ Search

Biomedical subjects

Yixue Li

Publications and source records attributed to Yixue Li.

At least 19 recordsLinked to original sources

Phylogenetic profiles for the prediction of protein-protein interactions: how to select reference organisms?

The phylogenetic profile method has been widely applied in the prediction of protein-protein interactions (PPIs). Studies often use all of the available complete genomes for this method. With more than 400 genomes complete and new ones on the horizon, it remains unclear how to select reference organisms for profile construction and then influence the PPI prediction. Here, we performed a systematic assessment of reference organism selection from 225 complete genomes with their evolutionary tree. Our results suggest that reference organisms should be selected from moderately and highly genetically distant organisms, from all three domains (Bacteria, Archaea, and Eukarya), and by their even distribution at the fifth hierarchical level in the evolutionary tree. Our study provides important guidance on the construction of phylogenetic profiles for PPI prediction and functional genomics, which has become challenging due to the large and increasing number of available candidate organisms.

Algorithms↗

Analysis of the dermatophyte Trichophyton rubrum expressed sequence tags.

BACKGROUND: Dermatophytes are the primary causative agent of dermatophytoses, a disease that affects billions of individuals worldwide. Trichophyton rubrum is the most common of the superficial fungi. Although T. rubrum is a recognized pathogen for humans, little is known about how its transcriptional pattern is related to development of the fungus and establishment of disease. It is therefore necessary to identify genes whose expression is relevant to growth, metabolism and virulence of T. rubrum. RESULTS: We generated 10 cDNA libraries covering nearly the entire growth phase and used them to isolate 11,085 unique expressed sequence tags (ESTs), including 3,816 contigs and 7,269 singletons. Comparisons with the GenBank non-redundant (NR) protein database revealed putative functions or matched homologs from other organisms for 7,764 (70%) of the ESTs. The remaining 3,321 (30%) of ESTs were only weakly similar or not similar to known sequences, suggesting that these ESTs represent novel genes. CONCLUSION: The present data provide a comprehensive view of fungal physiological processes including metabolism, sexual and asexual growth cycles, signal transduction and pathogenic mechanisms.

Arthrodermataceae↗

Identification and analysis of the mouse basic/Helix-Loop-Helix transcription factor family.

The basic/Helix-Loop-Helix (bHLH) proteins are a family of transcription factors that regulates a variety of biological processes. Based on a previously defined consensus motif, we identified the complete set of bHLH protein family from the mouse proteome databases and carried out a series of bioinformatics analysis. As results, 124 mouse bHLH proteins were identified in this study, and 28 of them were additional bHLH proteins beyond the previous report. These 124 mouse bHLH proteins were classified into groups from A to F by the nomenclature and phylogenetic analysis. Statistic analysis of the Gene Ontology annotation of these proteins showed that the bHLH proteins tend to perform functions related to cell differentiation and development. Gene function enrichment analysis among six groups illuminated that the proteins in certain group tend to have special biology functions, so that the molecular function of the uncharacterized proteins in groups could be inferred.

Amino Acid Sequence↗

A novel computational method to predict transcription factor DNA binding preference.

Transcription factor binds to sequence specific sites in regulatory region to control nearby gene's expression. It is termed as the major regulator of transcription. However, identifying DNA binding preference of transcription factors systematically is still a challenge. By using the nearest neighbor algorithm, a novel computational approach was developed to predict transcription factor DNA binding preference based on the gene ontology [M. Ashburner, C.A. Ball, J.A. Blake, D. Botstein, H. Butler, J.M. Cherry, A.P. Davis, K. Dolinski, S.S. Dwight, J.T. Eppig, M.A. Harris, D.P. Hill, L. Issel-Tarver, A. Kasarskis, S. Lewis, J.C. Matese, J.E. Richardson, M. Ringwald, G.M. Rubin, G. Sherlock, Gene Ontology: tool for the unification of biology, Nat. Genet. 25 (2000) 25-29.] and 0/1 encoding system of nucleotide. The overall success rate of Jackknife cross-validation test for our predictor reaches 76.6%, which indicates the DNA binding preference is closely correlated with its biological functions and computational method developed in this contribution could be a powerful tool to investigate transcription factor DNA binding preference, especially for those novel transcription factors with little prior knowledge on its DNA binding preference.

Algorithms↗

Reconstruction and in silico analysis of the MAPK signaling pathways in the human blood fluke, Schistosoma japonicum.

At present, little is known about signal transduction mechanisms in schistosomes, which cause the disease of schistosomiasis. The mitogen-activated protein kinase (MAPK) signaling pathways, which are evolutionarily conserved from yeast to Homo sapiens, play key roles in multiple cellular processes. Here, we reconstructed the hypothetical MAPK signaling pathways in Schistosoma japonicum and compared the schistosome pathways with those of model eukaryote species. We identified 60 homologous components in the S. japonciumMAPK signaling pathways. Among these, 27 were predicted to be full-length sequences. Phylogenetic analysis of these proteins confirmed the evolutionary conservation of the MAPK signaling pathways. Remarkably, we identified S. japonicum homologues of GTP-binding protein beta and alpha-I subunits in the yeast mating pathway, which might be involved in the regulation of different life stages and female sexual maturation processes as well in schistosomes. In addition, several pathway member genes, including ERK, JNK, Sja-DSP, MRAS and RAS, were determined through quantitative PCR analysis to be expressed in a stage-specific manner, with ERK, JNK and their inhibitor Sja-DSP markedly upregulated in adult female schistosomes.

Amino Acid Sequence↗

Insights into the coupling of duplication events and macroevolution from an age profile of animal transmembrane gene families.

The evolution of new gene families subsequent to gene duplication may be coupled to the fluctuation of population and environment variables. Based upon that, we presented a systematic analysis of the animal transmembrane gene duplication events on a macroevolutionary scale by integrating the palaeontology repository. The age of duplication events was calculated by maximum likelihood method, and the age distribution was estimated by density histogram and normal kernel density estimation. We showed that the density of the duplicates displays a positive correlation with the estimates of maximum number of cell types of common ancestors, and the oxidation events played a key role in the major transitions of this density trace. Next, we focused on the Phanerozoic phase, during which more macroevolution data are available. The pulse mass extinction timepoints coincide with the local peaks of the age distribution, suggesting that the transmembrane gene duplicates fixed frequently when the environment changed dramatically. Moreover, a 61-million-year cycle is the most possible cycle in this phase by spectral analysis, which is consistent with the cycles recently detected in biodiversity. Our data thus elucidate a strong coupling of duplication events and macroevolution; furthermore, our method also provides a new way to address these questions.

Animals↗

Automatic transcription factor classifier based on functional domain composition.

To understand the transcriptional regulatory mechanism, it is indispensable to identify transcription factors (TF) from the whole genome and to classify transcription factors into different classes. New computational approaches have been developed to identify TFs/non-TFs, and furthermore to classify TFs into four different classes, based on the protein functional domain composition [K.C. Chou, Y.D. Cai, Using functional domain composition and support vector machines for prediction of protein subcellular location, J. Biol. Chem. 277 (2002) 45765-45769]. We trained and tested our method on a non-redundancy dataset consisting of 74 transcription factors collected from TRANSFAC v7.0 [V. Matys, O.V. Kel-Margoulis, E. Fricke, I. Liebich, S. Land, A. Barre-Dirrie, I. Reuter, D. Chekmenev, M. Krull, K. Hornischer, N. Voss, P. Stegmaier, B. Lewicki-Potapov, H. Saxel, A.E. Kel, E. Wingender, TRANSFAC(R) and its module TRANSCompel(R): transcriptional gene regulation in eukaryotes, Nucleic Acids Res. 34 (2006) D108-D110] and 1558 non-transcription factors from UniProtKB/Swiss-Prot Release 49.3 of 21-Mar-2006. The overall success rates of jackknife cross-validation tests reached 98.4% for TF/non-TF identification and 97.2% for classifications of TF classes: basic domains, zinc-coordinating DNA-binding domains, helix-turn-helix, and beta-scaffold factors.

Artificial Intelligence↗

Demonstration of two novel methods for predicting functional siRNA efficiency.

BACKGROUND: siRNAs are small RNAs that serve as sequence determinants during the gene silencing process called RNA interference (RNAi). It is well know that siRNA efficiency is crucial in the RNAi pathway, and the siRNA efficiency for targeting different sites of a specific gene varies greatly. Therefore, there is high demand for reliable siRNAs prediction tools and for the design methods able to pick up high silencing potential siRNAs. RESULTS: In this paper, two systems have been established for the prediction of functional siRNAs: (1) a statistical model based on sequence information and (2) a machine learning model based on three features of siRNA sequences, namely binary description, thermodynamic profile and nucleotide composition. Both of the two methods show high performance on the two datasets we have constructed for training the model. CONCLUSION: Both of the two methods studied in this paper emphasize the importance of sequence information for the prediction of functional siRNAs. The way of denoting a bio-sequence by binary system in mathematical language might be helpful in other analysis work associated with fixed-length bio-sequence.

Algorithms↗

Exploring photosynthesis evolution by comparative analysis of metabolic networks between chloroplasts and photosynthetic bacteria.

BACKGROUND: Chloroplasts descended from cyanobacteria and have a drastically reduced genome following an endosymbiotic event. Many genes of the ancestral cyanobacterial genome have been transferred to the plant nuclear genome by horizontal gene transfer. However, a selective set of metabolism pathways is maintained in chloroplasts using both chloroplast genome encoded and nuclear genome encoded enzymes. As an organelle specialized for carrying out photosynthesis, does the chloroplast metabolic network have properties adapted for higher efficiency of photosynthesis? We compared metabolic network properties of chloroplasts and prokaryotic photosynthetic organisms, mostly cyanobacteria, based on metabolic maps derived from genome data to identify features of chloroplast network properties that are different from cyanobacteria and to analyze possible functional significance of those features. RESULTS: The properties of the entire metabolic network and the sub-network that consists of reactions directly connected to the Calvin Cycle have been analyzed using hypergraph representation. Results showed that the whole metabolic networks in chloroplast and cyanobacteria both possess small-world network properties. Although the number of compounds and reactions in chloroplasts is less than that in cyanobacteria, the chloroplast's metabolic network has longer average path length, a larger diameter, and is Calvin Cycle -centered, indicating an overall less-dense network structure with specific and local high density areas in chloroplasts. Moreover, chloroplast metabolic network exhibits a better modular organization than cyanobacterial ones. Enzymes involved in the same metabolic processes tend to cluster into the same module in chloroplasts. CONCLUSION: In summary, the differences in metabolic network properties may reflect the evolutionary changes during endosymbiosis that led to the improvement of the photosynthesis efficiency in higher plants. Our findings are consistent with the notion that since the light energy absorption, transfer and conversion is highly efficient even in photosynthetic bacteria, the further improvements in photosynthetic efficiency in higher plants may rely on changes in metabolic network properties.

Arabidopsis↗

Genomic characterization of ribitol teichoic acid synthesis in Staphylococcus aureus: genes, genomic organization and gene duplication.

BACKGROUND: Staphylococcus aureus or MRSA (Methicillin Resistant S. aureus), is an acquired pathogen and the primary cause of nosocomial infections worldwide. In S. aureus, teichoic acid is an essential component of the cell wall, and its biosynthesis is not yet well characterized. Studies in Bacillus subtilis have discovered two different pathways of teichoic acid biosynthesis, in two strains W23 and 168 respectively, namely teichoic acid ribitol (tar) and teichoic acid glycerol (tag). The genes involved in these two pathways are also characterized, tarA, tarB, tarD, tarI, tarJ, tarK, tarL for the tar pathway, and tagA, tagB, tagD, tagE, tagF for the tag pathway. With the genome sequences of several MRSA strains: Mu50, MW2, N315, MRSA252, COL as well as methicillin susceptible strain MSSA476 available, a comparative genomic analysis was performed to characterize teichoic acid biosynthesis in these S. aureus strains. RESULTS: We identified all S. aureus tar and tag gene orthologs in the selected S. aureus strains which would contribute to teichoic acids sythesis. Based on our identification of genes orthologous to tarI, tarJ, tarL, which are specific to tar pathway in B. subtilis W23, we also concluded that tar is the major teichoic acid biogenesis pathway in S. aureus. Further analyses indicated that the S. aureus tar genes, different from the divergon organization in B. subtilis, are organized into several clusters in cis. Most interesting, compared with genes in B. subtilis tar pathway, the S. aureus tar specific genes (tarI,J,L) are duplicated in all six S. aureus genomes. CONCLUSION: In the S. aureus strains we analyzed, tar (teichoic acid ribitol) is the main teichoic acid biogenesis pathway. The tar genes are organized into several genomic groups in cis and the genes specific to tar (relative to tag): tarI, tarJ, tarL are duplicated. The genomic organization of the S. aureus tar pathway suggests their regulations are different when compared to B. subtilis tar or tag pathway, which are grouped in two operons in a divergon structure.

Amino Acid Sequence↗

Classification of protein quaternary structure by functional domain composition.

BACKGROUND: The number and the arrangement of subunits that form a protein are referred to as quaternary structure. Quaternary structure is an important protein attribute that is closely related to its function. Proteins with quaternary structure are called oligomeric proteins. Oligomeric proteins are involved in various biological processes, such as metabolism, signal transduction, and chromosome replication. Thus, it is highly desirable to develop some computational methods to automatically classify the quaternary structure of proteins from their sequences. RESULTS: To explore this problem, we adopted an approach based on the functional domain composition of proteins. Every protein was represented by a vector calculated from the domains in the PFAM database. The nearest neighbor algorithm (NNA) was used for classifying the quaternary structure of proteins from this information. The jackknife cross-validation test was performed on the non-redundant protein dataset in which the sequence identity was less than 25%. The overall success rate obtained is 75.17%. Additionally, to demonstrate the effectiveness of this method, we predicted the proteins in an independent dataset and achieved an overall success rate of 84.11% CONCLUSION: Compared with the amino acid composition method and Blast, the results indicate that the domain composition approach may be a more effective and promising high-throughput method in dealing with this complicated problem in bioinformatics.

Algorithms↗

Recent progresses in the application of machine learning approach for predicting protein functional class independent of sequence similarity.

Protein sequence contains clues to its function. Functional prediction from sequence presents a challenge particularly for proteins that have low or no sequence similarity to proteins of known function. Recently, machine learning methods have been explored for predicting functional class of proteins from sequence-derived properties independent of sequence similarity, which showed promising potential for low- and non-homologous proteins. These methods can thus be explored as potential tools to complement alignment- and clustering-based methods for predicting protein function. This article reviews the strategies, current progresses, and underlying difficulties in using machine learning methods for predicting the functional class of proteins. The relevant software and web-servers are described. The reported prediction performances in the application of these methods are also presented, which need to be interpreted with caution as they are dependent on such factors as datasets used and choice of parameters.

Algorithms↗

Predicting O-glycosylation sites in mammalian proteins by using SVMs.

O-glycosylation is one of the most important, frequent and complex post-translational modifications. This modification can activate and affect protein functions. Here, we present three support vector machines models based on physical properties, 0/1 system, and the system combining the above two features. The prediction accuracies of the three models have reached 0.82, 0.85 and 0.85, respectively. The accuracies of the three SVMs methods were evaluated by 'leave-one-out' cross validation. This approach provides a useful tool to help identify the O-glycosylation sites in mammalian proteins. An online prediction web server is available at http://www.biosino.org/Oglyc.

Animals↗

Bioinformatics research on the SARS coronavirus (SARS_CoV) in China.

Severe acute respiratory syndrome (SARS) first appeared in 2002 in China, which fastly affected about 8000 patients over 29 countries and caused 774 fatalities. As its pathogen was identified as a new kind of coronavirus (SARS_CoV), its genome was quickly sequenced on several isolates. Studies on its functional genomics were performed by combinatorial application of all the available bioinformatics tools and the development of new programs. In this way, it was found that the four proteins were absolutely responsible for nosogenesis of SARS, i.e. spike (S) protein; small envelop (E) protein; membrane (M) protein; and nucleocaspid (N) protein. Molecular evolution studies have revealed that SARS must be originated from wild animals, and it was demonstrated that the major genetic variations in some critical genes, particularly the Spike gene, was essential for the transition from animal-to-human transmission to human-to-human transmission. Theoretical models, either Logistic model or SIR model, were developed to describe the transmission of SARS. The recorded difference of SARS spreading in Beijing and Hong Kong was also reasonably analyzed according to these models. The whole process of fruitful bioinformatics studies, along with other related scientific investigations have set up an unprecedented paradigm for human of how to battle against sudden-breaking and catastrophic epidemics.

China↗

Predicting rRNA-, RNA-, and DNA-binding proteins from primary structure with support vector machines.

In the post-genome era, the prediction of protein function is one of the most demanding tasks in the study of bioinformatics. Machine learning methods, such as the support vector machines (SVMs), greatly help to improve the classification of protein function. In this work, we integrated SVMs, protein sequence amino acid composition, and associated physicochemical properties into the study of nucleic-acid-binding proteins prediction. We developed the binary classifications for rRNA-, RNA-, DNA-binding proteins that play an important role in the control of many cell processes. Each SVM predicts whether a protein belongs to rRNA-, RNA-, or DNA-binding protein class. Self-consistency and jackknife tests were performed on the protein data sets in which the sequences identity was < 25%. Test results show that the accuracies of rRNA-, RNA-, DNA-binding SVMs predictions are approximately 84%, approximately 78%, approximately 72%, respectively. The predictions were also performed on the ambiguous and negative data set. The results demonstrate that the predicted scores of proteins in the ambiguous data set by RNA- and DNA-binding SVM models were distributed around zero, while most proteins in the negative data set were predicted as negative scores by all three SVMs. The score distributions agree well with the prior knowledge of those proteins and show the effectiveness of sequence associated physicochemical properties in the protein function prediction. The software is available from the author upon request.

Amino Acid Sequence↗

Knowledge guided analysis of microarray data.

To microarray expression data analysis, it is well accepted that biological knowledge-guided clustering techniques show more advantages than pure mathematical techniques. In this paper, Gene Ontology is introduced to guide the clustering process, and thus a new algorithm capturing both expression pattern similarities and biological function similarities is developed. Our algorithm was validated on two well-known public data sets and the results were compared with some previous works. It is shown that our method has advantages in both the quality of clusters and the precision of biological annotations. Furthermore, the clustering results can be adjusted according to different stringency requirements. It is expected that our algorithm can be extended to other biological knowledge, for example, metabolic networks.

Algorithms↗

Detecting correlation between sequence and expression divergences in a comparative analysis of human serpin genes.

Physiological functions and characteristic structures of the serpin gene superfamily have been studied extensively, yet the evolution of the serpin genes remains unclear. Gene duplication in this superfamily may shed light on this issue. Two models are used to predict the preservation of duplicated genes: the classical model and the duplication-degeneration-complementation (DDC) model. In this study, we analyzed the phylogenetic relationships of 33 human serpin genes and the expression data of some members of the serpin superfamily from a DNA microarray of human leukemia U937 cells with stably inducible expression of the leukemia-related AML1-ETO gene. We then determined the utility of the DDC model by mapping serpin superfamily expression data to the phylogenetic tree. The correlation between sequence and expression divergences as measured by the Pearson correlation coefficient indicated that human serpin genes evolved under the DDC model. Our study provides a new strategy for comparative analysis of gene sequences and microarray data.

Cell Line, Tumor↗

Refined phylogenetic profiles method for predicting protein-protein interactions.

MOTIVATION: The increasing availability of complete genome sequences provides excellent opportunity for the further development of tools for functional studies in proteomics. Several experimental approaches and in silico algorithms have been developed to cluster proteins into networks of biological significance that may provide new biological insights, especially into understanding the functions of many uncharacterized proteins. Among these methods, the phylogenetic profiles method has been widely used to predict protein-protein interactions. It involves the selection of reference organisms and identification of homologous proteins. Up to now, no published report has systematically studied the effects of the reference genome selection and the identification of homologous proteins upon the accuracy of this method. RESULTS: In this study, we optimized the phylogenetic profiles method by integrating phylogenetic relationships among reference organisms and sequence homology information to improve prediction accuracy. Our results revealed that the selection of the reference organisms set and the criteria for homology identification significantly are two critical factors for the prediction accuracy of this method. Our refined phylogenetic profiles method shows greater performance and potentially provides more reliable functional linkages compared with previous methods.

Algorithms↗