Search PubMed⌕ Search

Biomedical subjects

Xun Gu

Publications and source records attributed to Xun Gu.

35 records · Page 2Linked to original sources

How much expression divergence after yeast gene duplication could be explained by regulatory motif evolution?

We used the yeast genome sequences of gene families, microarray profiles and regulatory motif data to test the current wisdom that there is a strong correlation between regulatory motif structure and gene expression profile. Our results suggest that duplicate genes tend to be co-expressed but the correlation between motif content and expression similarity is generally poor, only approximately 2-3% of expression variation can be explained by the motif divergence. Our observations suggest that, in addition to the cis-regulatory motif structure in the upstream region of the gene, multiple trans-acting factors in the gene network can influence the pattern of gene expression significantly.

Evolution, Molecular↗

Further statistical analysis for genome-wide expression evolution in primate brain/liver/fibroblast tissues.

In spite of only a 1-2 per cent genomic DNA sequence difference, humans and chimpanzees differ considerably in behaviour and cognition. Affymetrix microarray technology provides a novel approach to addressing a long-term debate on whether the difference between humans and chimpanzees results from the alteration of gene expressions. Here, we used several statistical methods (distance method, two-sample t-tests, regularised t-tests, ANOVA and bootstrapping) to detect the differential expression pattern between humans and great apes. Our analysis shows that the pattern we observed before is robust against various statistical methods; that is, the pronounced expression changes occurred on the human lineage after the split from chimpanzees, and that the dramatic brain expression alterations in humans may be mainly driven by a set of genes with increased expression (up-regulated) rather than decreased expression (down-regulated).

Animals↗

Statistical framework for phylogenomic analysis of gene family expression profiles.

Microarray technology has produced massive expression data that are invaluable for investigating the genome-wide evolutionary pattern of gene expression. To this end, phylogenetic expression analysis is highly desirable. On the basis of the Brownian process, we developed a statistical framework (called the E(0) model), assuming the independent expression of evolution between lineages. Several evolutionary mechanisms are integrated to characterize the pattern of expression diversity after gene duplications, including gradual drift and dramatic shift (punctuated equilibrium). When the phylogeny of a gene family is given, we show that the likelihood function follows a multivariate normal distribution; the variance-covariance matrix is determined by the phylogenetic topology and evolutionary parameters. Maximum-likelihood methods for multiple microarray experiments are developed, and likelihood-ratio tests are designed for testing the evolutionary pattern of gene expression. To reconstruct the evolutionary trace of expression diversity after gene (or genome) duplications, we developed a Bayesian-based method and use the posterior mean as predictors. Potential applications in evolutionary genomics are discussed.

Algorithms↗

Natural history and functional divergence of protein tyrosine kinases.

Cellular signaling is important for many biological processes including growth, differentiation, adhesion, motility and apoptosis. The protein tyrosine kinase (PTK) supergene family is the key mediator in cellular signaling in metazoans, directly associated with a variety of human diseases. All PTKs contain a highly conserved catalytic kinase domain, in spite of variable multi-domain structures. Within each PTK gene family, members exhibit functional divergence in substrate-specificity or temporal/tissue-specific expression, although their primary function is conserved. After conducting phylogenetic analysis on major PTK gene families, we found that the expanding of each PTK family was likely caused by gene or genome duplication event(s) that occurred before the emergence of teleosts but after the vertebrate-amphioxus split. We further investigated the evolutionary pattern of functional divergence after gene duplication in those gene families. Our results show that site-specific shifted evolutionary rate (altered functional constraint) is a common pattern in PTK gene family evolution.

Animals↗

Role of duplicate genes in genetic robustness against null mutations.

Deleting a gene in an organism often has little phenotypic effect, owing to two mechanisms of compensation. The first is the existence of duplicate genes: that is, the loss of function in one copy can be compensated by the other copy or copies. The second mechanism of compensation stems from alternative metabolic pathways, regulatory networks, and so on. The relative importance of the two mechanisms has not been investigated except for a limited study, which suggested that the role of duplicate genes in compensation is negligible. The availability of fitness data for a nearly complete set of single-gene-deletion mutants of the Saccharomyces cerevisiae genome has enabled us to carry out a genome-wide evaluation of the role of duplicate genes in genetic robustness against null mutations. Here we show that there is a significantly higher probability of functional compensation for a duplicate gene than for a singleton, a high correlation between the frequency of compensation and the sequence similarity of two duplicates, and a higher probability of a severe fitness effect when the duplicate copy that is more highly expressed is deleted. We estimate that in S. cerevisiae at least a quarter of those gene deletions that have no phenotype are compensated by duplicate genes.

Evolution, Molecular↗

Induced gene expression in human brain after the split from chimpanzee.

Despite only approximately 1% difference in genomic DNA sequence, humans and chimpanzees differ considerably in mental and linguistic capabilities, and in susceptibility to some diseases. A recent comparison of gene expression in human and great apes cast some light on the genetic basis of these differences, but more rigorous study is required. Our statistical reanalysis of these microarray data shows that there have indeed been dramatic alterations in the expression of genes in the human brain since the split from chimpanzees, mainly caused by a set of genes with increased (rather than decreased) expression in the human brain.

Animals↗

Algorithms for multiple genome rearrangement by signed reversals.

We discuss a multiple genome rearrangement problem by signed reversals: Given a collection of genomes, we generate them in the minimum number of signed reversals. It is NP-hard and equivalent to finding an optimal Steiner tree to connect the genomes by reversal paths. We design two algorithms to find the optimal Steiner nodes of the problem: Neighbor-perturbing algorithm and branch-and-bound algorithm. The first one is a polynomial running time approximation algorithm. It searches for the optimal Steiner nodes by perturbing initial Steiner nodes nearby their neighborhoods and improving them better and better until convergence. The second one is an exact exponential running time algorithm for a median problem. It finds the optimal Steiner node by checking all candidates that satisfy the necessary conditions for optimal Steiner nodes. We implement the algorithms into two programs respectively and show by experimental examples that they are more efficient than other similar ones, such as GRAPPA, BPAnalysis, and MGR, etc.

Algorithms↗

Functional divergence in protein (family) sequence evolution.

As widely used today to infer 'function', the homology search is based on the neutral theory that sites of greatest functional significance are under the strongest selective constraints as well as lowest evolutionary rates, and vice versa. Therefore, site-specific rate changes (or altered selective constraints) are related to functional divergence during protein (family) evolution. In this paper, we review our recent work about this issue. We show a great deal of functional information can be obtained from the evolutionary perspective, which can in turn be used to facilitate high throughput functional assays. The emergence of evolutionary functional genomics is also indicated. The related software DIVERGE can be obtained from http://xgu1.zool.iastate.edu.

Amino Acid Sequence↗

Age distribution of human gene families shows significant roles of both large- and small-scale duplications in vertebrate evolution.

The classical (two-round) hypothesis of vertebrate genome duplication proposes two successive whole-genome duplication(s) (polyploidizations) predating the origin of fishes, a view now being seriously challenged. As the debate largely concerns the relative merits of the 'big-bang mode' theory (large-scale duplication) and the 'continuous mode' theory (constant creation by small-scale duplications), we tested whether a significant proportion of paralogous genes in the contemporary human genome was indeed generated in the early stage of vertebrate evolution. After an extensive search of major databases, we dated 1,739 gene duplication events from the phylogenetic analysis of 749 vertebrate gene families. We found a pattern characterized by two waves (I, II) and an ancient component. Wave I represents a recent gene family expansion by tandem or segmental duplications, whereas wave II, a rapid paralogous gene increase in the early stage of vertebrate evolution, supports the idea of genome duplication(s) (the big-bang mode). Further analysis indicated that large- and small-scale gene duplications both make a significant contribution during the early stage of vertebrate evolution to build the current hierarchy of the human proteome.

Animals↗

Evolutionary analysis for functional divergence of Jak protein kinase domains and tissue-specific genes.

Jak (Janus kinase) is a nonreceptor tyrosine kinase, which plays important roles in signal transduction pathways. The unique feature of Jak is that, in addition to a fully functional tyrosine kinase domain (JH1), Jak possesses a pseudokinase domain (JH2). Although JH2 lost its catalytic function, experimental evidence has shown that this domain may have acquired some new but unknown functions. This apparent functional divergence after the (internal) domain duplication may result in dramatic changes of selective constraints at some sites. We conducted a data analysis to test this hypothesis. Our result shows that shifted selective constraints (or shifted evolutionary rates) between the JH1 and the JH2 domains are statistically significant. Predicted amino acid sites by posterior analysis can be classified into two groups: very conserved in JH1 but highly variable in JH2, and vice versa. Moreover, we have studied the evolutionary pattern of four tissue-specific genes, Jak1, Jak2, Jak3, and Tyk2, which were generated in the early stages of vertebrates. We found that after the (first) gene duplication, site-specific rate shifts between Jak2/Jak3 and Jak1/Tyk are significant, presumably as a consequence of functional divergence among these genes. The implication of our study for functional genomics is discussed.

Amino Acids↗

Predicting functional divergence in protein evolution by site-specific rate shifts.

Most modern tools that analyze protein evolution allow individual sites to mutate at constant rates over the history of the protein family. However, Walter Fitch observed in the 1970s that, if a protein changes its function, the mutability of individual sites might also change. This observation is captured in the "non-homogeneous gamma model", which extracts functional information from gene families by examining the different rates at which individual sites evolve. This model has recently been coupled with structural and molecular biology to identify sites that are likely to be involved in changing function within the gene family. Applying this to multiple gene families highlights the widespread divergence of functional behavior among proteins to generate paralogs and orthologs.

Amino Acid Sequence↗

DIVERGE: phylogeny-based analysis for functional-structural divergence of a protein family.

SUMMARY: DetectIng Variability in Evolutionary Rates among GEnes (DIVERGE) is a software system to study functional divergence of a protein family by detecting site-specific change in evolutionary rate using a multiple alignment of amino acid sequences for a given phylogenetic tree. The program first conducts a statistical test for site-specific rate shifts along the tree, and predicting candidate amino acid residues responsible for functional divergence based on posterior analysis. These results can then be mapped on the 3D protein structure if available. AVAILABILITY: DIVERGE is available free of charge from http://xgu1.zool.iastate.edu/. Distribution packages for both Linux and Microsoft Windows operating systems are available, including manual and example files.

Amino Acid Sequence↗

Identification of essential amino acid changes in paired domain evolution using a novel combination of evolutionary analysis and in vitro and in vivo studies.

Pax genes are defined by the presence of a paired box that encodes a DNA-binding domain of 128 amino acids. They are involved in the development of the central nervous system, organogenesis, and oncogenesis. The known Pax genes are divided into five groups within two supergroups. By means of a novel combination of evolutionary analysis, in vitro binding assays and in vivo functional analyses, we have identified the key residues that determine the differing DNA-binding properties of the two supergroups and of the Pax-2, 5, 8 and Pax-6 subgroups within supergroup I. The differences in binding properties between the two supergroups are largely caused by amino acid changes at residues 20 and 121 of the paired domain. Although the paired domains of the Pax-2, 5, 8 and the Pax-6 group differ by >19 amino acids, their distinct DNA-binding properties are determined almost completely by a single amino acid change. Thus, a small number of amino acid changes can account in large part for the divergence in binding properties among the known paired domains. Our approach for selecting candidate sites responsible for the functional divergence between genes should also be useful for studying other gene families.

Amino Acid Sequence↗

Novel PAX6 binding sites in the human genome and the role of repetitive elements in the evolution of gene regulation.

Pax6 is a critical transcription factor in the development of the eye, pancreas, and central nervous system. It is composed of two DNA-binding domains, the paired domain (PD), which has two helix-turn-helix (HTH) motifs, and the homeodomain (HD), made up from another HTH motif. Each HTH motif can bind to DNA separately or in combination with the others. We identified three novel binding sites that are specific for the PD and HD domains of human PAX6 from single-copy human genomic DNA libraries using cyclic amplification of protein binding sequences (CAPBS) and electrophoretic mobility shift assays (EMSAs). One of the binding sites was found within sequences of repetitive Alu elements. However, most of the Alu sequences were unable to bind to PAX6 because of a small number of mismatches (mostly in CpG dinucleotide hot spots) in the consensus Alu sequences. PAX6 binding Alu elements are found primarily in old and intermediate-aged Alu subfamilies. These data along with our previously identified B1-type Pax6 binding site showed that evolutionarily conserved Pax6 has target sites that are disparate in primates and rodents. This difference indicates that human and mouse Pax6-regulated gene networks may have evolved through these lineage-specific repeat elements.

Alu Elements↗

Multiple genome rearrangement by reversals.

In this paper, we discuss a multiple genome rearrangement problem: Given a collection of genomes represented by permutations, we generate the collection from some fixed genome, e.g., the identity permutation, in a minimum number of signed reversals. It is NP-hard, so efficient heuristics is important for finding its optimal solution. We at first discuss how to generate two and three genomes from a fixed genome by polynomial algorithms for some special cases. Then based on the polynomial algorithms, we obtain some approximation algorithms for generating two and three genomes in general, respectively. Finally, we apply these approximation algorithms to design a new approximation algorithm for generating more genomes. We also show by some experimental examples that the algorithms are efficient.

Algorithms↗