Search PubMed⌕ Search

Biomedical subjects

Kenneth H Wolfe

Publications and source records attributed to Kenneth H Wolfe.

35 records · Page 2Linked to original sources

Widespread paleopolyploidy in model plant species inferred from age distributions of duplicate genes.

It is often anticipated that many of today's diploid plant species are in fact paleopolyploids. Given that an ancient large-scale duplication will result in an excess of relatively old duplicated genes with similar ages, we analyzed the timing of duplication of pairs of paralogous genes in 14 model plant species. Using EST contigs (unigenes), we identified pairs of paralogous genes in each species and used the level of synonymous nucleotide substitution to estimate the relative ages of gene duplication. For nine of the investigated species (wheat [Triticum aestivum], maize [Zea mays], tetraploid cotton [Gossypium hirsutum], diploid cotton [G. arboretum], tomato [Lycopersicon esculentum], potato [Solanum tuberosum], soybean [Glycine max], barrel medic [Medicago truncatula], and Arabidopsis thaliana), the age distributions of duplicated genes contain peaks corresponding to short evolutionary periods during which large numbers of duplicated genes were accumulated. Large-scale duplications (polyploidy or aneuploidy) are strongly suspected to be the cause of these temporal peaks of gene duplication. However, the unusual age profile of tandem gene duplications in Arabidopsis indicates that other scenarios, such as variation in the rate at which duplicated genes are deleted, must also be considered.

Computational Biology↗

Functional divergence of duplicated genes formed by polyploidy during Arabidopsis evolution.

To study the evolutionary effects of polyploidy on plant gene functions, we analyzed functional genomics data for a large number of duplicated gene pairs formed by ancient polyploidy events in Arabidopsis thaliana. Genes retained in duplicate are not distributed evenly among Gene Ontology or Munich Information Center for Protein Sequences functional categories, which indicates a nonrandom process of gene loss. Genes involved in signal transduction and transcription have been preferentially retained, and those involved in DNA repair have been preferentially lost. Although the two members of each gene pair must originally have had identical transcription profiles, less than half of the pairs formed by the most recent polyploidy event still retain significantly correlated profiles. We identified several cases where groups of duplicated gene pairs have diverged in concert, forming two parallel networks, each containing one member of each gene pair. In these cases, the expression of each gene is strongly correlated with the other nonhomologous genes in its network but poorly correlated with its paralog in the other network. We also find that the rate of protein sequence evolution has been significantly asymmetric in >20% of duplicate pairs. Together, these results suggest that functional diversification of the surviving duplicated genes is a major feature of the long-term evolution of polyploids.

Arabidopsis↗

Evolution of the MAT locus and its Ho endonuclease in yeast species.

The genetics of the mating-type (MAT) locus have been studied extensively in Saccharomyces cerevisiae, but relatively little is known about how this complex system evolved. We compared the organization of MAT and mating-type-like (MTL) loci in nine species spanning the hemiascomycete phylogenetic tree. We inferred that the system evolved in a two-step process in which silent HMR/HML cassettes appeared, followed by acquisition of the Ho endonuclease from a mobile genetic element. Ho-mediated switching between an active MAT locus and silent cassettes exists only in the Saccharomyces sensu stricto group and their closest relatives: Candida glabrata, Kluyveromyces delphensis, and Saccharomyces castellii. We identified C. glabrata MTL1 as the ortholog of the MAT locus of K. delphensis and show that switching between C. glabrata MTL1a and MTL1alpha genotypes occurs in vivo. The more distantly related species Kluyveromyces lactis has silent cassettes but switches mating type without the aid of Ho endonuclease. Very distantly related species such as Candida albicans and Yarrowia lipolytica do not have silent cassettes. In Pichia angusta, a homothallic species, we found MATalpha2, MATalpha1, and MATa1 genes adjacent to each other on the same chromosome. Although some continuity in the chromosomal location of the MAT locus can be traced throughout hemiascomycete evolution and even to Neurospora, the gene content of the locus has changed with the loss of an HMG domain gene (MATa2) from the MATa idiomorph shortly after HO was recruited.

Deoxyribonucleases, Type II Site-Specific↗

Divergence of spatial gene expression profiles following species-specific gene duplications in human and mouse.

To examine the process by which duplicated genes diverge in function, we studied how the gene expression profiles of orthologous gene sets in human and mouse are affected by the presence of additional recent species-specific paralogs. Gene expression profiles were compared across 16 homologous tissues in human and mouse using microarray data from the Gene Expression Atlas for 1575 sets of orthologs including 250 with species-specific paralogs. We find that orthologs that have undergone recent duplication are less likely to have strongly correlated expression profiles than those that remain in a one-to-one relationship between human and mouse. There is a general trend for paralogous genes to become more specialized in their expression patterns, with decreased breadth and increased specificity of expression as gene family size increases. Despite this trend, detailed examination of some particular gene families where species-specific duplications have occurred indicated several examples of apparent neofunctionalization of duplicated genes, but only one case of subfunctionalization. Often, the expression of both copies of a duplicated gene appears to have changed relative to the ancestral state. Our results suggest that gene expression profiles are surprisingly labile and that expression in a particular tissue may be gained or lost repeatedly during the evolution of even small gene families. We conclude that gene duplication is a major driving force behind the emergence of divergent gene expression patterns.

Animals↗

Congruence of tissue expression profiles from Gene Expression Atlas, SAGEmap and TissueInfo databases.

BACKGROUND: Extracting biological knowledge from large amounts of gene expression information deposited in public databases is a major challenge of the postgenomic era. Additional insights may be derived by data integration and cross-platform comparisons of expression profiles. However, database meta-analysis is complicated by differences in experimental technologies, data post-processing, database formats, and inconsistent gene and sample annotation. RESULTS: We have analysed expression profiles from three public databases: Gene Expression Atlas, SAGEmap and TissueInfo. These are repositories of oligonucleotide microarray, Serial Analysis of Gene Expression and Expressed Sequence Tag human gene expression data respectively. We devised a method, Preferential Expression Measure, to identify genes that are significantly over- or under-expressed in any given tissue. We examined intra- and inter-database consistency of Preferential Expression Measures. There was good correlation between replicate experiments of oligonucleotide microarray data, but there was less coherence in expression profiles as measured by Serial Analysis of Gene Expression and Expressed Sequence Tag counts. We investigated inter-database correlations for six tissue categories, for which data were present in the three databases. Significant positive correlations were found for brain, prostate and vascular endothelium but not for ovary, kidney, and pancreas. CONCLUSION: We show that data from Gene Expression Atlas, SAGEmap and TissueInfo can be integrated using the UniGene gene index, and that expression profiles correlate relatively well when large numbers of tags are available or when tissue cellular composition is simple. Finally, in the case of brain, we demonstrate that when PEM values show good correlation, predictions of tissue-specific expression based on integrated data are very accurate.

Brain↗

Positive selection and subfunctionalization of duplicated CCT chaperonin subunits.

To reach a functional and energetically stable conformation, many proteins need molecular helpers called chaperonins. Among the group II chaperonins, CCT proteins provide crucial machinery for the stabilization and proper folding of several proteins in the cytosol of eukaryotic cells through interactions that are subunit-specific and geometry-dependent. CCT proteins are made up of eight different subunits, all with similar sequences, positioned in a precise arrangement. Each subunit has been proposed to have a specialized function during the binding and folding of the CCT protein substrate. Here, we demonstrate that functional divergence occurred after several CCT duplication events due to the fixation of amino acid substitutions by positive selection. Sites critical for ATP binding and substrate binding were found to have undergone positive selection and functional divergence predominantly in subunits that bind tubulin but not actin. Furthermore, we show clear functional divergence between CCT subunits that bind the C-terminal domains of actin and tubulin and those that bind the N-terminal domains. Phylogenetic analyses could not resolve the deep relationships between most subunits, except for the groups alpha/beta/eta and delta/epsilon, suggesting several almost simultaneous ancient duplication events. Together, the results support the idea that, in contrast to homo-oligomeric chaperonins such as GroEL, the high divergence level between CCT subunits is the result of positive selection after each duplication event to provide a specialized role for each CCT subunit in the different steps of protein folding.

Amino Acid Substitution↗

Wrapping up BLAST and other applications for use on Unix clusters.

UNLABELLED: We have developed two programs that speed up common bioinformatic applications by spreading them across a UNIX cluster.(1) BLAST.pm, a new module for the 'MOLLUSC' package. (2) WRAPID, a simple tool for parallelizing large numbers of small instances of programs such as BLAST, FASTA and CLUSTALW. AVAILABILITY: The packages were developed in Perl on a 20-node Linux cluster and are provided together with a configuration script and documentation. They can be freely downloaded from http://wolfe.gen.tcd.ie/wrapper.

Computer Communication Networks↗

Evidence from comparative genomics for a complete sexual cycle in the 'asexual' pathogenic yeast Candida glabrata.

BACKGROUND: Candida glabrata is a pathogenic yeast of increasing medical concern. It has been regarded as asexual since it was first described in 1917, yet phylogenetic analyses have revealed that it is more closely related to sexual yeasts than other Candida species. We show here that the C. glabrata genome contains many genes apparently involved in sexual reproduction. RESULTS: By genome survey sequencing, we find that genes involved in mating and meiosis are as numerous in C. glabrata as in the sexual species Kluyveromyces delphensis, which is its closest known relative. C. glabrata has a putative mating-type (MAT) locus and a pheromone gene (MFALPHA2), as well as orthologs of at least 31 other Saccharomyces cerevisiae genes that have no known roles apart from mating or meiosis, including FUS3, IME1 and SMK1. CONCLUSIONS: We infer that C. glabrata is likely to have an undiscovered sexual stage in its life cycle, similar to that recently proposed for C. albicans. The two Candida species represent two distantly related yeast lineages that have independently become both pathogenic and 'asexual'. Parallel evolution in the two lineages as they adopted mammalian hosts resulted in separate but analogous switches from overtly sexual to cryptically sexual life cycles, possibly in response to defense by the host immune system.

Candida glabrata↗

Molecular evolution meets the genomics revolution.

Changes in technology in the past decade have had such an impact on the way that molecular evolution research is done that it is difficult now to imagine working in a world without genomics or the Internet. In 1992, GenBank was less than a hundredth of its current size and was updated every three months on a huge spool of tape. Homology searches took 30 minutes and rarely found a hit. Now it is difficult to find sequences with only a few homologs to use as examples for teaching bioinformatics. For molecular evolution researchers, the genomics revolution has showered us with raw data and the information revolution has given us the wherewithal to analyze it. In broad terms, the most significant outcome from these changes has been our newfound ability to examine the evolution of genomes as a whole, enabling us to infer genome-wide evolutionary patterns and to identify subsets of genes whose evolution has been in some way atypical.

Animals↗

A recent polyploidy superimposed on older large-scale duplications in the Arabidopsis genome.

The Arabidopsis genome contains numerous large duplicated chromosomal segments, but the different approaches used in previous analyses led to different interpretations regarding the number and timing of ancestral large-scale duplication events. Here, using more appropriate methodology and a more recent version of the genome sequence annotation, we investigate the scale and timing of segmental duplications in Arabidopsis. We used protein sequence similarity searches to detect duplicated blocks in the genome, used the level of synonymous substitution between duplicated genes to estimate the relative ages of the blocks containing them, and analyzed the degree of overlap between adjacent duplicated blocks. We conclude that the Arabidopsis lineage underwent at least two distinct episodes of duplication. One was a polyploidy that occurred much more recently than estimated previously, before the Arabidopsis/Brassica rapa split and probably during the early emergence of the crucifer family (24-40 Mya). An older set of duplicated blocks was formed after the monocot/dicot divergence, and the relatively low level of overlap among these blocks indicates that at least some of them are remnants of a larger duplication such as a polyploidy or aneuploidy.

Arabidopsis↗

The 2R hypothesis and the human genome sequence.

One theory formalised in 1970 proposes that the complexity of vertebrate genomes originated by means of genome duplication at the base of the vertebrate lineage. Since then, the theory has remained both popular and controversial. Here we review the theory, and present preliminary results from our analysis of duplications in the draft human genome sequence. We find evidence for extensive duplication of parts of the genome. We also question the validity of the 'parsimony test' that has been used in other analyses.

Animals↗

Evolutionary re-organisation of a large operon in adzuki bean chloroplast DNA caused by inverted repeat movement.

We have sequenced two sections of chloroplast DNA from adzuki bean (Vigna angularis), containing the junctions between the inverted repeat (IR) and large single copy (LSC) regions of the genome. The gene order at both junctions is different from that described for other members of the legume family, such as Lotus japonicus and soybean. These differences have been attributed to an apparent 78-kb inversion that spans nearly the entire LSC region and which is present in adzuki and its close relative, the common bean. This 78-kb rearrangement broke the large S10 operon of ribosomal proteins into two smaller operons, one at each end of the LSC, without affecting the gene content of the genome. It disrupted the physical and transcriptional relationship between the six-gene rpl23-rpl14 cluster and the four-gene rps8-rpoA cluster that is conserved in most land plants. Analysis of the endpoints of the rearrangement indicates that it probably occurred by means of a two-step process of expansion and contraction of the IR and not by a 78-kb inversion.

Base Sequence↗

Gene order evolution and paleopolyploidy in hemiascomycete yeasts.

The wealth of comparative genomics data from yeast species allows the molecular evolution of these eukaryotes to be studied in great detail. We used "proximity plots" to visually compare chromosomal gene order information from 14 hemiascomycetes, including the recent Génolevures survey, to Saccharomyces cerevisiae. Contrary to the original reports, we find that the Génolevures data strongly support the hypothesis that S. cerevisiae is a degenerate polyploid. Using gene order information alone, 70% of the S. cerevisiae genome can be mapped into "sister" regions that tile together with almost no overlap. This map confirms and extends the map of sister regions that we constructed previously by using duplicated genes, an independent source of information. Combining gene order and gene duplication data assigns essentially the whole genome into sister regions, the largest gap being only 36 genes long. The 16 centromere regions of S. cerevisiae form eight pairs, indicating that an ancestor with eight chromosomes underwent complete doubling; alternatives such as segmental duplications can be ruled out. Gene arrangements in Kluyveromyces lactis and four other species agree quantitatively with what would be expected if they diverged from S. cerevisiae before its polyploidization. In contrast, Saccharomyces exiguus, Saccharomyces servazzii, and Candida glabrata show higher levels of gene adjacency conservation, and more cases of imperfect conservation, suggesting that they split from the S. cerevisiae lineage after polyploidization. This finding is confirmed by sequences around the C. glabrata TRP1 and IPP1 loci, which show that it contains sister regions derived from the same duplication event as that of S. cerevisiae.

Ascomycota↗

Extensive genomic duplication during early chordate evolution.

Opinions on the hypothesis that ancient genome duplications contributed to the vertebrate genome range from strong skepticism to strong credence. Previous studies concentrated on small numbers of gene families or chromosomal regions that might not have been representative of the whole genome, or used subjective methods to identify paralogous genes and regions. Here we report a systematic and objective analysis of the draft human genome sequence to identify paralogous chromosomal regions (paralogons) formed during chordate evolution and to estimate the ages of duplicate genes. We found that the human genome contains many more paralogons than would be expected by chance. Molecular clock analysis of all protein families in humans that have orthologs in the fly and nematode indicated that a burst of gene duplication activity took place in the period 350 650 Myr ago and that many of the duplicate genes formed at this time are located within paralogons. Our results support the contention that many of the gene families in vertebrates were formed or expanded by large-scale DNA duplications in an early chordate. Considering the incompleteness of the sequence data and the antiquity of the event, the results are compatible with at least one round of polyploidy.

Animals↗

Genomic differences between Candida glabrata and Saccharomyces cerevisiae around the MRPL28 and GCN3 loci.

We report the sequences of two genomic regions from the pathogenic yeast Candida glabrata and their comparison to Saccharomyces cerevisiae. A 3 kb region from C. glabrata was sequenced that contains homologues of the S. cerevisiae genes TFB3, MRPL28 and STP1. The equivalent region in S. cerevisiae includes a fourth gene, MFA1, coding for mating factor a. The absence of MFA1 is consistent with C. glabrata's asexual life cycle, although we cannot exclude the possibility that a-factor gene(s) are located somewhere else in its genome. We also report the sequence of a 16 kb region from C. glabrata that contains a five-gene cluster similar to S. cerevisiae chromosome XI (including GCN3) followed by a four-gene cluster similar to chromosome XV (including HIS3). A small-scale rearrangement of gene order has occurred in the chromosome XI-like section.

Candida↗

Nucleotide substitution rates in legume chloroplast DNA depend on the presence of the inverted repeat.

The chloroplast genomes of some species of legumes lack the large inverted repeat (IR) that is a trademark of most land-plant chloroplasts. Our analysis of chloroplast genes in legume species that have an IR shows that the synonymous (silent) substitution rate in IR genes is 2.3-fold lower than in single-copy (SC) genes, which is largely in agreement with earlier findings. Given that all genes in species that lack the IR are single-copy, what level of synonymous substitution exists in these genes? We report a uniform substitution rate in IR-less genomes, and moreover, we find this rate to be at the level otherwise reserved for SC genes. In other words, the synonymous substitution rate has accelerated in the remaining copy of the duplicate region. We propose that this acceleration is a direct result of the decrease in the copy number of the sequence, rather than an intrinsic property of the genes normally located in the IR.

DNA, Chloroplast↗

Fourfold faster rate of genome rearrangement in nematodes than in Drosophila.

We compared the genome of the nematode Caenorhabditis elegans to 13% of that of Caenorhabditis briggsae, identifying 252 conserved segments along their chromosomes. We detected 517 chromosomal rearrangements, with the ratio of translocations to inversions to transpositions being approximately 1:1:2. We estimate that the species diverged 50-120 million years ago, and that since then there have been 4030 rearrangements between their whole genomes. Our estimate of the rearrangement rate, 0.4-1.0 chromosomal breakages/Mb per Myr, is at least four times that of Drosophila, which was previously reported to be the fastest rate among eukaryotes. The breakpoints of translocations are strongly associated with dispersed repeats and gene family members in the C. elegans genome.

Animals↗