Search PubMed⌕ Search

Biomedical subjects

Kenneth H Wolfe

Publications and source records attributed to Kenneth H Wolfe.

At least 19 recordsLinked to original sources

Birth of a metabolic gene cluster in yeast by adaptive gene relocation.

Although most eukaryotic genomes lack operons, they contain some physical clusters of genes that are related in function despite being unrelated in sequence. How these clusters are formed during evolution is unknown. The DAL cluster is the largest metabolic gene cluster in yeast and consists of six adjacent genes encoding proteins that enable Saccharomyces cerevisiae to use allantoin as a nitrogen source. We show here that the DAL cluster was assembled, quite recently in evolutionary terms, through a set of genomic rearrangements that happened almost simultaneously. Six of the eight genes involved in allantoin degradation, which were previously scattered around the genome, became relocated to a single subtelomeric site in an ancestor of S. cerevisiae and Saccharomyces castellii. These genomic rearrangements coincided with a biochemical reorganization of the purine degradation pathway, which switched to importing allantoin instead of urate. This change eliminated urate oxidase, one of several oxygen-consuming enzymes that were lost by yeasts that can grow vigorously in anaerobic conditions. The DAL cluster is located in a domain of modified chromatin involving both H2A.Z histone exchange and Hst1-Sum1-mediated histone deacetylation, and it may be a coadapted gene complex formed by epistatic selection.

Allantoin↗

A genome sequence survey shows that the pathogenic yeast Candida parapsilosis has a defective MTLa1 allele at its mating type locus.

Candida parapsilosis is responsible for ca. 15% of Candida infections and is of particular concern in neonates and surgical intensive care patients. The related species Candida albicans has recently been shown to possess a functional mating pathway. To analyze the analogous pathway in C. parapsilosis, we carried out a genome sequence survey of the type strain. We identified ca. 3,900 genes, with an average amino acid identity of 59% with C. albicans. Of these, 23 are predicted to be predominantly involved in mating. We identified a genomic locus homologous to the MTLa mating type locus of C. albicans, but the C. parapsilosis type strain has at least two internal stop codons in the MTLa1 open reading frame, and two predicted introns are not spliced. These stop codons were present in MTLa1 of all eight C. parapsilosis isolates tested. Furthermore, we found that all isolates of C. parapsilosis tested appear to contain only the MTLa idiomorph at the presumptive mating locus, unlike C. albicans and C. dubliniensis. MTLalpha sequences are present but at a different chromosomal location. It is therefore likely that all (or at least the majority) of C. parapsilosis isolates have a mating pathway that is either defective or substantially different from that of C. albicans.

Alleles↗

Clusters of co-expressed genes in mammalian genomes are conserved by natural selection.

Genes that belong to the same functional pathways are often packaged into operons in prokaryotes. However, aside from examples in nematode genomes, this form of transcriptional regulation appears to be absent in eukaryotes. Nevertheless, a number of recent studies have shown that gene order in eukaryotic genomes is not completely random, and that genes with similar expression patterns tend to be clustered together. What remains unclear is whether co-expressed genes have been gathered together by natural selection to facilitate their regulation, or if the genes are co-expressed simply by virtue of their being close together in the genome. Here, we show that gene expression clusters tend to contain fewer chromosomal breakpoints between human and mouse than expected by chance, which indicates that they are being held together by natural selection. This conclusion applies to clusters defined on the basis of broad (housekeeping) expression, or on the basis of correlated transcription profiles across tissues. Contrary to previous reports, we find that genes with high expression are not clustered to a greater extent than expected by chance and are not conserved during evolution.

Animals↗

Allele-specific transcript isoforms in human.

Estimates of the number of human genes that produce more than one transcript isoform through alternative mRNA splicing depend on the assumption that the observation of multiple transcripts from a gene can be attributed entirely to alternative splicing. It is possible, however, that a substantial proportion of cases where multiple transcripts have been observed for a gene result from differences between alleles. Many examples of genes that are spliced differently from different alleles have been reported but no systematic estimate of the proportion of alternatively spliced genes that are affected by such polymorphisms has been carried out. We find that alternative transcript isoforms are non-randomly associated with closely linked nucleotide polymorphisms, based on an integrated analysis of the dbSNP, dbEST and ASAP databases. From the observed level of association between transcript isoforms and polymorphisms, we estimate that 21% of alternatively spliced genes are affected by polymorphisms that either completely determine which form of the transcript is observed or alter the relative abundances of some of the alternative isoforms. We provide a conservative lower bound of 6% on this estimate and point out that alternative splicing cannot be confirmed absolutely unless more than one transcript is observed from the same allele.

Alleles↗

Origins of recently gained introns in Caenorhabditis.

The genomes of the nematodes Caenorhabditis elegans and Caenorhabditis briggsae both contain approximately 100,000 introns, of which >6,000 are unique to one or the other species. To study the origins of new introns, we used a conservative method involving phylogenetic comparisons to animal orthologs and nematode paralogs to identify cases where an intron content difference between C. elegans and C. briggsae was caused by intron insertion rather than deletion. We identified 81 recently gained introns in C. elegans and 41 in C. briggsae. Novel introns have a stronger exon splice site consensus sequence than the general population of introns and show the same preference for phase 0 sites in codons over phases 1 and 2. More of the novel introns are inserted in genes that are expressed in the C. elegans germ line than expected by chance. Thirteen of the 122 gained introns are in genes whose protein products function in premRNA processing, including three gains in the gene for spliceosomal protein SF3B1 and two in the nonsense-mediated decay gene smg-2. Twenty-eight novel introns have significant DNA sequence identity to other introns, including three that are similar to other introns in the same gene. All of these similarities involve minisatellites or palindromes in the intron sequences. Our results suggest that at least some of the intron gains were caused by reverse splicing of a preexisting intron.

Amino Acid Sequence↗

PubCrawler: keeping up comfortably with PubMed and GenBank.

The free PubCrawler web service (http://www.pubcrawler.ie) has been operating for five years and so far has brought literature and sequence updates to over 22 000 users. It provides information on a personalized web page whenever new articles appear in PubMed or when new sequences are found in GenBank that are specific to customized queries. The server also acts as an automatic alerting system by sending out short notifications or emails with the latest updates as soon as they become available. A new output format and more flexibility for the email formatting help PubCrawler cope with increasing challenges arising from browser incompatibilities and mail filters, therefore making it suitable for a wide range of users.

Databases, Nucleic Acid↗

Widespread paleopolyploidy in model plant species inferred from age distributions of duplicate genes.

It is often anticipated that many of today's diploid plant species are in fact paleopolyploids. Given that an ancient large-scale duplication will result in an excess of relatively old duplicated genes with similar ages, we analyzed the timing of duplication of pairs of paralogous genes in 14 model plant species. Using EST contigs (unigenes), we identified pairs of paralogous genes in each species and used the level of synonymous nucleotide substitution to estimate the relative ages of gene duplication. For nine of the investigated species (wheat [Triticum aestivum], maize [Zea mays], tetraploid cotton [Gossypium hirsutum], diploid cotton [G. arboretum], tomato [Lycopersicon esculentum], potato [Solanum tuberosum], soybean [Glycine max], barrel medic [Medicago truncatula], and Arabidopsis thaliana), the age distributions of duplicated genes contain peaks corresponding to short evolutionary periods during which large numbers of duplicated genes were accumulated. Large-scale duplications (polyploidy or aneuploidy) are strongly suspected to be the cause of these temporal peaks of gene duplication. However, the unusual age profile of tandem gene duplications in Arabidopsis indicates that other scenarios, such as variation in the rate at which duplicated genes are deleted, must also be considered.

Computational Biology↗

Functional divergence of duplicated genes formed by polyploidy during Arabidopsis evolution.

To study the evolutionary effects of polyploidy on plant gene functions, we analyzed functional genomics data for a large number of duplicated gene pairs formed by ancient polyploidy events in Arabidopsis thaliana. Genes retained in duplicate are not distributed evenly among Gene Ontology or Munich Information Center for Protein Sequences functional categories, which indicates a nonrandom process of gene loss. Genes involved in signal transduction and transcription have been preferentially retained, and those involved in DNA repair have been preferentially lost. Although the two members of each gene pair must originally have had identical transcription profiles, less than half of the pairs formed by the most recent polyploidy event still retain significantly correlated profiles. We identified several cases where groups of duplicated gene pairs have diverged in concert, forming two parallel networks, each containing one member of each gene pair. In these cases, the expression of each gene is strongly correlated with the other nonhomologous genes in its network but poorly correlated with its paralog in the other network. We also find that the rate of protein sequence evolution has been significantly asymmetric in >20% of duplicate pairs. Together, these results suggest that functional diversification of the surviving duplicated genes is a major feature of the long-term evolution of polyploids.

Arabidopsis↗

Evolution of the MAT locus and its Ho endonuclease in yeast species.

The genetics of the mating-type (MAT) locus have been studied extensively in Saccharomyces cerevisiae, but relatively little is known about how this complex system evolved. We compared the organization of MAT and mating-type-like (MTL) loci in nine species spanning the hemiascomycete phylogenetic tree. We inferred that the system evolved in a two-step process in which silent HMR/HML cassettes appeared, followed by acquisition of the Ho endonuclease from a mobile genetic element. Ho-mediated switching between an active MAT locus and silent cassettes exists only in the Saccharomyces sensu stricto group and their closest relatives: Candida glabrata, Kluyveromyces delphensis, and Saccharomyces castellii. We identified C. glabrata MTL1 as the ortholog of the MAT locus of K. delphensis and show that switching between C. glabrata MTL1a and MTL1alpha genotypes occurs in vivo. The more distantly related species Kluyveromyces lactis has silent cassettes but switches mating type without the aid of Ho endonuclease. Very distantly related species such as Candida albicans and Yarrowia lipolytica do not have silent cassettes. In Pichia angusta, a homothallic species, we found MATalpha2, MATalpha1, and MATa1 genes adjacent to each other on the same chromosome. Although some continuity in the chromosomal location of the MAT locus can be traced throughout hemiascomycete evolution and even to Neurospora, the gene content of the locus has changed with the loss of an HMG domain gene (MATa2) from the MATa idiomorph shortly after HO was recruited.

Deoxyribonucleases, Type II Site-Specific↗

Divergence of spatial gene expression profiles following species-specific gene duplications in human and mouse.

To examine the process by which duplicated genes diverge in function, we studied how the gene expression profiles of orthologous gene sets in human and mouse are affected by the presence of additional recent species-specific paralogs. Gene expression profiles were compared across 16 homologous tissues in human and mouse using microarray data from the Gene Expression Atlas for 1575 sets of orthologs including 250 with species-specific paralogs. We find that orthologs that have undergone recent duplication are less likely to have strongly correlated expression profiles than those that remain in a one-to-one relationship between human and mouse. There is a general trend for paralogous genes to become more specialized in their expression patterns, with decreased breadth and increased specificity of expression as gene family size increases. Despite this trend, detailed examination of some particular gene families where species-specific duplications have occurred indicated several examples of apparent neofunctionalization of duplicated genes, but only one case of subfunctionalization. Often, the expression of both copies of a duplicated gene appears to have changed relative to the ancestral state. Our results suggest that gene expression profiles are surprisingly labile and that expression in a particular tissue may be gained or lost repeatedly during the evolution of even small gene families. We conclude that gene duplication is a major driving force behind the emergence of divergent gene expression patterns.

Animals↗

Congruence of tissue expression profiles from Gene Expression Atlas, SAGEmap and TissueInfo databases.

BACKGROUND: Extracting biological knowledge from large amounts of gene expression information deposited in public databases is a major challenge of the postgenomic era. Additional insights may be derived by data integration and cross-platform comparisons of expression profiles. However, database meta-analysis is complicated by differences in experimental technologies, data post-processing, database formats, and inconsistent gene and sample annotation. RESULTS: We have analysed expression profiles from three public databases: Gene Expression Atlas, SAGEmap and TissueInfo. These are repositories of oligonucleotide microarray, Serial Analysis of Gene Expression and Expressed Sequence Tag human gene expression data respectively. We devised a method, Preferential Expression Measure, to identify genes that are significantly over- or under-expressed in any given tissue. We examined intra- and inter-database consistency of Preferential Expression Measures. There was good correlation between replicate experiments of oligonucleotide microarray data, but there was less coherence in expression profiles as measured by Serial Analysis of Gene Expression and Expressed Sequence Tag counts. We investigated inter-database correlations for six tissue categories, for which data were present in the three databases. Significant positive correlations were found for brain, prostate and vascular endothelium but not for ovary, kidney, and pancreas. CONCLUSION: We show that data from Gene Expression Atlas, SAGEmap and TissueInfo can be integrated using the UniGene gene index, and that expression profiles correlate relatively well when large numbers of tags are available or when tissue cellular composition is simple. Finally, in the case of brain, we demonstrate that when PEM values show good correlation, predictions of tissue-specific expression based on integrated data are very accurate.

Brain↗

Positive selection and subfunctionalization of duplicated CCT chaperonin subunits.

To reach a functional and energetically stable conformation, many proteins need molecular helpers called chaperonins. Among the group II chaperonins, CCT proteins provide crucial machinery for the stabilization and proper folding of several proteins in the cytosol of eukaryotic cells through interactions that are subunit-specific and geometry-dependent. CCT proteins are made up of eight different subunits, all with similar sequences, positioned in a precise arrangement. Each subunit has been proposed to have a specialized function during the binding and folding of the CCT protein substrate. Here, we demonstrate that functional divergence occurred after several CCT duplication events due to the fixation of amino acid substitutions by positive selection. Sites critical for ATP binding and substrate binding were found to have undergone positive selection and functional divergence predominantly in subunits that bind tubulin but not actin. Furthermore, we show clear functional divergence between CCT subunits that bind the C-terminal domains of actin and tubulin and those that bind the N-terminal domains. Phylogenetic analyses could not resolve the deep relationships between most subunits, except for the groups alpha/beta/eta and delta/epsilon, suggesting several almost simultaneous ancient duplication events. Together, the results support the idea that, in contrast to homo-oligomeric chaperonins such as GroEL, the high divergence level between CCT subunits is the result of positive selection after each duplication event to provide a specialized role for each CCT subunit in the different steps of protein folding.

Amino Acid Substitution↗

Wrapping up BLAST and other applications for use on Unix clusters.

UNLABELLED: We have developed two programs that speed up common bioinformatic applications by spreading them across a UNIX cluster.(1) BLAST.pm, a new module for the 'MOLLUSC' package. (2) WRAPID, a simple tool for parallelizing large numbers of small instances of programs such as BLAST, FASTA and CLUSTALW. AVAILABILITY: The packages were developed in Perl on a 20-node Linux cluster and are provided together with a configuration script and documentation. They can be freely downloaded from http://wolfe.gen.tcd.ie/wrapper.

Computer Communication Networks↗

Evidence from comparative genomics for a complete sexual cycle in the 'asexual' pathogenic yeast Candida glabrata.

BACKGROUND: Candida glabrata is a pathogenic yeast of increasing medical concern. It has been regarded as asexual since it was first described in 1917, yet phylogenetic analyses have revealed that it is more closely related to sexual yeasts than other Candida species. We show here that the C. glabrata genome contains many genes apparently involved in sexual reproduction. RESULTS: By genome survey sequencing, we find that genes involved in mating and meiosis are as numerous in C. glabrata as in the sexual species Kluyveromyces delphensis, which is its closest known relative. C. glabrata has a putative mating-type (MAT) locus and a pheromone gene (MFALPHA2), as well as orthologs of at least 31 other Saccharomyces cerevisiae genes that have no known roles apart from mating or meiosis, including FUS3, IME1 and SMK1. CONCLUSIONS: We infer that C. glabrata is likely to have an undiscovered sexual stage in its life cycle, similar to that recently proposed for C. albicans. The two Candida species represent two distantly related yeast lineages that have independently become both pathogenic and 'asexual'. Parallel evolution in the two lineages as they adopted mammalian hosts resulted in separate but analogous switches from overtly sexual to cryptically sexual life cycles, possibly in response to defense by the host immune system.

Candida glabrata↗

Molecular evolution meets the genomics revolution.

Changes in technology in the past decade have had such an impact on the way that molecular evolution research is done that it is difficult now to imagine working in a world without genomics or the Internet. In 1992, GenBank was less than a hundredth of its current size and was updated every three months on a huge spool of tape. Homology searches took 30 minutes and rarely found a hit. Now it is difficult to find sequences with only a few homologs to use as examples for teaching bioinformatics. For molecular evolution researchers, the genomics revolution has showered us with raw data and the information revolution has given us the wherewithal to analyze it. In broad terms, the most significant outcome from these changes has been our newfound ability to examine the evolution of genomes as a whole, enabling us to infer genome-wide evolutionary patterns and to identify subsets of genes whose evolution has been in some way atypical.

Animals↗

A recent polyploidy superimposed on older large-scale duplications in the Arabidopsis genome.

The Arabidopsis genome contains numerous large duplicated chromosomal segments, but the different approaches used in previous analyses led to different interpretations regarding the number and timing of ancestral large-scale duplication events. Here, using more appropriate methodology and a more recent version of the genome sequence annotation, we investigate the scale and timing of segmental duplications in Arabidopsis. We used protein sequence similarity searches to detect duplicated blocks in the genome, used the level of synonymous substitution between duplicated genes to estimate the relative ages of the blocks containing them, and analyzed the degree of overlap between adjacent duplicated blocks. We conclude that the Arabidopsis lineage underwent at least two distinct episodes of duplication. One was a polyploidy that occurred much more recently than estimated previously, before the Arabidopsis/Brassica rapa split and probably during the early emergence of the crucifer family (24-40 Mya). An older set of duplicated blocks was formed after the monocot/dicot divergence, and the relatively low level of overlap among these blocks indicates that at least some of them are remnants of a larger duplication such as a polyploidy or aneuploidy.

Arabidopsis↗

The 2R hypothesis and the human genome sequence.

One theory formalised in 1970 proposes that the complexity of vertebrate genomes originated by means of genome duplication at the base of the vertebrate lineage. Since then, the theory has remained both popular and controversial. Here we review the theory, and present preliminary results from our analysis of duplications in the draft human genome sequence. We find evidence for extensive duplication of parts of the genome. We also question the validity of the 'parsimony test' that has been used in other analyses.

Animals↗

Evolutionary re-organisation of a large operon in adzuki bean chloroplast DNA caused by inverted repeat movement.

We have sequenced two sections of chloroplast DNA from adzuki bean (Vigna angularis), containing the junctions between the inverted repeat (IR) and large single copy (LSC) regions of the genome. The gene order at both junctions is different from that described for other members of the legume family, such as Lotus japonicus and soybean. These differences have been attributed to an apparent 78-kb inversion that spans nearly the entire LSC region and which is present in adzuki and its close relative, the common bean. This 78-kb rearrangement broke the large S10 operon of ribosomal proteins into two smaller operons, one at each end of the LSC, without affecting the gene content of the genome. It disrupted the physical and transcriptional relationship between the six-gene rpl23-rpl14 cluster and the four-gene rps8-rpoA cluster that is conserved in most land plants. Analysis of the endpoints of the rearrangement indicates that it probably occurred by means of a two-step process of expansion and contraction of the IR and not by a 78-kb inversion.

Base Sequence↗