Search PubMed⌕ Search

Biomedical subjects

Walter Gilbert

Publications and source records attributed to Walter Gilbert.

14 recordsLinked to original sources

The evolution of spliceosomal introns: patterns, puzzles and progress.

The origins and importance of spliceosomal introns comprise one of the longest-abiding mysteries of molecular evolution. Considerable debate remains over several aspects of the evolution of spliceosomal introns, including the timing of intron origin and proliferation, the mechanisms by which introns are lost and gained, and the forces that have shaped intron evolution. Recent important progress has been made in each of these areas. Patterns of intron-position correspondence between widely diverged eukaryotic species have provided insights into the origins of the vast differences in intron number between eukaryotic species, and studies of specific cases of intron loss and gain have led to progress in understanding the underlying molecular mechanisms and the forces that control intron evolution.

Animals↗

Rates of intron loss and gain: implications for early eukaryotic evolution.

We study the intron-exon structures of 684 groups of orthologs from seven diverse eukaryotic genomes and provide maximum likelihood estimates for rates and numbers of intron losses and gains in these same genes for a variety of lineages. Rates of intron loss vary from approximately 2 x 10(-9) to 2 x 10(-10) per year. Rates of gain vary from 6 x 10(-13) to 4 x 10(-12) per possible intron insertion site per year. There is an inverse correspondence between rates of intron loss and gain, leading to a 20-fold variation among lineages in the ratio of the rates of the two processes. The observed rates of intron gain are insufficient to explain the large number of introns estimated to have been present in the plant-animal ancestor, suggesting that introns present in early eukaryotes may have been created by a fundamentally different process than more recently gained introns.

Animals↗

Resolution of a deep animal divergence by the pattern of intron conservation.

The relationship between three biologically important groups, arthropods, nematodes, and deuterostomes, remains unresolved. It is unknown whether arthropods are more closely related to nematodes (consistent with the "ecdysozoa" hypothesis) or to deuterostomes (consistent with "coelomata"). We present a method in which we use the pattern of spliceosomal intron conservation to develop a series of inequalities that characterize each possible relationship. We find that only the ecdysozoa grouping satisfies these predictions, with P < 10(-6). Simulations show that our method, unlike some previous methods, is largely insensitive to rate variation between branches.

Animals↗

Complex early genes.

We use the pattern of intron conservation in 684 groups of orthologs from seven fully sequenced eukaryotic genomes to provide maximum likelihood estimates of the number of introns present in the same orthologs in various eukaryotic ancestors. We find: (i) intron density in the plant-animal ancestor was high, perhaps two-thirds that of humans and three times that of Drosophila; and (ii) intron density in the ancestral bilateran was also high, equaling that of humans and four times that of Drosophila. We further find that modern introns are generally very old, with two-thirds of modern bilateran introns dating to the ancestral bilateran and two-fifths of modern plant, animal, and fungus introns dating to the plant-animal ancestor. Intron losses outnumber gains over a large range of eukaryotic lineages. These results show that early eukaryotic gene structures were very complex, and that simplification, not embellishment, has dominated subsequent evolution.

Animals↗

The pattern of intron loss.

We studied intron loss in 684 groups of orthologous genes from seven fully sequenced eukaryotic genomes. We found that introns closer to the 3' ends of genes are preferentially lost, as predicted if introns are lost through gene conversion with a reverse transcriptase product of a spliced mRNA. Adjacent introns tend to be lost in concert, as expected if such events span multiple intron positions. Directly contrary to the expectations of some, introns that do not interrupt codons (phase zero) are more, not less, likely to be lost, an intriguing and previously unappreciated result. Adjacent introns with matching phases are not more likely to be retained, as would be expected if they enjoyed a relative selective advantage. The findings of 3' and phase zero intron loss biases are in direct contradiction to an extremely recent study of fungi intron evolution. All patterns are less pronounced in the lineage leading to Caenorhabditis elegans, suggesting that the process of intron loss may be qualitatively different in nematodes. Our results support a reverse transcriptase-mediated model of intron loss.

3' Flanking Region↗

Mystery of intron gain.

For nearly 15 years, it has been widely believed that many introns were recently acquired by the genes of multicellular organisms. However, the mechanism of acquisition has yet to be described for a single animal intron. Here, we report a large-scale computational analysis of the human, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana genomes. We divided 147,796 human intron sequences into batches of similar lengths and aligned them with each other. Different types of homologies between introns were found, but none showed evidence of simple intron transposition. Also, 106,902 plant, 39,624 Drosophila, and 6021 C. elegans introns were examined. No single case of homologous introns in nonhomologous genes was detected. Thus, we found no example of transposition of introns in the last 50 million years in humans, in 3 million years in Drosophila and C. elegans, or in 5 million years in Arabidopsis. Either new introns do not arise via transposition of other introns or intron transposition must have occurred so early in evolution that all traces of homology have been lost.

Animals↗

Large-scale comparison of intron positions in mammalian genes shows intron loss but no gain.

We compared intron-exon structures in 1,560 human-mouse orthologs and 360 mouse-rat orthologs. The origin of differences in intron positions between species was inferred by comparison with an outgroup, Fugu for human-mouse and human for mouse-rat. Among 10,020 intron positions in the human-mouse comparison, we found unequivocal evidence for five independent intron losses in the mouse lineage but no evidence for intron loss in humans or for intron gain in either lineage. Among 1,459 positions in rat-mouse comparisons, we found evidence for one loss in rat but neither loss in mouse nor gain in either lineage. In each case, the intron losses were exact, without change in the surrounding coding sequence, and involved introns that are extremely short, with an average of 200 bp, an order of magnitude shorter than the mammalian average. These results favor a model whereby introns are lost through gene conversion with intronless copies of the gene. In addition, the finding of widespread conservation of intron-exon structure, even over large evolutionary distances, suggests that comparative methods employing information about gene structures should be very successful in correctly predicting exon boundaries in genomic sequences.

Amino Acid Sequence↗

Phylogenetically older introns strongly correlate with module boundaries in ancient proteins.

The hypothesis that some (but not all) introns were used to construct ancient genes by exon shuffling of modules at the earliest stages of evolution is supported by the finding of an excess of phase-zero intron positions in the boundary regions of such modules in 276 ancient proteins (defined as common to eukaryotes and prokaryotes). Here we show further that as phase-zero intron positions are shared by distant taxa, and thus are truly phylogenetically ancient, their excess in the boundaries becomes greater, rising to an 80% excess if shared by four out of the five taxa: vertebrates, invertebrates, fungi, plants, and protists.

Computational Biology↗

The universe of exons revisited.

We study the distribution of exons in eukaryotic genes to determine whether one can detect the reuse of exon sequences and to use the frequency of such reuse to estimate how many ancestral exon sequences there might have been. We use two databases of exons. One contained 56,276 internal exons from putatively unrelated genes (less than 20% sequence identity) and the second contained 8917 internal exons from regions of these genes that are homologous and colinear with prokaryotic genes; these are ancient conserved regions (ACRs). At the 95% significance level we find 3500 exon-sequence matches in the large database and 500 matches in the ACR database. These matches correspond to groups of similar sequences. The size-rank relationship for these groups follows a power law, the size falling off as the inverse square root of the rank. This form of the power law distribution leads us to make an estimate for the size of a possible universe of ancestral exons. Using the data corresponding to the ACR regions, that universe is estimated to be about 15,000-30,000 in size.

Amino Acid Sequence↗

Large-scale comparison of intron positions among animal, plant, and fungal genes.

We purge large databases of animal, plant, and fungal intron-containing genes to a 20% similarity level and then identify the most similar animal-plant, animal-fungal, and plant-fungal protein pairs. We identify the introns in each BLAST 2.0 alignment and score matched intron positions and slid (near-matched, within six nucleotides) intron positions automatically. Overall we find that 10% of the animal introns match plant positions, and a further 7% are "slides." Fifteen percent of fungal introns match animal positions, and 13% match plant positions. Furthermore, the number of alignments with high numbers of matches deviates greatly from the Poisson expectation. The 30 animal-plant alignments with the highest matches (for which 44% of animal introns match plant positions) when aligned with fungal genes are also highly enriched for triple matches: 39% of the fungal introns match both animal and plant positions. This is strong evidence for ancestral introns predating the animal-plant-fungal divergence, and in complete opposition to any expectations based on random insertion. In examining the slid introns, we show that at least half are caused by imperfections in the alignments, and are most likely to be actual matches at common positions. Thus, our final estimates are that approximately equal 14% of animal introns match plant positions, and that approximately equal 17-18% of fungal introns match animal or plant positions, all of these being likely to be ancestral in the eukaryotes.

Amino Acid Sequence↗

The signal of ancient introns is obscured by intron density and homolog number.

In ancient genes whose products have known 3-dimensional structures, an excess of phase zero introns (those that lie between the codons) appear in the boundaries of modules, compact regions of the polypeptide chain. These excesses are highly significant and could support the hypothesis that ancient genes were assembled by exon shuffling involving compact modules. (Phase one and two introns, and many phase zero introns, appear to arise later.) However, as more genes, with larger numbers of homologs and intron positions, were examined, the effects became smaller, dropping from a 40% excess to an 8% excess as the number of intron positions increased from 570 to 3,328, even though the statistical significance remained strong. An interpretation of this behavior is that novel inserted positions appearing in homologs washed out the signal from a finite number of ancient positions. Here we show that this is likely to be the case. Analyses of intron positions restricted to those in genes for which relatively few intron positions from homologs are known, or to those in genes with a small number of known homologous gene structures, show a significant correlation of phase zero intron positions with the module structure, which weakens as the density of attributed intron positions or the number of homologs increases. These effects do not appear for phase one and phase two introns. This finding matches the expectation of the mixed model of intron origin, in which a fraction of phase zero introns are left from the assembly of the first genes, while other introns have been added in the course of evolution.

Evolution, Molecular↗

Regularities of context-dependent codon bias in eukaryotic genes.

Nucleotides surrounding a codon influence the choice of this particular codon from among the group of possible synonymous codons. The strongest influence on codon usage arises from the nucleotide immediately following the codon and is known as the N1 context. We studied the relative abundance of codons with N1 contexts in genes from four eukaryotes for which the entire genomes have been sequenced: Homo sapiens, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana. For all the studied organisms it was found that 90% of the codons have a statistically significant N1 context-dependent codon bias. The relative abundance of each codon with an N1 context was compared with the relative abundance of the same 4mer oligonucleotide in the whole genome. This comparison showed that in about half of all cases the context-dependent codon bias could not be explained by the sequence composition of the genome. Ranking statistics were applied to compare context-dependent codon biases for codons from different synonymous groups. We found regularities in N1 context-dependent codon bias with respect to the codon nucleotide composition. Codons with the same nucleotides in the second and third positions and the same N1 context have a statistically significant correlation of their relative abundances.

Animals↗

Do introns favor or avoid regions of amino acid conservation?

Are intron positions correlated with regions of high amino acid conservation? For a set of ancient conserved proteins, with intronless prokaryotic but intron-containing eukaryotic homologs, multiple sequence alignments identified residues invariant throughout evolution. Intron positions between codons show no preferences. However, introns lying after the first base of a codon prefer conserved regions, markedly in glycines. Because glycines are in excess in conserved regions, this behavior could reflect phase-one introns entering glycine residues randomly in the ancestral sequences. Examination of intron positions within codons of evolutionarily invariable amino acids showed that roughly 50% of these introns are bordered by guanines at both 5'- and 3'-ends, 25% have a G only before the intron, and 5% have a G only after the intron, whereas about 20% are bordered by nonguanine bases.

Alternative Splicing↗