Search PubMed⌕ Search

Biomedical subjects

P M Sharp

Publications and source records attributed to P M Sharp.

At least 91 records · Page 5Linked to original sources

Codon usage in Aspergillus nidulans.

Synonymous codon usage in genes from the ascomycete (filamentous) fungus Aspergillus nidulans has been investigated. A total of 45 gene sequences has been analysed. Multivariate statistical analysis has been used to identify a single major trend among genes. At one end of this trend are lowly expressed genes, whereas at the other extreme lie genes known or expected to be highly expressed. The major trend is from nearly random codon usage (in the lowly expressed genes) to codon usage that is highly biased towards a set of 19-20 "optimal" codons. The G + C content of the A. nidulans genome is close to 50%, indicating little overall mutational bias, and so the codon usage of lowly expressed genes is as expected in the absence of selection pressure at silent sites. Most of the optimal codons are C- or G- ending, making highly expressed genes more G + C-rich at silent sites.

Aspergillus nidulans↗

Determinants of DNA sequence divergence between Escherichia coli and Salmonella typhimurium: codon usage, map position, and concerted evolution.

The nature and extent of DNA sequence divergence between homologous protein-coding genes from Escherichia coli and Salmonella typhimurium have been examined. The degree of divergence varies greatly among genes at both synonymous (silent) and nonsynonymous sites. Much of the variation in silent substitution rates can be explained by natural selection on synonymous codon usage, varying in intensity with gene expression level. Silent substitution rates also vary significantly with chromosomal location, with genes near oriC having lower divergence. Certain genes have been examined in more detail. In particular, the duplicate genes encoding elongation factor Tu, tufA and tufB, from S. typhimurium have been compared to their E. coli homologues. As expected these very highly expressed genes have high codon usage bias and have diverged very little between the two species. Interestingly, these genes, which are widely spaced on the bacterial chromosome, also appear to be undergoing concerted evolution, i.e., there has been exchange between the loci subsequent to the divergence of the two species.

Base Sequence↗

Mitochondrial DNA sequence divergence in the Melanogaster and oriental species subgroups of Drosophila.

The nucleotide sequence of a segment of the mitochondrial DNA from three Drosophila species (D. erecta, D. eugracilis, and D. takahashii), belonging to different subgroups of the melanogaster group has been determined. The segment encompasses three complete tRNA genes (tRNAtrp, tRNAcys, and tRNAtyr) and portions of two protein-coding genes: the subunit 2 of the NADH dehydrogenase (ND2) and the subunit 1 of the cytochrome oxidase (COI). Comparisons also involve homologous sequences already known for four other Drosophila species of the melanogaster group. Length differences were confined in the intergenic region where a long stretch of AT repeats was observed in one of the species analyzed. The three tRNA genes exhibit very different evolutionary rates, the most slowly evolving one, tRNAtyr, is adjacent to the 5' end of COI; tRNAs in similar positions have been previously shown to evolve slowly because they are probably involved in transcript processing. Although the rate of synonymous substitutions was very similar between ND2 and COI genes there were strong discrepancies between them in terms of the number of nonsynonymous substitutions. Differences have also been found in G + C content of the genes, which are likely to be linked to different selective pressures. There is a reduction in G + C content in the region where selective constraints are reduced. This suggests the existence of different levels of constraints along the sequenced segment. An overall analysis of the types of substitutions showed a decrease in A + T content during the course of evolution of the species.

Animals↗

ERIC sequences: a novel family of repetitive elements in the genomes of Escherichia coli, Salmonella typhimurium and other enterobacteria.

We describe a family of highly conserved, Enterobacterial Repetitive Intergenic Consensus (ERIC) sequences, 14 of which have been identified in Escherichia coli and Salmonella typhimurium and a further three in other enterobacterial species (Yersinia pseudotuberculosis, Klebsiella pneumoniae and Vibrio cholerae). ERIC sequences are 126 bp long and appear to be restricted to transcribed regions of the genome, either in intergenic regions of polycistronic operons or in untranslated regions upstream or downstream of open reading frames. ERIC sequences are highly conserved at the nucleotide sequence level but their chromosomal locations differ between species. Several features of ERIC sequences resemble those of REP sequences (Stern et al., 1984) although the nucleotide sequence is entirely different. The question of whether ERICs have a specific function, or represent a form of 'selfish' DNA, is discussed.

Base Sequence↗

Molecular phylogeny of Rodentia, Lagomorpha, Primates, Artiodactyla, and Carnivora and molecular clocks.

Phylogenetic analysis of DNA sequences from primates, rodents, lagomorphs, artiodactyls, carnivores, and birds strongly suggests that the order Rodentia is an outgroup to the other four mammalian orders and that Artiodactyla and Carnivora belong to a superordinal clade. Further, there is strong evidence against the Glires concept, which unites Lagomorpha and Rodentia. The radiation among Lagomorpha, Primates, and Artiodactyla--Carnivora is very bush-like, but there is some evidence that Lagomorpha has branched off first. Thus, the branching sequence for these five orders of mammals seems to be Rodentia, Lagomorpha, Primates, Artiodactyla, and Carnivora. The branching date for Rodentia could be as early as 100 million years ago. The rate of nucleotide substitution in the rodent lineage is shown to be at least 1.5 times higher than those in the other four mammalian lineages.

Animals↗

Processes of genome evolution reflected by base frequency differences among Serratia marcescens genes.

The G + C content of silent sites in codons varies greatly among Serratia marcescens genes; the value in any one gene seems to reflect a balance between mutation pressure towards high G + C content and natural selection constraining choice among synonymous codons. Interestingly, non-coding sequences have substantially lower G + C content than silent sites thought to be under little selective constraint.

Base Composition↗

Chromosomal location and evolutionary rate variation in enterobacterial genes.

The basal rate of DNA sequence evolution in enterobacteria, as seen in the extent of divergence between Escherichia coli and Salmonella typhimurium, varies greatly among genes, even when only "silent" sites are considered. The degree of divergence is clearly related to the level of gene expression, reflecting constraints on synonymous codon choice. However, where this constraint is weak, among genes not expressed at high levels, divergence is also related to the chromosomal location of the gene; it appears that genes furthest away from oriC, the origin of replication, have a mutation rate approximately two times that of genes near oriC.

Bias↗

Codon usage and gene expression level in Dictyostelium discoideum: highly expressed genes do 'prefer' optimal codons.

Codon usage patterns in the slime mould Dictyostelium discoideum have been re-examined (a total of 58 genes have been analysed). Considering the extreme A + T-richness of this genome (G + C = 22%), there is a surprising degree of codon usage variation among genes. For example, G + C content at silent sites varies from less than 10% to greater than 30%. It was previously suggested [Warrick, H.M. and Spudich, J.A. (1988) Nucleic Acids Res. 16: 6617-6635] that highly expressed genes contain fewer 'optimal' codons than genes expressed at lower levels. However, it appears that the optimal codons were misidentified. Multivariate statistical analysis shows that the greatest variation among genes is in relative usage of a particular subset of codons (about one per amino acid), many of which are C-ending. We have identified these as optimal codons, since (i) their frequency is positively correlated with gene expression level, and (ii) there is a strong mutation bias in this genome towards A and T nucleotides. Thus, codon usage in D. discoideum can be explained by a balance between the forces of mutational bias and translational selection.

Codon↗

Evidence that mutation patterns vary among Drosophila transposable elements.

In Drosophila melanogaster, codon usage in the open reading frames (ORFs) of transposable elements (TEs) differs greatly from that in other ORFs. In addition, while the ORFs from a single element are similar, there is considerable variation among elements. In the TE ORFs there are no indications of selection for the codons prevalent in the other D. melanogaster genes, but rather codon usage can be succinctly summarized in terms of the base composition at silent sites. We suggest that the particular silent site base composition of each TE is determined by an individual pattern of mutation. In many of the TEs there is an ORF encoding a protein with homology to reverse transcriptase; the amino acid sequences of these are quite divergent, and so it is possible that each of these incorporates certain mismatched bases at different frequencies during replication.

Animals↗

Mutation rates differ among regions of the mammalian genome.

In the traditional view of molecular evolution, the rate of point mutation is uniform over the genome of an organism and variation in the rate of nucleotide substitution among DNA regions reflects differential selective constraints. Here we provide evidence for significant variation in mutation rate among regions in the mammalian genome. We show first that substitutions at silent (degenerate) sites in protein-coding genes in mammals seem to be effectively neutral (or nearly so) as they do not occur significantly less frequently than substitutions in pseudogenes. We then show that the rate of silent substitution varies among genes and is correlated with the base composition of genes and their flanking DNA. This implies that the variation in both silent substitution rate and base composition can be attributed to systematic differences in the rate and pattern of mutation over regions of the genome. We propose that the differences arise because mutation patterns vary with the timing of replication of different chromosomal regions in the germline. This hypothesis can account for both the origin of isochores in mammalian genomes and the observation that silent nucleotide substitutions in different mammalian genes do not have the same molecular clock.

Animals↗

On the rate of DNA sequence evolution in Drosophila.

Analysis of the rate of nucleotide substitution at silent sites in Drosophila genes reveals three main points. First, the silent rate varies (by a factor of two) among nuclear genes; it is inversely related to the degree of codon usage bias, and so selection among synonymous codons appears to constrain the rate of silent substitution in some genes. Second, mitochondrial genes may have evolved only as fast as nuclear genes with weak codon usage bias (and two times faster than nuclear genes with high codon usage bias); this is quite different from the situation in mammals where mitochondrial genes evolve approximately 5-10 times faster than nuclear genes. Third, the absolute rate of substitution at silent sites in nuclear genes in Drosophila is about three times higher than the average silent rate in mammals.

Animals↗

Date of the monocot-dicot divergence estimated from chloroplast DNA sequence data.

The divergence between monocots and dicots represents a major event in higher plant evolution, yet the date of its occurrence remains unknown because of the scarcity of relevant fossils. We have estimated this date by reconstructing phylogenetic trees from chloroplast DNA sequences, using two independent approaches: the rate of synonymous nucleotide substitution was calibrated from the divergence of maize, wheat, and rice, whereas the rate of nonsynonymous substitution was calibrated from the divergence of angiosperms and bryophytes. Both methods lead to an estimate of the monocot-dicot divergence at 200 million years (Myr) ago (with an uncertainty of about 40 Myr). This estimate is also supported by analyses of the nuclear genes encoding large and small subunit ribosomal RNAs. These results imply that the angiosperm lineage emerged in Jurassic-Triassic time, which considerably predates its appearance in the fossil record (approximately 120 Myr ago). We estimate the divergence between cycads and angiosperms to be approximately 340 Myr, which can be taken as an upper bound for the age of angiosperms.

Animals↗

Fast and sensitive multiple sequence alignments on a microcomputer.

A strategy is described for the rapid alignment of many long nucleic acid or protein sequences on a microcomputer. The program described can handle up to 100 sequences of 1200 residues each. The approach is based on progressively aligning sequences according to the branching order in an initial phylogenetic tree. The results obtained using the package appear to be as sensitive as those from any other available method.

Algorithms↗

CLUSTAL: a package for performing multiple sequence alignment on a microcomputer.

An approach for performing multiple alignments of large numbers of amino acid or nucleotide sequences is described. The method is based on first deriving a phylogenetic tree from a matrix of all pairwise sequence similarity scores, obtained using a fast pairwise alignment algorithm. Then the multiple alignment is achieved from a series of pairwise alignments of clusters of sequences, following the order of branching in the tree. The method is sufficiently fast and economical with memory to be easily implemented on a microcomputer, and yet the results obtained are comparable to those from packages requiring mainframe computer facilities.

Algorithms↗

Codon usage patterns in Escherichia coli, Bacillus subtilis, Saccharomyces cerevisiae, Schizosaccharomyces pombe, Drosophila melanogaster and Homo sapiens; a review of the considerable within-species diversity.

The genetic code is degenerate, but alternative synonymous codons are generally not used with equal frequency. Since the pioneering work of Grantham's group it has been apparent that genes from one species often share similarities in codon frequency; under the "genome hypothesis" there is a species-specific pattern to codon usage. However, it has become clear that in most species there are also considerable differences among genes. Multivariate analyses have revealed that in each species so far examined there is a single major trend in codon usage among genes, usually from highly biased to more nearly even usage of synonymous codons. Thus, to represent the codon usage pattern of an organism it is not sufficient to sum over all genes as this conceals the underlying heterogeneity. Rather, it is necessary to describe the trend among genes seen in that species. We illustrate these trends for six species where codon usage has been examined in detail, by presenting the pooled codon usage for the 10% of genes at either end of the major trend. Closely-related organisms have similar patterns of codon usage, and so the six species in Table 1 are representative of wider groups. For example, with respect to codon usage, Salmonella typhimurium closely resembles E. coli, while all mammalian species so far examined (principally mouse, rat and cow) largely resemble humans.

Amino Acids↗