Search PubMed⌕ Search

Biomedical subjects

P M Sharp

Publications and source records attributed to P M Sharp.

At least 109 records · Page 6Linked to original sources

Identification of functional open reading frames in chloroplast genomes.

We have used a rapid computer dot-matrix comparison method to identify all DNA regions which have been evolutionarily conserved between the completely sequenced chloroplast genomes of tobacco and a liverwort. Analysis of these regions reveals 74 homologous open reading frames (ORFs) which have been conserved as to length and amino acid sequence; these ORFs also have an excess of nucleotide substitutions at silent sites of codons. Since the nonfunctional parts of these genomes have become saturated with mutations and show no sequence similarity whatsoever, the homologous ORFs are almost certainly functional. A further four pairs of ORFs show homology limited to only a short part of their putative gene products. Amino acid sequence identities range between 50 and 99%; some chloroplast proteins are seen to be among the most slowly evolving of all known proteins. A search of the nucleotide and amino acid sequence databanks has revealed several previously unidentified genes in chloroplast sequences from other species, but no new homologies to prokaryotic genes.

Biological Evolution↗

Selective differences among translation termination codons.

The frequency of use of the three alternative translation termination codons has been examined in 165 Escherichia coli, 52 Bacillus subtilis and 106 Saccharomyces cerevisiae genes. Genes were first categorised according to their degree of bias in sense codon usage. In each species there is a very strong bias in favour of UAA (over UAG and UGA) in genes where sense codon usage is highly biased. This bias declines, principally with an increase in the use of UGA, in genes with lower sense codon bias. It appears that selection operating during translation may maintain the bias in stop codon usage. Such selection could result from the greater availability of UAA-cognate release factor(s), or from a lower frequency of translational readthrough at UAA.

Bacillus subtilis↗

Rates and dates of divergence between AIDS virus nucleotide sequences.

The acquired immune deficiency syndrome (AIDS), caused by a retrovirus called human immunodeficiency virus (HIV), has become a pandemic. A knowledge of the rate of nucleotide substitution in HIV and of the history and pattern of spread of the virus is important for understanding the epidemiology and pathogenesis of AIDS and for developing therapies and vaccine strategies. A new model has been developed and used to estimate the substitution rates in various regions in the HIV genome. The rate of nonsynonymous (amino acid-changing) substitution is lowest in the regions coding for the capsid proteins and the reverse transcriptase, being approximately 1.7 X 10(-3) nucleotide substitutions/site/year. The nonsynonymous rate is extremely high (14 X 10(-3] in the hypervariable regions of the envelope gene, suggesting extremely rapid change in viral antigenicity. The nonsynonymous rates in the other coding regions are between 3 X 10(-3) and 7 X 10(-3). The average synonymous rate for the HIV genome is 10 X 10(-3). These rates are 10(6) times greater than the rates in DNA genomes and at least as high as the rates in other RNA viruses. Evidence is provided for a case of recombination between different HIV strains. Our analysis suggests that the AIDS virus had existed in central Africa before 1960 and spread to North America before the mid 1970s. The evolutionary relationships among HIV isolates are inferred from nucleotide sequence data, and the result is consistent with the view that AIDS spread from Haiti to the United States.

Base Sequence↗

"Silent" sites in Drosophila genes are not neutral: evidence of selection among synonymous codons.

The patterns of synonymous codon usage in 91 Drosophila melanogaster genes have been examined. Codon usage varies strikingly among genes. This variation is associated with differences in G+C content at silent sites, but (unlike the situation in mammalian genes) these differences are not correlated with variation in intron base composition and so are not easily explicable in terms of mutational biases. Instead, those genes with high G+C content at silent sites, resulting from a strong "preference" for a particular subset of the codons that are mostly C-ending, appear to be the more highly expressed genes. This suggests that G+C content is reduced in sequences where selective constraints are weaker, as indeed seen in a pseudogene. These and other data discussed are consistent with the effects of translational selection among synonymous codons, as seen in unicellular organisms. The existence of selective constraints on silent substitutions, which may vary in strength among genes, has implications for the use of silent molecular clocks.

Animals↗

Synonymous codon usage in Bacillus subtilis reflects both translational selection and mutational biases.

Codon usage data for 56 Bacillus subtilis genes show that synonymous codon usage in B. subtilis is less biased than in Escherichia coli, or in Saccharomyces cerevisiae. Nevertheless, certain genes with a high codon bias can be identified by correspondence analysis, and also by various indices of codon bias. These genes are very highly expressed, and a general trend (a decrease) in codon bias across genes seems to correspond to decreasing expression level. This, then, may be a general phenomenon in unicellular organisms. The unusually small effect of translational selection on the pattern of codon usage in lowly expressed genes in B. subtilis yields similar dinucleotide frequencies among different codon positions, and on complementary strands. These patterns could arise through selection on DNA structure, but more probably are largely determined by mutation. This prevalence of mutational bias could lead to difficulties in assessing whether open reading frames encode proteins.

Bacillus subtilis↗

The codon Adaptation Index--a measure of directional synonymous codon usage bias, and its potential applications.

A simple, effective measure of synonymous codon usage bias, the Codon Adaptation Index, is detailed. The index uses a reference set of highly expressed genes from a species to assess the relative merits of each codon, and a score for a gene is calculated from the frequency of use of all codons in that gene. The index assesses the extent to which selection has been effective in moulding the pattern of codon usage. In that respect it is useful for predicting the level of expression of a gene, for assessing the adaptation of viral genes to their hosts, and for making comparisons of codon usage in different organisms. The index may also give an approximate indication of the likely success of heterologous gene expression.

Animals↗

Ubiquitin genes as a paradigm of concerted evolution of tandem repeats.

Ubiquitin is remarkable for its ubiquitous distribution and its extreme protein sequence conservation. Ubiquitin genes comprise direct repeats of the ubiquitin coding unit with no spacers. The nucleotide sequences of several ubiquitin repeats from each of humans, chicken, Xenopus, Drosophila, barley, and yeast have recently been determined. By analysis of these data we show that ubiquitin is evolving more slowly than any other known protein, and that this (together with its gene organization) contributes to an ideal situation for the occurrence of concerted evolution of tandem repeats. By contrast, there is little evidence of between-cluster concerted evolution. We deduce that in ubiquitin genes, concerted evolution involves both unequal crossover and gene conversion, and that the average time since two repeated units within the polyubiquitin locus most recently shared a common ancestor is approximately 38 million years (Myr) in mammals, but perhaps only 11 Myr in Drosophila. The extreme conservatism of ubiquitin evolution also allows the inference that certain synonymous serine codons differing at the first two positions were probably mutated at single steps.

Animals↗

An evaluation of the molecular clock hypothesis using mammalian DNA sequences.

A statistical analysis of extensive DNA sequence data from primates, rodents, and artiodactyls clearly indicates that no global molecular clock exists in mammals. Rates of nucleotide substitution in rodents are estimated to be four to eight times higher than those in higher primates and two to four times higher than those in artiodactyls. There is strong evidence for lower substitution rates in apes and humans than in monkeys, supporting the hominoid slowdown hypothesis. There is also evidence for lower rates in humans than in apes, suggesting a further rate slowdown in the human lineage after the separation of humans from apes. By contrast, substitution rates are nearly equal in mouse and rat. These results suggest that differences in generation time or, more precisely, in the number of germline DNA replications per year are the primary cause of rate differences in mammals. Further, these differences are more in line with the neutral mutation hypothesis than if the rates are the same for short- and long-living mammals.

Animals↗

Rates of nucleotide substitution vary greatly among plant mitochondrial, chloroplast, and nuclear DNAs.

Comparison of plant mitochondrial (mt), chloroplast (cp) and nuclear (n) DNA sequences shows that the silent substitution rate in mtDNA is less than one-third that in cpDNA, which in turn evolves only half as fast as plant nDNA. The slower rate in mtDNA than in cpDNA is probably due to a lower mutation rate. Silent substitution rates in plant and mammalian mtDNAs differ by one or two orders of magnitude, whereas the rates in nDNAs may be similar. In cpDNA, the rate of substitution both at synonymous sites and in noncoding sequences in the inverted repeat is greatly reduced in comparison to single-copy sequences. The rate of cpDNA evolution appears to have slowed in some dicot lineages following the monocot/dicot split, and the slowdown is more conspicuous at nonsynonymous sites than at synonymous sites.

Biological Evolution↗

The rate of synonymous substitution in enterobacterial genes is inversely related to codon usage bias.

Genes sequences from Escherichia coli, Salmonella typhimurium, and other members of the Enterobacteriaceae show a negative correlation between the degree of synonymous-codon usage bias and the rate of nucleotide substitution at synonymous sites. In particular, very highly expressed genes have very biased codon usage and accumulate synonymous substitutions very slowly. In contrast, there is little correlation between the degree of codon bias and the rate of protein evolution. It is concluded that both the rate of synonymous substitution and the degree of codon usage bias largely reflect the intensity of selection at the translational level. Because of the high variability among genes in rates of synonymous substitution, separate molecular clocks of synonymous substitution might be required for different genes.

Biological Evolution↗

Codon usage in regulatory genes in Escherichia coli does not reflect selection for 'rare' codons.

It has often been suggested that differential usage of codons recognized by rare tRNA species, i.e. "rare codons", represents an evolutionary strategy to modulate gene expression. In particular, regulatory genes are reported to have an extraordinarily high frequency of rare codons. From E. coli we have compiled codon usage data for highly expressed genes, moderately/lowly expressed genes, and regulatory genes. We have identified a clear and general trend in codon usage bias, from the very high bias seen in very highly expressed genes and attributed to selection, to a rather low bias in other genes which seems to be more influenced by mutation than by selection. There is no clear tendency for an increased frequency of rare codons in the regulatory genes, compared to a large group of other moderately/lowly expressed genes with low codon bias. From this, as well as a consideration of evolutionary rates of regulatory genes, and of experimental data on translation rates, we conclude that the pattern of synonymous codon usage in regulatory genes reflects primarily the relaxation of natural selection.

Base Sequence↗

Codon usage in yeast: cluster analysis clearly differentiates highly and lowly expressed genes.

Codon usage data has been compiled for 110 yeast genes. Cluster analysis on relative synonymous codon usage revealed two distinct groups of genes. One group corresponds to highly expressed genes, and has much more extreme synonymous codon preference. The pattern of codon usage observed is consistent with that expected if a need to match abundant tRNAs, and intermediacy of tRNA-mRNA interaction energies are important selective constraints. Thus codon usage in the highly expressed group shows a higher correlation with tRNA abundance, a greater degree of third base pyrimidine bias, and a lesser tendency to the A+T richness which is characteristic of the yeast genome. The cluster analysis can be used to predict the likely level of gene expression of any gene, and identifies the pattern of codon usage likely to yield optimal gene expression in yeast.

Base Composition↗

An evolutionary perspective on synonymous codon usage in unicellular organisms.

Observed patterns of synonymous codon usage are explained in terms of the joint effects of mutation, selection, and random drift. Examination of the codon usage in 165 Escherichia coli genes reveals a consistent trend of increasing bias with increasing gene expression level. Selection on codon usage appears to be unidirectional, so that the pattern seen in lowly expressed genes is best explained in terms of an absence of strong selection. A measure of directional synonymous-codon usage bias, the Codon Adaptation Index, has been developed. In enterobacteria, rates of synonymous substitution are seen to vary greatly among genes, and genes with a high codon bias evolve more slowly. A theoretical study shows that the patterns of extreme codon bias observed for some E. coli (and yeast) genes can be generated by rather small selective differences. The relative plausibilities of various theoretical models for explaining nonrandom codon usage are discussed.

Amino Acid Sequence↗

Molecular evolution of bacteriophages: evidence of selection against the recognition sites of host restriction enzymes.

Restriction enzymes produced by bacteria serve as a defense against invading bacteriophages, and so phages without other protection would be expected to undergo selection to eliminate recognition sites for these enzymes from their genomes. The observed frequencies of all restriction sites in the genomes of all completely sequenced DNA phages (T7, lambda, phi X174, G4, M13, f1, fd, and IKe) have been compared to expected frequencies derived from trinucleotide frequencies. Attention was focused on 6-base palindromes since they comprise the typical recognition sites for type II restriction enzymes. All of these coliphages, with the exception of lambda and G4, exhibit significant avoidance of the particular sequences that are enterobacterial restriction sites. As expected, the sequenced fraction of the genome of phi 29, a Bacillus subtilis phage, lacks Bacillus restriction sites. By contrast, the RNA phage MS2, several viruses that infect eukaryotes (EBV, adenovirus, papilloma, and SV40), and three mitochondrial genomes (human, mouse, and cow) were found not to lack restriction sites. Because the particular palindromes avoided correspond closely with the recognition sites for host enzymes and because other viruses and small genomes do not show this avoidance, it is concluded that the effect indeed results from natural selection.

Bacillus subtilis↗

Does the 'non-coding' strand code?

The hypothesis that DNA strands complementary to the coding strand contain in phase coding sequences has been investigated. Statistical analysis of the 50 genes of bacteriophage T7 shows no significant correlation between patterns of codon usage on the coding and non-coding strands. In Bacillus and yeast genes the correlation observed is not different from that expected with random synonymous codon usage, while a high correlation seen in 52 E. coli genes can be explained in terms of an excess of RNY codons. A deficiency of UUA, CUA and UCA codons (complementary to termination) seems to be restricted to the E. coli genes, and may be due to low abundance of the relevant cognate tRNA species. Thus the analysis shows that the non-coding strand has the properties expected of a sequence complementary to a coding strand, with no indications that it encodes, or may have encoded, proteins.

Bacillus↗