Search PubMedSearch

Biomedical subjects

P M Sharp

Publications and source records attributed to P M Sharp.

At least 19 recordsLinked to original sources

Evolution of codon usage patterns: the extent and nature of divergence between Candida albicans and Saccharomyces cerevisiae.

Codon usage in a sample of 28 genes from the pathogenic yeast Candida albicans has been analysed using multivariate statistical analysis. A major trend among genes, correlated with gene expression level, was identified. We have focussed on the extent and nature of divergence between C.albicans and the closely related yeast Saccharomyces cerevisiae. It was recently suggested that significant differences exist between the subsets of preferred codons in these two species [Brown et al. (1991) Nucleic Acids Res. 19, 4293]. Overall, the genes of C.albicans are more A + T-rich, reflecting the lower genomic G + C content of that species, and presumably resulting from a different pattern of mutational bias. However, in both species highly expressed genes preferentially use the same subset of 'optimal' codons. A suggestion that the low frequency of NCG codons in both yeast species results from selection against the presence of codons that are potentially highly mutable is discounted. Codon usage in C.albicans, as in other unicellular species, can be interpreted as the result of a balance between the processes of mutational bias and translational selection. Codon usage in two related Candida species, C.maltosa and C.tropicalis, is briefly discussed.

Biological Evolution

Roles of selection and recombination in the evolution of type I restriction-modification systems in enterobacteria.

Restriction-modification systems can protect bacteria against viral infection. Sequences of the hsdM gene, encoding one of the three subunits of type I restriction-modification systems, have been determined for four strains of enterobacteria. Comparison with the known sequences of EcoK and EcoR124 indicates that all are homologous, though they fall into three families (exemplified by EcoK, EcoA, and EcoR124), the first two of which are apparently allelic. The extent of amino acid sequence identity between EcoK and EcoA is so low that the genes encoding them might be better termed pseudoalleles; this almost certainly reflects genetic exchange among highly divergent species. Within the EcoK family the ratio of intra- to interspecific divergence is very high. The extent of divergence between the genes from Escherichia coli K-12 and Salmonella typhimurium LT2 is similar to that for other genes with the same level of codon usage bias. In contrast, intraspecific divergence (between E. coli strains B and K-12) is extremely high and may reflect the action of frequency-dependent selection mediated by bacteriophages. There is also evidence of lateral transfer of a short sequence between E. coli and S. typhimurium.

Amino Acid Sequence

Human infection by genetically diverse SIVSM-related HIV-2 in west Africa.

Our understanding of the biology and origins of human immunodeficiency virus type 2 (HIV-2) derives from studies of cultured isolates from urban populations experiencing epidemic infection and disease. To test the hypothesis that such isolates might represent only a subset of a larger, genetically more diverse group of viruses, we used nested polymerase chain reactions to characterize HIV-2 sequences in uncultured mononuclear blood cells of two healthy Liberian agricultural workers, from whom virus isolation was repeatedly unsuccessful, and from a culture-positive symptomatic urban dweller. Analysis of pol, env and long terminal repeat regions revealed the presence of three highly divergent HIV-2 strains, one of which (from one of the healthy subjects) was significantly more closely related to simian immunodeficiency viruses infecting sooty mangabeys and rhesus macaques (SIVSM/SIVMAC) than to any virus of human derivation. This subject also harboured multiply defective viral genotypes that resulted from hypermutation of G to A bases. Our results indicate that HIV-2, SIVSM and SIVMAC comprise a single, highly diverse group of lentiviruses which cannot be separated into distinct phylogenetic lineages according to species of origin.

Adult

GCWIND: a microcomputer program for identifying open reading frames according to codon positional G+C content.

GCWIND is a microcomputer (IBM-PC compatible) program for the identification of protein-coding open reading frames. The program is similar to the FRAME program, but the latter has only been implemented for a specialized graphics package. The base compositions (%G+C) for each of the three possible reading phases through the DNA sequence are displayed separately, together with the positions of potential translation initiation and termination codons (on the leading and complementary strands), to provide an immediate representation of those regions within the sequence that have coding potential.

Codon

Molecular population genetics of Escherichia coli: DNA sequence diversity at the celC, crr, and gutB loci of natural isolates.

The DNA sequences of three genes--celC, crr, and gutB--have been determined for each of 11 or 12 natural isolates of Escherichia coli from the ECOR collection. These genes encode the phosphoenolpyruvate-dependent phosphotransferase-system enzyme III proteins specific for beta-glucoside sugars (celC), glucose (crr), and glucitol (gutB), respectively. There is little evidence of recombination at or among these loci; among these strains, relationships inferred from each gene are largely consistent with each other and with the relationship inferred from multilocus enzyme electrophoresis. DNA sequence diversity is similar for all three genes, particularly when silent (synonymous) sites only are considered. This is surprising because there is much stronger codon usage bias at crr than at celC or gutB. The extent of divergence in the protein sequences encoded by these three genes varies considerably. The constitutively expressed glucose-specific enzyme is completely conserved. It is surprising that the inducible glucitol-specific enzyme, which is functional, is more variable than the cellobiose-specific enzyme, which is cryptic; the latter might be expected to be under less (if any) purifying selection.

Amino Acid Sequence

DNA sequence variability at the rplX locus of Bacillus subtilis.

The pattern and extent of DNA sequence variability at the rplX locus (encoding ribosomal protein L24) has been investigated in nine strains of Bacillus subtilis. Overall, there is a very low level of nucleotide diversity, even at silent sites, which is probably due to selection among synonymous codons. By analogy with Escherichia coli, there may also be some effect of the relative proximity of rplX to the chromosomal origin of replication. The small number of nucleotide substitutions are non-randomly distributed: all of the synonymous changes are in valine codons. From the sequence differences the strains can be divided into two groups, which are not coincident with their previous classification; this observation is consistent with recombination among strains.

Amino Acid Sequence

Complete nucleotide sequence, genome organization, and biological properties of human immunodeficiency virus type 1 in vivo: evidence for limited defectiveness and complementation.

Previous studies of the genetic and biologic characteristics of human immunodeficiency virus type 1 (HIV-1) have by necessity used tissue culture-derived virus. We recently reported the molecular cloning of four full-length HIV-1 genomes directly from uncultured human brain tissue (Y. Li, J. C. Kappes, J. A. Conway, R. W. Price, G. M. Shaw, and B. H. Hahn, J. Virol. 65:3973-3985, 1991). In this report, we describe the biologic properties of these four clones and the complete nucleotide sequences and genome organization of two of them. Clones HIV-1YU-2 and HIV-1YU-10 were 9,174 and 9,176 nucleotides in length, differed by 0.26% in nucleotide sequence, and except for a frameshift mutation in the pol gene in HIV-1YU-10, contained open reading frames corresponding to 5'-gag-pol-vif-vpr-tat-rev-vpu-env-nef-3' flanked by long terminal repeats. HIV-1YU-2 was fully replication competent, while HIV-1YU-10 and two other clones, HIV-1YU-21 and HIV-1YU-32, were defective. All three defective clones, however, when transfected into Cos-1 cells in any pairwise combination, yielded virions that were replication competent and transmissible by cell-free passage. The cellular host range of HIV-1YU-2 was strictly limited to primary T lymphocytes and monocyte-macrophages, a property conferred by its external envelope glycoprotein. Phylogenetic analyses of HIV-1YU-2 gene sequences revealed this virus to be a member of the North American/European HIV-1 subgroup, with specific similarity to other monocyte-tropic viruses in its V3 envelope amino acid sequence. These results indicate that HIV-1 infection of brain is characterized by the persistence of mixtures of fully competent, minimally defective, and more substantially altered viral forms and that complementation among them is readily attainable. In addition, the limited degree of genotypic heterogeneity observed among HIV-1YU and other brain-derived viruses and their preferential tropism for monocyte-macrophages suggest that viral replication within the central nervous system may differ from that within the peripheral lymphoid compartment in significant and clinically important ways. The availability of genetically and biologically well characterized HIV-1 clones from uncultured human tissue should facilitate future studies of virus-cell interactions relevant to viral pathogenesis and drug and vaccine development.

AIDS Dementia Complex

The salmon gene encoding apolipoprotein A-I: cDNA sequence, tissue expression and evolution.

A cDNA encoding an apolipoprotein (Apo) has been isolated from the Atlantic salmon (Salmo salar) and sequenced. It encodes a peptide of 258 amino acids (aa), including a signal peptide of 18 aa, with 5'- and 3'-untranslated regions of the mRNA of 12 and 329 nucleotides, respectively. The protein has structural features in common with other Apo's of human and avian origin, including conserved sequences in the signal peptide and a series of internal repeats of 22 aa. The sequence has been identified as salmon Apo A-I (sApoA-I), and has 23% aa identity with human ApoA-I. Northern-blot analysis using the sApoA-I cDNA probe against total RNA prepared from several salmon tissues detects the expression of this gene in liver, intestine and muscle. A phylogenetic analysis reveals that the mammalian ApoA-I, ApoA-IV and Apo-E aa sequences are more closely related to each other than any of them are to sApoA-I. This suggests that the duplication events, from which A-I, A-IV and E arose, occurred after the divergence of the tetrapod and teleost ancestors.

Amino Acid Sequence

Synonymous nucleotide substitution rates in mammalian genes: implications for the molecular clock and the relationship of mammalian orders.

Synonymous substitution rates have been estimated for 58 genes compared among primates, artiodactyls, and rodents. Although silent sites might be expected to be neutral, there is substantial rate variation among genes within each lineage. Some of the rate variation is associated with G + C content: genes with intermediate G + C values have the highest rates. Nevertheless, considerable heterogeneity remains after correcting for G + C content. Synonymous substitution rates also vary among lineages, but the relative rates of genes are well conserved in different lineages. Certain genes have also been sequenced in a fourth order (lagomorph or carnivore), and these data have been used to investigate mammalian phylogeny. Data on lagomorphs are consistent with a star phylogeny, but there is evidence that carnivores and artiodactyls are sister groups. Genes sequenced in both rat and mouse suggest that the increased substitution rate in rodents has occurred since the rat/mouse divergence.

Animals

Codon usage in Aspergillus nidulans.

Synonymous codon usage in genes from the ascomycete (filamentous) fungus Aspergillus nidulans has been investigated. A total of 45 gene sequences has been analysed. Multivariate statistical analysis has been used to identify a single major trend among genes. At one end of this trend are lowly expressed genes, whereas at the other extreme lie genes known or expected to be highly expressed. The major trend is from nearly random codon usage (in the lowly expressed genes) to codon usage that is highly biased towards a set of 19-20 "optimal" codons. The G + C content of the A. nidulans genome is close to 50%, indicating little overall mutational bias, and so the codon usage of lowly expressed genes is as expected in the absence of selection pressure at silent sites. Most of the optimal codons are C- or G- ending, making highly expressed genes more G + C-rich at silent sites.

Aspergillus nidulans

Determinants of DNA sequence divergence between Escherichia coli and Salmonella typhimurium: codon usage, map position, and concerted evolution.

The nature and extent of DNA sequence divergence between homologous protein-coding genes from Escherichia coli and Salmonella typhimurium have been examined. The degree of divergence varies greatly among genes at both synonymous (silent) and nonsynonymous sites. Much of the variation in silent substitution rates can be explained by natural selection on synonymous codon usage, varying in intensity with gene expression level. Silent substitution rates also vary significantly with chromosomal location, with genes near oriC having lower divergence. Certain genes have been examined in more detail. In particular, the duplicate genes encoding elongation factor Tu, tufA and tufB, from S. typhimurium have been compared to their E. coli homologues. As expected these very highly expressed genes have high codon usage bias and have diverged very little between the two species. Interestingly, these genes, which are widely spaced on the bacterial chromosome, also appear to be undergoing concerted evolution, i.e., there has been exchange between the loci subsequent to the divergence of the two species.

Base Sequence

Mitochondrial DNA sequence divergence in the Melanogaster and oriental species subgroups of Drosophila.

The nucleotide sequence of a segment of the mitochondrial DNA from three Drosophila species (D. erecta, D. eugracilis, and D. takahashii), belonging to different subgroups of the melanogaster group has been determined. The segment encompasses three complete tRNA genes (tRNAtrp, tRNAcys, and tRNAtyr) and portions of two protein-coding genes: the subunit 2 of the NADH dehydrogenase (ND2) and the subunit 1 of the cytochrome oxidase (COI). Comparisons also involve homologous sequences already known for four other Drosophila species of the melanogaster group. Length differences were confined in the intergenic region where a long stretch of AT repeats was observed in one of the species analyzed. The three tRNA genes exhibit very different evolutionary rates, the most slowly evolving one, tRNAtyr, is adjacent to the 5' end of COI; tRNAs in similar positions have been previously shown to evolve slowly because they are probably involved in transcript processing. Although the rate of synonymous substitutions was very similar between ND2 and COI genes there were strong discrepancies between them in terms of the number of nonsynonymous substitutions. Differences have also been found in G + C content of the genes, which are likely to be linked to different selective pressures. There is a reduction in G + C content in the region where selective constraints are reduced. This suggests the existence of different levels of constraints along the sequenced segment. An overall analysis of the types of substitutions showed a decrease in A + T content during the course of evolution of the species.

Animals

ERIC sequences: a novel family of repetitive elements in the genomes of Escherichia coli, Salmonella typhimurium and other enterobacteria.

We describe a family of highly conserved, Enterobacterial Repetitive Intergenic Consensus (ERIC) sequences, 14 of which have been identified in Escherichia coli and Salmonella typhimurium and a further three in other enterobacterial species (Yersinia pseudotuberculosis, Klebsiella pneumoniae and Vibrio cholerae). ERIC sequences are 126 bp long and appear to be restricted to transcribed regions of the genome, either in intergenic regions of polycistronic operons or in untranslated regions upstream or downstream of open reading frames. ERIC sequences are highly conserved at the nucleotide sequence level but their chromosomal locations differ between species. Several features of ERIC sequences resemble those of REP sequences (Stern et al., 1984) although the nucleotide sequence is entirely different. The question of whether ERICs have a specific function, or represent a form of 'selfish' DNA, is discussed.

Base Sequence

Molecular phylogeny of Rodentia, Lagomorpha, Primates, Artiodactyla, and Carnivora and molecular clocks.

Phylogenetic analysis of DNA sequences from primates, rodents, lagomorphs, artiodactyls, carnivores, and birds strongly suggests that the order Rodentia is an outgroup to the other four mammalian orders and that Artiodactyla and Carnivora belong to a superordinal clade. Further, there is strong evidence against the Glires concept, which unites Lagomorpha and Rodentia. The radiation among Lagomorpha, Primates, and Artiodactyla--Carnivora is very bush-like, but there is some evidence that Lagomorpha has branched off first. Thus, the branching sequence for these five orders of mammals seems to be Rodentia, Lagomorpha, Primates, Artiodactyla, and Carnivora. The branching date for Rodentia could be as early as 100 million years ago. The rate of nucleotide substitution in the rodent lineage is shown to be at least 1.5 times higher than those in the other four mammalian lineages.

Animals

Processes of genome evolution reflected by base frequency differences among Serratia marcescens genes.

The G + C content of silent sites in codons varies greatly among Serratia marcescens genes; the value in any one gene seems to reflect a balance between mutation pressure towards high G + C content and natural selection constraining choice among synonymous codons. Interestingly, non-coding sequences have substantially lower G + C content than silent sites thought to be under little selective constraint.

Base Composition

Chromosomal location and evolutionary rate variation in enterobacterial genes.

The basal rate of DNA sequence evolution in enterobacteria, as seen in the extent of divergence between Escherichia coli and Salmonella typhimurium, varies greatly among genes, even when only "silent" sites are considered. The degree of divergence is clearly related to the level of gene expression, reflecting constraints on synonymous codon choice. However, where this constraint is weak, among genes not expressed at high levels, divergence is also related to the chromosomal location of the gene; it appears that genes furthest away from oriC, the origin of replication, have a mutation rate approximately two times that of genes near oriC.

Bias