Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96Linked to original sources

Molecular markers of serine protease evolution.

The evolutionary history of serine proteases can be accounted for by highly conserved amino acids that form crucial structural and chemical elements of the catalytic apparatus. These residues display non- random dichotomies in either amino acid choice or serine codon usage and serve as discrete markers for tracking changes in the active site environment and supporting structures. These markers categorize serine proteases of the chymotrypsin-like, subtilisin-like and alpha/beta-hydrolase fold clans according to phylogenetic lineages, and indicate the relative ages and order of appearance of those lineages. A common theme among these three unrelated clans of serine proteases is the development or maintenance of a catalytic tetrad, the fourth member of which is a Ser or Cys whose side chain helps stabilize other residues of the standard catalytic triad. A genetic mechanism for mutation of conserved markers, domain duplication followed by gene splitting, is suggested by analysis of evolutionary markers from newly sequenced genes with multiple protease domains.

Chymotrypsin↗

Nucleotide sequence of the Adh gene region of Drosophila pseudoobscura: evolutionary change and evidence for an ancient gene duplication.

The alcohol dehydrogenase (Adh) locus (ADH; alcohol: NAD+ oxidoreductase, EC 1.1.1.1) of Drosophila pseudoobscura was cloned and sequenced. Forty-five percent of the "effectively silent sites" have changed between Adh in D. pseudoobscura of the obscura species group and the homologous DNA sequence in D. mauritiana, the latter representing the melanogaster species group. The untranslated leader sequence of the adult transcript of D. pseudoobscura has two deletions relative to the D. mauritiana message. The ADH protein sequences of D. pseudoobscura is missing the third and fourth amino acids at the N-terminus relative to the D. mauritiana enzyme. Of the remaining 254 amino acid positions, 27 (10.64%) differ between the two species. Amino acid replacements are randomly distributed into hydrophilic and hydrophobic domains of ADH. However, replacement substitutions are distributed nonrandomly across the three exons among D. pseudoobscura and members of the melanogaster subgroup, suggesting that functional constraints across the exons are different. Surprisingly, silent substitutions are also nonrandomly distributed with the third exon being the most divergent. This pattern suggests possible selective constraints on supposedly neutral silent substitutions and/or variation in underlying mutation rates across the gene. The presence of transcriptional and translational signals at the beginning and end of conserved sequences 3' to Adh implies the existence of a previously undescribed gene. Codon usage and patterns of nucleotide divergence are consistent with a protein coding function for this gene. In addition, conservation of nucleotide and amino acid sequence and similarity in hydropathy plots suggests that the gene 3' to Adh represents an ancient duplication of the Adh gene.

Alcohol Dehydrogenase↗

Rates of DNA evolution in Drosophila depend on function and developmental stage of expression.

DNA-sequence divergence of genes expressed in the embryonic stage was compared with the divergence of genes expressed in adults for 13 species of Drosophila representing various degrees of relatedness. DNA-DNA hybridization experiments were conducted using as tracers complementary DNA (cDNA) reversed transcribed from poly(A)+ mRNA isolated from different developmental stages. The results indicate: (1) cDNA is less diverged than total single-copy DNA; (2) cDNA sequences are not in the rapidly evolving fraction of the single-copy genome of Drosophila; (3) early in evolutionary divergence embryonic messages are about half as diverged as adult messages; sequence data from some of the species compared indicate this is likely due to differences in rates of silent substitutions in genes expressed at different stages of development; and (4) at greater evolutionary distance, the differences in embryonic and adult messages disappear; this could be due to lineage-specific shifts in codon usage.

Animals↗

Complete sequence of the mitochondrial DNA of the annelid worm Lumbricus terrestris.

We have determined the complete nucleotide (nt) sequence of the mitochondrial genome of an oligochaete annelid, the earthworm Lumbricus terrestris. This genome contains the 37 genes typical of metazoan mitochondrial DNA (mtDNA), including ATPase8, which is missing from some invertebrate mtDNAs. ATPase8 is not immediately upstream of ATPase6, a condition found previously only in the mtDNA of snails. All genes are transcribed from the same DNA strand. The largest noncoding region is 384 nt and is characterized by several homopolymer runs, a tract of alternating TA pairs, and potential secondary structures. All protein-encoding genes either overlap the adjacent downstream gene or end at an abbreviated stop codon. In Lumbricus mitochondria, the variation of the genetic code that is typical of most invertebrate mitochondrial genomes is used. Only the codon ATG is used for translation initiation. Lumbricus mtDNA is A + T rich, which appears to affect the codon usage pattern. The DHU arm appears to be unpaired not only in tRNAser(AGN), as is typical for metazoans, but perhaps also in tRNAser(UCN), a condition found previously only in a chiton and among nematodes. Relating the Lumbricus gene organization to those of other major protostome groups requires numerous rearrangements.

Amino Acid Sequence↗

A eubacterial gene conferring spectinomycin resistance on Chlamydomonas reinhardtii: integration into the nuclear genome and gene expression.

We have constructed a dominant selectable marker for nuclear transformation of C. reinhardtii, composed of the coding sequence of the eubacterial aadA gene (conferring spectinomycin resistance) fused to the 5' and 3' untranslated regions of the endogenous RbcS2 gene. Spectinomycin-resistant transformants isolated by direct selection (1) contain the chimeric gene(s) stably integrated into the nuclear genome, (2) show cosegregation of the resistance phenotype with the introduced DNA, and (3) synthesize the expected mRNA and protein. Small linearized plasmids appeared to be inserted into the nuclear genome preferentially through their ends, with relatively few large deletions and/or rearrangements. Multiple copy transformants often integrated concatemers of transforming DNA. Our detailed analysis of the complex integration patterns of plasmid DNA in C. reinhardtii nuclear transformants should be useful for improving the technique of insertional mutagenesis. We also found that the spectinomycin-resistance phenotype was unstable in about half of the transformants. When maintained under nonselective conditions, neither the aadA mRNA nor the AadA protein were detected in these subclones. Moreover, since the integrated transforming DNA was not altered or lost expression of the RbcS2::aadA::RbcS2 gene(s) appears to be repressed. Measurements of transcriptional activity, mRNA accumulation, and mRNA stability suggest that expression of this chimeric gene(s) may also be affected by rapid RNA degradation, presumably due to defects in mRNA processing and, or nuclear export. Thus, both gene silencing and transcript instability, rather than biased codon usage, may explain the difficulties encountered in the expression of foreign genes in the nuclear genome of Chlamydomonas.

Animals↗

The complete DNA sequence of the mitochondrial genome of a "living fossil," the coelacanth (Latimeria chalumnae).

The complete nucleotide sequence of the 16,407-bp mitochondrial genome of the coelacanth (Latimeria chalumnae) was determined. The coelacanth mitochondrial genome order is identical to the consensus vertebrate gene order which is also found in all ray-finned fishes, the lungfish, and most tetrapods. Base composition and codon usage also conform to typical vertebrate patterns. The entire mitochondrial genome was PCR-amplified with 24 sets of primers that are expected to amplify homologous regions in other related vertebrate species. Analyses of the control region of the coelacanth mitochondrial genome revealed the existence of four 22-bp tandem repeats close to its 3' end. The phylogenetic analyses of a large data set combining genes coding for rRNAs, tRNAs, and proteins (16,140 characters) confirmed the phylogenetic position of the coelacanth as a lobe-finned fish; it is more closely related to tetrapods than to ray-finned fishes. However, different phylogenetic methods applied to this largest available molecular data set were unable to resolve unambiguously the relationship of the coelacanth to the two other groups of extant lobe-finned fishes, the lungfishes and the tetrapods. Maximum parsimony favored a lungfish/coelacanth or a lungfish/tetrapod sistergroup relationship depending on which transversion:transition weighting is assumed. Neighbor-joining and maximum likelihood supported a lungfish/tetrapod sistergroup relationship.

Amino Acid Sequence↗

Rate variation of DNA sequence evolution in the Drosophila lineages.

Rate constancy of DNA sequence evolution was examined for three species of Drosophila, using two samples: the published sequences of eight genes from regions of the normal recombination rates and new data of the four AS-C (ac, sc, l'sc and ase) and ci genes. The AS-C and ci genes were chosen because these genes are located in the regions of very reduced recombination in Drosophila melanogaster and their locations remain unchanged throughout the entire lineages involved, yielding less effect of ancestral polymorphism in the study of rate constancy. The synonymous substitution pattern of the three lineages was found to be erratic in both samples. The dispersion index for replacement substitution was relatively high for the per, G6pd and ac genes. A significant heterogeneity was found in the number of synonymous substitutions in the three lineages between the two samples of genes with different recombination rates. This is partly due to a lack of the lineage effect in the D. melanogaster and Drosophila simulans lineages in the AS-C and ci genes in contrast to Akashi's observation of genes in regions of normal recombination. The higher codon bias in Drosophila yakuba as compared with D. melanogaster and D. simulans was observed in the four AS-C genes, which suggests change(s) in action of natural selection involved in codon usage on these genes. Fluctuating selection intensity may also be responsible for the observed locus-lineage interaction effects in synonymous substitution.

Animals↗

The effect of tandem substitutions on the correlation between synonymous and nonsynonymous rates in rodents.

Nonsynonymous substitutions in DNA cause amino acid substitutions while synonymous substitutions in DNA leave amino acids unchanged. The cause of the correlation between the substitution rates at nonsynonymous (K(A)) and synonymous (K(S)) sites in mammals is a contentious issue, and one that impacts on many aspects of molecular evolution. Here we use a large set of orthologous mammalian genes to investigate the causes of the K(A)-K(S) correlation in rodents. The strength of the K(A)-K(S) correlation exceeds the neutral theory expectation when substitution rates are estimated using algorithmic methods, but not when substitution rates are estimated by maximum likelihood. Irrespective of this methodological uncertainty the strength of the K(A)-K(S) correlation appears mostly due to tandem substitutions, an excess of which is generated by substitutional nonindependence. Doublet mutations cannot explain the excess of tandem synonymous-nonsynonymous substitutions, and substitution patterns indicate that selection on silent sites is the likely cause. We find no evidence for selection on codon usage. The nature of the relationship between synonymous divergence and base composition is unclear because we find a significant correlation if we use maximum-likelihood methods but not if we use algorithmic methods. Finally, we find that K(S) is reduced at the start of genes, which suggests that selection for RNA structure may affect silent sites in mammalian protein-coding genes.

Algorithms↗

The effects of Hill-Robertson interference between weakly selected mutations on patterns of molecular evolution and variation.

Associations between selected alleles and the genetic backgrounds on which they are found can reduce the efficacy of selection. We consider the extent to which such interference, known as the Hill-Robertson effect, acting between weakly selected alleles, can restrict molecular adaptation and affect patterns of polymorphism and divergence. In particular, we focus on synonymous-site mutations, considering the fate of novel variants in a two-locus model and the equilibrium effects of interference with multiple loci and reversible mutation. We find that weak selection Hill-Robertson (wsHR) interference can considerably reduce adaptation, e.g., codon bias, and, to a lesser extent, levels of polymorphism, particularly in regions of low recombination. Interference causes the frequency distribution of segregating sites to resemble that expected from more weakly selected mutations and also generates specific patterns of linkage disequilibrium. While the selection coefficients involved are small, the fitness consequences of wsHR interference across the genome can be considerable. We suggest that wsHR interference is an important force in the evolution of nonrecombining genomes and may explain the unexpected constancy of codon bias across species of very different census population sizes, as well as several unusual features of codon usage in Drosophila.

Alleles↗

Purification of human beta2-adrenergic receptor expressed in methylotrophic yeast Pichia pastoris.

Human beta2-adrenergic receptor is a G-protein-coupled receptor with seven transmembrane helices, and is important in pharmaceutical targeting on pulmonary and cardiovascular diseases. N-terminal histidine-tagged gene constructs with optimized codon usage were designed so as to obtain Pichia pastoris transformants with a high expression level. The constructs were inserted into the pPIC9 vector, and then electroporated into the SMD1168 strains. The highest expression level obtained was about 4 mg/liter-culture broth. The dissociation constant of the receptor in the membrane fraction was 1.2 nM toward CGP-12177 antagonist. The receptor was solubilized with sucrose monolaurate and purified with a series of chromatography steps including anion-exchange, Ni-Sepharose, alprenolol-Agarose, and hydroxyapatite columns. The receptor was heterogeneously glycosylated, showing broad SDS-PAGE bands around 70-90 kDa. After endoglycosidase treatment, the receptor appeared as a single band around 45 kDa, and was further purified with hydroxyapatite and gel-filtration columns. The receptor was eluted as a sharp peak at the gel-filtration elution volume corresponding to a molecular mass of 117 kDa. The saccharide-trimmed receptor thus purified is homogeneous as analyzed with SDS-PAGE, shows the dissociation constant of 4.7 nM toward CGP-12177 antagonist, and is suitable for crystallization experiments.

Chromatography, Gel↗

Fructan biosynthesis in transgenic plants.

Data from plants transformed to accumulate fructan are assessed in the context of natural concentrations of reserve carbohydrates and natural fluxes of carbon in primary metabolism: Transgenic fructan accumulation is universally reported as an instantaneous endpoint concentration. In exceptional cases, concentrations of 60-160 mg g(-1) fresh mass were reported and compare favourably with naturally occurring maximal starch and fructan content in leaves and storage organs. Generally, values were less than 20 mg g(-1) for plants transformed with bacterial genes and <9 mg g(-1) for plant-plant transformants. Superficially, the results indicate a marked modification of carbon partitioning. However, transgenic fructan accumulation was generally constitutive and involved accumulation over time-scales of weeks or months. When calculated as a function of accumulation period, fluxes into the transgenic product were low, in the range 0.00002-0.03 nkat g(-1). By comparison with an estimated minimum daily carbohydrate flux in leaves for a natural fructan-accumulating plant in field conditions (37 nkat g(-1)), transgenic fructan accumulation was only 0.00005-0.08% of primary carbohydrate flux and does not indicate radical modification of carbon partitioning, but rather, a quantitatively minor leakage into transgenic fructan. Possible mechanisms for this low fructan accumulation in the transformants are considered and include: (i) rare codon usage in bacterial genes compared with eukaryotes, (ii) low transgene mRNA concentrations caused by low expression and/or high turnover, (iii) resultant low expression of enzyme protein, (iv) resultant low total enzyme activity, (v) inappropriate kinetic properties of the gene products with respect to substrate concentrations in the host, (vi) in situ product hydrolysis, and (vii) levan toxicity. Transformants expressing bacterial fructan synthesis exhibited a number of aberrant phenotypes such as stunting, leaf bleaching, necrosis, reduced tuber number and mass, tuber cortex discoloration, reduction in starch accumulation, and chloroplast agglutination. In severe cases of developmental aberration, potato tubers were replaced by florets. Possible mechanisms to explain these aberrations are discussed. In most instances, the attempted subcellular targeting of the transgene product was not demonstrated. Where localization was attempted, the transgene product generally mis-localized, for example, to the cell perimeter or to the endomembrane system, instead of the intended target, the vacuole. Fructosyltransferases exhibited different product specificities in planta than in vitro, expression in planta generally favouring the formation of larger fructan oligomers and polymers. This implies a direct influence of the intracellular environment on the capacity for polymerization of fructosyltransferases and may have implications for the mechanism of natural fructan polymerization in vivo.

Bacteria↗

GBRAP: A Comprehensive Database and Tool for Exploring Genomic Diversity Across All Domains of Life.

Evolutionary studies require extensive examination of genomic information across all domains of life. Despite the availability of a large number of genomes through GenBank, the effective visualization or comparison of the information they contain is challenging due to many reasons, including their size. We introduce genome-based retrieval and analysis parser, a comprehensive software tool to analyze genome files, and an online database housing an extensive collection of carefully curated, high-quality genome statistics for all the organisms available in the RefSeq database of National Center for Biotechnology Information. Users can either directly search, or select from precategorized groups, the organisms of their choice and retrieve data, and the output is generated as tables containing more than 200 columns of useful genomic information (base counts, GC content, Shannon entropy, codon usage, etc.) separately calculated for different genomic elements (e.g. coding sequences, introns, transfer RNA, ribosomal RNA, noncoding RNA, etc.). The data are independently displayed (if applicable) for each chromosomal, mitochondrial, plastid, or plasmid sequence. All the data can be visualized on the database or downloaded as comma-separated value or Excel files. The genome-based retrieval and analysis parser database is free to access without any registration and is publicly available at http://tacclab.org/gbrap/.

Software↗

Mammalian housekeeping genes evolve more slowly than tissue-specific genes.

Do housekeeping genes, which are turned on most of the time in almost every tissue, evolve more slowly than genes that are turned on only at specific developmental times or tissues? Recent large-scale gene expression studies enable us to have a better definition of housekeeping genes and to address the above question in detail. In this study, we examined 1581 human-mouse orthologous gene pairs for their patterns of sequence evolution, contrasting housekeeping genes with tissue-specific genes. Our results show that, in comparison to tissue-specific genes, housekeeping genes on average evolve more slowly and are under stronger selective constraints as reflected by significantly smaller values of Ka/Ks. Besides stronger purifying selection, we explored several other factors that can possibly slow down nonsynonymous rates in housekeeping genes. Although mutational bias might slightly slow the nonsynonymous rates in housekeeping genes, it is unlikely to be the major cause of the rate difference between the two types of genes. The codon usage pattern of housekeeping genes does not seem to differ from that of tissue-specific genes. Moreover, contrary to the old textbook concept, we found that approximately 74% of the housekeeping genes in our study belong to multigene families, not significantly different from that of the tissue-specific genes ( approximately 70%). Therefore, the stronger selective constraints on housekeeping genes are not due to a lower degree of genetic redundancy.

Animals↗

Nonneutral evolution of the transcribed pseudogene Makorin1-p1 in mice.

Pseudogenes are nonfunctional relics of formerly functional genes and are thought to evolve neutrally. In some pseudogenes, however, the molecular evolutionary patterns are atypical of neutrally evolving sequences, exhibiting sequence conservation, codon-usage bias, and other features associated with functional genes. Makorin1-p1 is a transcribed pseudogene first identified in the mouse Mus musculus. The transcript of Makorin1-p1 can regulate the stability of the transcript of its paralogous functional gene Makorin1. Specifically, the half-life of Makorin1 mRNA increases significantly in the presence of Makorin1-p1 transcript, and targeted deletion of Makorin1-p1 is lethal in mice. Here, we show that Makorin1-p1 originated after the separation of Mus and Rattus but before the divergence of M. musculus and M. pahari. The transcribed region of Makorin1-p1 exhibits rates of point and indel substitutions that are two to four times lower than those in the untranscribed region, suggesting that the transcribed region is under functional constraint and is not evolving neutrally. Although the transcript of Makorin1-p1 likely functions by its sequence similarity to Makorin1, we find no evidence of gene conversion between them, indicating that functional conservation alone is sufficient to maintain their coordinated evolution. A duplication-degeneration model is proposed to explain how Makorin1-p1 was co-opted into the regulatory system of Makorin1. There are over 10,000 pseudogenes in a typical mammalian genome, and it is plausible that many functional but untranslatable pseudogenes exist. Our results illustrate the potential of using evolutionary analysis to identify such pseudogenes from genome sequences.

Animals↗

Evolution of the AID/APOBEC family of polynucleotide (deoxy)cytidine deaminases.

The AID/APOBEC family (comprising AID, APOBEC1, APOBEC2, and APOBEC3 subgroups) contains members that can deaminate cytidine in RNA and/or DNA and exhibit diverse physiological functions (AID and APOBEC3 deaminating DNA to trigger pathways in adaptive and innate immunity; APOBEC1 mediating apolipoprotein B RNA editing). The founder member APOBEC1, which has been used as a paradigm, is an RNA-editing enzyme with proposed antecedents in yeast. Here, we have undertaken phylogenetic analysis to glean insight into the primary physiological function of the AID/APOBEC family. We find that although the family forms part of a larger superfamily of deaminases distributed throughout the biological world, the AID/APOBEC family itself is restricted to vertebrates with homologs of AID (a DNA deaminase that triggers antibody gene diversification) and of APOBEC2 (unknown function) identifiable in sequence databases from bony fish, birds, amphibians, and mammals. The cloning of an AID homolog from dogfish reveals that AID extends at least as far back as cartilaginous fish. Like mammalian AID, the pufferfish AID homolog can trigger deoxycytidine deamination in DNA but, consistent with its cold-blooded origin, is thermolabile. The fine specificity of its mutator activity and the biased codon usage in pufferfish IgV genes appear broadly similar to that of their mammalian counterparts, consistent with a coevolution of the antibody mutator and its substrate for the optimal targeting of somatic mutation during antibody maturation. By contrast, APOBEC1 and APOBEC3 are later evolutionary arrivals with orthologs not found in pufferfish (although synteny with mammals is maintained in respect of the flanking loci). We conclude that AID and APOBEC2 are likely to be the ancestral members of the AID/APOBEC family (going back to the beginning of vertebrate speciation) with both APOBEC1 and APOBEC3 being mammal-specific derivatives of AID and a complex set of domain shuffling underpinning the expansion and evolution of the primate APOBEC3s.

APOBEC-1 Deaminase↗

Alternatively and constitutively spliced exons are subject to different evolutionary forces.

There has been a controversy on whether alternatively spliced exons (ASEs) evolve faster than constitutively spliced exons (CSEs). Although it has been noted that ASEs are subject to weaker selective constraints than CSEs, so they evolve faster, there have also been studies that indicated slower evolution in ASEs than in CSEs. In this study, we retrieve more than 5,000 human-mouse orthologous exons and calculate the synonymous (KS) and nonsynonymous (KA) substitution rates in these exons. Our results show that ASEs have higher KA values and higher KA/KS ratios than CSEs, indicating faster amino acid-level evolution in ASEs. The faster evolution may be in part due to weaker selective constraints. It is also possible that the faster rate is in part due to faster functional evolution in ASEs. On the other hand, the majority of ASEs have lower KS values than CSEs. With reference to the substitution rate in introns, we show that the KS values in ASEs are close to the neutral substitution rate, whereas the synonymous substitution rate in CSEs has likely been accelerated. The elevated synonymous rate in CSEs is not related to CpG dinucleotides or low-complexity regions of protein but may be weakly related to codon usage bias. The overall trends of higher KA and lower KS in ASEs than in CSEs are also observed in human-rat and mouse-rat comparisons. Therefore, our observations hold for mammals of different molecular clocks.

Animals↗

The rate of adaptive evolution in enteric bacteria.

Here we estimate the rate of adaptive substitution in a set of 410 genes that are present in 6 Escherichia coli and 6 Salmonella enterica genomes. We estimate that more than 50% of amino acid substitutions in this set of genes have been fixed by positive selection between the E. coli and S. enterica lineages. We also show that the proportion of adaptive substitutions is uncorrelated with the rate of amino acid substitution or gene function but that it may be correlated with levels of synonymous codon usage bias.

Adaptation, Biological↗

Microcomputer programs for DNA sequence analysis.

Computer programs are described which allow (a) analysis of DNA sequences to be performed on a laboratory microcomputer or (b) transfer of DNA sequences between a laboratory microcomputer and another computer system, such as a DNA library. The sequence analysis programs are interactive, do not require prior experience with computers and in many other respects resemble programs which have been written for larger computer systems (1-7). The user enters sequence data into a text file, accesses this file with the programs, and is then able to (a) search for restriction enzyme sites or other specified sequences, (b) translate in one or more reading frames in one or both directions in order to find open reading frames, or (c) determine codon usage in the sequence in one or more given reading frames. The results are given in table format and a restriction map is generated. The modem program permits collection of large amounts of data from a sequence library into a permanent file on the microcomputer disc system, or transfer of laboratory data in the reverse direction to a remote computer system.

Base Sequence↗