Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Sequence divergence of an archaebacterial gene cloned from a mesophilic and a thermophilic methanogen.

A 1.6-kb fragment of DNA from the thermophilic, methane-producing, anaerobic archaebacterium Methanobacterium thermoautotrophicum delta H has been cloned and sequenced. This DNA complements mutations in both the purE1 and purE2 loci of Escherichia coli. The sequence of the M. thermoautotrophicum DNA predicts that complementation in E. coli results from the synthesis of a polypeptide with a molecular weight of 36,249. A polypeptide apparently of this molecular weight is synthesized in E. coli minicells containing recombinant plasmids that carry the cloned fragment of methanogen DNA. We have previously cloned and sequenced a purE-complementing gene from the mesophilic methanogen Methanobrevibacter smithii. The two methanogen-derived purE-complementing genes are 53% homologous and encode polypeptides that are 45% homologous in their amino acid sequences but would be 74% homologous if conservative amino acid substitutions were considered as maintaining sequence homology. The genome of M. thermoautotrophicum has a molar G + C content of 49.7%, whereas the genome of M. smithii is 30.6% G + C. Conservation of encoded amino acids while accommodating the very different G + C contents is accomplished by use of different codons that encode the same amino acid. The majority of base changes occur at the third codon position. The intergenic regions of the cloned M. thermoautotrophicum DNA contain sequences previously identified as ribosome binding sites and as putative methanogen promoters. Although the two purE-complementing genes are apparently derived from a common ancestor, only the gene from M. smithii maintains a codon usage that conforms to the RNY rule.

Amino Acid Sequence↗

Synthetic cryIIIA gene from Bacillus thuringiensis improved for high expression in plants.

A 1974 bp synthetic gene was constructed from chemically synthesized oligonucleotides in order to improve transgenic protein expression of the cryIIIA gene from Bacillus thuringiensis var. tenebrionis in transgenic tobacco. The crystal toxin genes (cry) from B. thuringiensis are difficult to express in plants even when under the control of efficient plant regulatory sequences. We identified and eliminated five classes of sequence found throughout the cryIIIA gene that mimic eukaryotic processing signals and which may be responsible for the low levels of transcription and translation. Furthermore, the GC content of the gene was raised from 36% to 49% and the codon usage was changed to be more plant-like. When the synthetic gene was placed behind the cauliflower mosaic virus 35S promoter and the alfalfa mosaic virus translational enhancer, up to 0.6% of the total protein in transgenic tobacco plants was cryIIIA as measured from immunoblot analysis. Bioassay data using potato beetle larvae confirmed this estimate.

Animals↗

Phylogenetic affinity of mitochondria of Euglena gracilis and kinetoplastids using cytochrome oxidase I and hsp60.

The mitochondrial DNA-encoded cytochrome oxidase subunit I (COI) gene and the nuclear DNA-encoded hsp60 gene from the euglenoid protozoan Euglena gracilis were cloned and sequenced. The COI sequence represents the first example of a mitochondrial genome-encoded gene from this organism. This gene contains seven TGG tryptophan codons and no TGA tryptophan codons, suggesting the use of the universal genetic code. This differs from the situation in the mitochondrion of the related kinetoplastid protozoa, in which TGA codes for tryptophan. In addition, a complete absence of CGN triplets may imply the lack of the corresponding tRNA species. COI cDNAs from E. gracilis possess short 5' and 3' untranslated transcribed sequences and lack a 3' poly[A] tail. The COI gene does not require uridine insertion/ deletion RNA editing, as occurs in kinetoplastid mitochondria, to be functional, and no short guide RNA-like molecules could be visualized by labeling total mitochondrial RNA with [alpha-32P]GTP and guanylyl transferase. In spite of the differences in codon usage and the 3' end structures of mRNAs, phylogenetic analysis using the COI and hsp60 protein sequences suggests a monophyletic relationship between the mitochondrial genomes of E. gracilis and of the kinetoplastids, which is consistent with the phylogenetic relationship of these groups previously obtained using nuclear ribosomal RNA sequences.

Animals↗

Synonymous substitution rates in Drosophila: mitochondrial versus nuclear genes.

Synonymous substitution rates in mitochondrial and nuclear genes of Drosophila were compared. To make accurate comparisons, we considered the following: (1) relative synonymous rates, which do not require divergence time estimates, should be used; (2) methods estimating divergence should take into account base composition; (3) only very closely related species should be used to avoid effects of saturation; (4) the heterogeneity of rates should be examined. We modified the methods estimating synonymous substitution numbers to account for base composition bias. By using these methods, we found that mitochondrial genes have 1.7-3.4 times higher synonymous substitution rates than the fastest nuclear genes or 4.5-9.0 times higher rates than the average nuclear genes. The average rate of synonymous transversions was 2.7 (estimated from the melanogaster species subgroup) or 2.9 (estimated from the obscura group) times higher in mitochondrial genes than in nuclear genes. Synonymous transversions in mitochondrial genes occurred at an approximately equivalent rate to those in the fastest nuclear genes. This last result is not consistent with the hypothesis that the difference in turnover rates between mitochondrial and nuclear genomes is the major factor determining higher synonymous substitution rates in mtDNA. We conclude that the difference in synonymous substitution rates is due to a combination of two factors: a higher transitional mutation rate in mtDNA and constraints on nuclear genes due to selection for codon usage.

Animals↗

Molecular characterization of the principal symbiotic bacteria of the weevil Sitophilus oryzae: a peculiar G + C content of an endocytobiotic DNA.

The principal intracellular symbiotic bacteria of the cereal weevil Sitophilus oryzae were characterized using the sequence of the 16S rDNA gene (rrs gene) and G + C content analysis. Polymerase chain reaction amplification with universal eubacterial primers of the rrs gene showed a single expected sequence of 1,501 bp. Comparison of this sequence with the available database sequences placed the intracellular bacteria of S. oryzae as members of the Enterobacteriaceae family, closely related to the free-living bacteria, Erwinia herbicola and Escherichia coli, and the endocytobiotic bacteria of the tsetse fly and aphids. Moreover, by high-performance liquid chromatography, we measured the genomic G + C content of the S. oryzae principal endocytobiotes (SOPE) as 54%, while the known genomic G + C content of most intracellular bacteria is about 39.5%. Furthermore, based on the third codon position G + C content and the rrs gene G + C content, we demonstrated that most intracellular bacteria except SOPE are A + T biased irrespective of their phylogenetic position. Finally, using the hsp60 gene sequence, the codon usage of SOPE was compared with that of two phylogenetically closely related bacteria: E. coli, a free-living bacterium, and Buchnera aphidicola, the intracellular symbiotic bacteria of aphids. Taken together, these results show a peculiar and distinctly different DNA composition of SOPE with respect to the other obligate intracellular bacteria, and, combined with biological and biochemical data, they elucidate the evolution of symbiosis in S. oryzae.

Animals↗

Enhanced evolvability in immunoglobulin V genes under somatic hypermutation.

Darwinian theory requires that mutations be produced in a nonanticipatory manner; it is nonetheless consistent to suggest that mutations that have repeatedly led to nonviable phenotypes would be introduced less frequently than others-if under appropriate genetic control. Immunoglobulins produced during infection acquire point mutations that are subsequently selected for improved binding to the eliciting antigen. We and others have speculated that an enhancement of mutability in the complementarity-determining regions (CDR; where mutations have a greater chance of being advantageous) and/or decrement of mutability in the framework regions (FR; where mutations are more likely to be lethal) may be accomplished by differential codon usage in concert with the known sequence specificity of the hypermutation mechanism. We have examined 115 nonproductively rearranged human Ig sequences. The mutation patterns in these unexpressed genes are unselected and therefore directly reflect inherent mutation biases. Using a chi2 test, we have shown that the number of mutations in the CDRs is significantly higher than the number of mutations found in the FRs, providing direct evidence for the hypothesis that mutations are preferentially targeted into the CDRs.

Arthritis↗

Phylogeny and the evolution of the Amylase multigenes in the Drosophila montium species subgroup.

To investigate the phylogenetic relationships and molecular evolution of alpha-amylase (Amy) genes in the Drosophila montium species subgroup, we constructed the phylogenetic tree of the Amy genes from 40 species from the montium subgroup. On our tree the sequences of the auraria, kikkawai, and jambulina complexes formed distinct tight clusters. However, there were a few inconsistencies between the clustering pattern of the sequences and taxonomic classification in the kikkawai and jambulina complexes. Sequences of species from other complexes (bocqueti, bakoue, nikananu, and serrata) often did not cluster with their respective taxonomic groups. This suggests that relationships among the Amy genes may be different from those among species due to their particular evolution. Alternatively, the current taxonomy of the investigated species is unreliable. Two types of divergent paralogous Amy genes, the so-called Amy1- and Amy3-type genes, previously identified in the D. kikkawai complex, were common in the montium subgroup, suggesting that the duplication event from which these genes originate is as ancient as the subgroup or it could even predate its differentiation. Thc Amy1-type genes were closer to the Amy genes of D. melanogaster and D. pseudoobscura than to the Amy3-type genes. In the Amy1-type genes, the loss of the ancestral intron occurred independently in the auraria complex and in several Afrotropical species. The GC content at synonymous third codon positions (GC3s) of the Amy1-type genes was higher than that of the Amy3-type genes. Furthermore, the Amy1-type genes had more biased codon usage than the Amy3-type genes. The correlations between GC3s and GC content in the introns (GCi) differed between these two Amy-type genes. These findings suggest that the evolutionary forces that have affected silent sites of the two Amy-type genes in the montium species subgroup may differ.

Amylases↗

Relative rates of nucleotide substitution in frogs.

Accurate estimation of relative mutation rates of mitochondrial DNA (mtDNA) and single-copy nuclear DNA (scnDNA) within lineages contributes to a general understanding of molecular evolutionary processes and facilitates making demographic inferences from population genetic data. The rate of divergence at synonymous sites ( K(s)) may be used as a surrogate for mutation rate. Such data are available for few organisms and no amphibians. Relative to mammals and birds, amphibian mtDNA is thought to evolve slowly, and the K(s) ratio of mtDNA to scnDNA would be expected to be low as well. Relative K(s) was estimated from a mitochondrial gene, ND2, and a nuclear gene, c-myc, using both "approximate" and likelihood methods. Three lineages of congeneric frogs were studied and this ratio was found to be approximately 16, the highest of previously reported ratios. No evidence of a low K(s) in the nuclear gene was found: c-myc codon usage was not biased, the K(s) was double the intron divergence rate, and the absolute K(s) was similar to estimates obtained here for other genes from other frog species. A high K(s) in mitochondrial vs. nuclear genes was unexpected in light of previous reports of a slow rate of mtDNA evolution in amphibians. These results highlight the need for further investigation of the effects of life history on mutation rates.

Animals↗

Comparison of the yeast proteome to other fungal genomes to find core fungal genes.

The purpose of this research was to search for evolutionarily conserved fungal sequences to test the hypothesis that fungi have a set of core genes that are not found in other organisms, as these genes may indicate what makes fungi different from other organisms. By comparing 6355 predicted or known yeast (Saccharomyces cerevisiae) genes to the genomes of 13 other fungi using Standalone TBLASTN at an e-value <1E-5, a list of 3340 yeast genes was obtained with homologs present in at least 12 of 14 fungal genomes. By comparing these common fungal genes to complete genomes of animals (Fugu rubripes, Caenorhabditis elegans), plants (Arabidopsis thaliana, Oryza sativa), and bacteria (Agrobacterium tumefaciens, Xylella fastidiosa), a list of common fungal genes with homologs in these plants, animals, and bacteria was produced (938 genes), as well as a list of exclusively fungal genes without homologs in these other genomes (60 genes). To ensure that the 60 genes were exclusively fungal, these were compared using TBLASTN to the major sequence databases at GenBank: NR (nonredundant), EST (expressed sequence tags), GSS (genome survey sequences), and HTGS (unfinished high-throughput genome sequences). This resulted in 17 yeast genes with homologs in other fungal genomes, but without known homologs in other organisms. These 17 core, fungal genes were not found to differ from other yeast genes in GC content or codon usage patterns. More intensive study is required of these 17 genes and other common fungal genes to discover unique features of fungi compared to other organisms.

Computational Biology↗

A comparative categorization of protein function encoded in bacterial or archeal genomic islands.

Genomes of prokaryotes harbor genomic islands (GIs), which are frequently acquired via horizontal gene transfer (HGT). Here I present an analysis of GIs with respect to gene-encoded functions. GIs were identified by statistical analysis of codon usage and clustering. Genes classified as putatively alien (pA) or putatively native (pN) were categorized according to the COG database. Among pA and pN genes, the distribution of COG functions and classes were studied for different groupings of prokaryotes. Groups were formed according to taxonomical relation or habitats. In all groups, genes related to class L (replication, recombination, and repair) were statistically significantly overrepresented in GIs. GIs of bacteria and archaea showed a distinct pattern of preferences. In archeal GIs, genes belonging to COG class M (cell wall/membrane/envelope biogenesis) or Q (secondary metabolites biosynthesis, transport, and catabolism) were more frequent. In bacterial GIs, genes of classes U (intracellular trafficking, secretion, and vesicular transport), N (cell motility), and V (defense mechanisms) were predominant. Underrepresentation was strongest for genes belonging to class J (translation, ribosomal structure, and biogenesis). Among single COG functions overrepresented in GIs were transferases and transporters. In both superkingdoms, HGT enhances genomic content by meeting demands that are independent of the studied habitats. These findings are in agreement with the complexity theory, which predicts the preferential import of operational genes. However, only specific subsets of operational genes were enriched in GIs. Modification of the cell envelope, cell motility, secretion, and protection of cellular DNA are major issues in HGT.

Amino Acid Sequence↗

Distribution and evolution of bacteriophage WO in Wolbachia, the endosymbiont causing sexual alterations in arthropods.

Wolbachia are obligatory intracellular and maternally inherited bacteria, known to infect many species of arthropod. In this study, we discovered a bacteriophage-like genetic element in Wolbachia, which was tentatively named bacteriophage WO. The phylogenetic tree based on phage WO genes of several Wolbachia strains was not congruent with that based on chromosomal genes of the same strains, suggesting that phage WO was active and horizontally transmitted among various Wolbachia strains. All the strains of Wolbachia used in this study were infected with phage WO. Although the phage genome contained genes of diverse origins, the average G+C content and codon usage of these genes were quite similar to those of a chromosomal gene of Wolbachia. These results raised the possibility that phage WO has been associated with Wolbachia for a very long time, conferring some benefit to its hosts. The evolution and possible roles of phage WO in various reproductive alterations of insects caused by Wolbachia are discussed.

Amino Acid Sequence↗

A survey of the molecular evolutionary dynamics of twenty-five multigene families from four grass taxa.

We surveyed the molecular evolutionary characteristics of 25 plant gene families, with the goal of better understanding general processes in plant gene family evolution. The survey was based on 247 GenBank sequences representing four grass species (maize, rice, wheat, and barley). For each gene family, orthology and paralogy relationships were uncertain. Recognizing this uncertainty, we characterized the molecular evolution of each gene family in four ways. First, we calculated the ratio of nonsynonymous to synonymous substitutions (d(N)/d(S)) both on branches of gene phylogenies and across codons. Our results indicated that the d(N)/d(S) ratio was statistically heterogeneous across branches in 17 of 25 (68%) gene families. The vast majority of d(N)/d(S) estimates were <<1.0, suggestive of selective constraint on amino acid replacements, and no estimates were >1.0, either across phylogenetic lineages or across codons. Second, we tested separately for nonsynonymous and synonymous molecular clocks. Sixty-eight percent of gene families rejected a nonsynonymous molecular clock, and 52% of gene families rejected a synonymous molecular clock. Thus, most gene families in this study deviated from clock-like evolution at either synonymous or nonsynonymous sites. Third, we calculated the effective number of codons and the proportion of G+C synonymous sites for each sequence in each gene family. One or both quantities vary significantly within 18 of 25 gene families. Finally, we tested for gene conversion, and only six gene families provided evidence of gene conversion events. Altogether, evolution for these 25 gene families is marked by selective constraint that varies among gene family members, a lack of molecular clock at both synonymous and nonsynonymous sites, and substantial variation in codon usage.

DNA↗

Production of active bovine cathepsin C (dipeptidyl aminopeptidase I) in the methylotrophic yeast Candida boidinii.

The heterologous production of active bovine cathepsin C (CTC; dipeptidyl aminopeptidase I) was investigated. Attempts to express CTC in Escherichia coli were hampered by formation of inclusion bodies that were partially degraded. To overcome this impediment, secretion of recombinant CTC was attempted in the methylotrophic yeast Candida boidinii. A DNA fragment encoding bovine procathepsin C was synthesized based on preferred codon usage in C. boidinii and placed downstream of the C. boidinii proteinase A signal sequence resulting in secretion of active CTC into the culture medium. The gene was expressed under the control of the methanol-inducible formate dehydrogenase gene promoter. Production levels were significantly improved by using a protease-deficient strain, changing medium composition, and by lowering the temperature of induction. When the recombinant C. boidinii was grown for 90 h in a jar-fermenter, active CTC was secreted with a yield of up to approximately 12 mg/l.

Amino Acid Sequence↗

A delta-endotoxin encoded in Pseudomonas fluorescens displays a high degree of insecticidal activity.

The short field-life of Bacillus thuringiensis (Bt) insecticidal crystal protein has limited its use. When the Bt toxin is produced in Pseudomonas fluorescens it can be encapsulated and retain its effectiveness for two to three times longer than other Bt formulations. In order to improve Bt expression, we have synthesized cryIA(c) Bt delta-endotoxin encoding region (GenBank AF537267) according to the usage codon of P. fluorescens and transformed the Bt toxin expression cassette into P. fluorescens strains. T7 RNA polymerase and the T7 promoter system were used to control expression of Bt toxin. SDS-PAGE and Western blotting assay revealed that the delta-endotoxin was expressed as 8% of the total protein in P. fluorescens. In in vitro tests, release of toxin from dead bacteria was demonstrated. Supplementation of diets with Bt toxin-containing Pseudomonas bacterium resulted in high mortality of cabbage butterfly ( Pieris brassicae) larvae.

Animals↗

Expression of the sweet-tasting plant protein brazzein in Escherichia coli and Lactococcus lactis: a path toward sweet lactic acid bacteria.

Brazzein is an intensely sweet-tasting plant protein with good stability, which makes it an attractive alternative to sucrose. A brazzein gene has been designed, synthesized, and expressed in Escherichia coli at 30 degrees C to yield brazzein in a soluble form and in considerable quantity. Antibodies have been produced using brazzein fused to His-tag. Brazzein without the tag was sweet and resembled closely the taste of its native counterpart. The brazzein gene was also expressed in Lactococcus lactis, using a nisin-controlled expression system, to produce sweet-tasting lactic acid bacteria. The low level of expression was detected with anti-brazzein antibodies. Secretion of brazzein into the medium has not led to significant yield increase. Surprisingly, optimizing the codon usage for Lactococcus lactis led to a decrease in the yield of brazzein.

Amino Acid Sequence↗

Overexpression and lack of degradation of thaumatin in an aspergillopepsin A-defective mutant of Aspergillus awamori containing an insertion in the pepA gene.

A gene encoding the sweet-tasting protein thaumatin (tha) with optimized codon usage was expressed in Aspergillus awamori. Mutants of A. awamori with reduced proteolytic activity were isolated. One of these mutants, named lpr66, contained an insertion of about 200 bp in the pepA gene, resulting in an inactive aspergillopepsin A. In vitro thaumatin degradation tests confirmed that culture broths of mutant lpr66 showed only a small thaumatin-degrading activity. A. awamori lpr66 has been used as host strain for thaumatin expression cassettes containing the tha gene under the control of either the cahB (cephalosporin acetylhydrolase) promoter of Acremonium chrysogenum or the gdhA (glutamate dehydrogenase) promoter of Aspergillus awamori. Residual proteolytic activities were repressed by using a mixture of glucose and sucrose as carbon sources and L-asparagine as nitrogen source. Degradation of thaumatin by acidic proteases was prevented by maintaining the pH value at 6.2 in the fermentor. Expression of cassettes containing the gdhA promoter was optimal in ammonium sulfate as nitrogen source, whereas transformants expressing the tha gene from the cahB promoter yielded higher thaumatin levels using L-asparagine as nitrogen source. Under optimal fermentation conditions, yields of 105 mg thaumatin/l were obtained, thus making this fermentation a process of industrial interest.

Aspartic Acid Endopeptidases↗

Cloning and expression of the delta 9 fatty acid desaturase gene from Cryptococcus curvatus ATCC 20509 containing histidine boxes and a cytochrome b5 domain.

To allow genetic modification of the fatty acid biosynthesis routes in the lipid-accumulating yeast Cryptococcus curvatus the delta 9 fatty acid desaturase gene was cloned and characterized. The 1668-bp gene encodes a protein of 556 amino acids with a calculated molecular mass of 62 kDa. The gene shows strong homology to previously cloned delta 9 fatty acid desaturase genes from yeast and rat. Homology includes three histidine boxes characteristic for membrane-bound desaturases and a cytochrome b5 domain responsible for electron transport. The delta 9 desaturase gene has a high G+C content of 61% and displays a codon usage different from that of Saccharomyces cerevisiae, but similar to that of the basidiomycete Schizophyllum commune. Expression of the delta 9 desaturase gene of C. curvatus ATCC 20509 was studied in the presence of different fatty acids in the growth medium. Repression of desaturase mRNA signals was found if fatty acids with a double bond at the delta 9 position were present. Fatty acids with a double bond at another position (delta 10 or delta 6) or saturated fatty acids had no effect on the transcription of the cloned gene.

Amino Acid Sequence↗

Identification and initial characterization of a putative Mycoplasma gallinarum leucine aminopeptidase gene.

Aminopeptidases (APN) may play a role in host colonization of M. gallinarum. Characterization of endogenous APN activity suggests that the leucine APN (LAP) of M. gallinarum is a metallo-aminopeptidase activated by Mn2+ and is present in the cytosol and possibly associated with the inner leaflet of the membrane. A 1.36-kb open reading frame (ORF) identified from overlapping genomic phage clones showed 68% nucleotide identity and 51% amino acid identity with the M. salivarium LAP gene. This ORF is expressed as a 1.5-kb monocistronic transcript and is present as a single copy in M. gallinarum. This gene sequence was modified to account for codon usage, and expression in E. coli produced a 51-kDa protein, which compares well with the product predicted from the ORF. This ORF is a strong candidate for contributing the LAP activity of M. gallinarum protein extracts.

Amino Acid Sequence↗