Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,027 records · Page 57Linked to original sources

Expression of a foreign gene in Chlamydomonas reinhardtii.

Genomic transformation of Chlamydomonas reinhardtii exposed to glass-bead abrasion was accomplished with a chimeric neomycin phosphotransferaseII (NPTII)-encoding gene (nos::npt) flanked by the nopaline synthase promoter and polyadenylation sequences obtained from the Ti plasmid of Agrobacterium tumefaciens. These sequences were in a plasmid (pGA482) which also contained gene nit1 encoding nitrate reductase of C. reinhardtii. Transformants were selected by their ability to grow on medium containing nitrate, and 52% of these was also resistant to kanamycin. Evidence for nos::npt expression includes: (1) hybridization with probes specific for npt, (2) demonstration of NPTII activity after electrophoresis of extracts, and (3) chromatographic identification of the reaction product of NPTII, kanamycin phosphate. The highly biased codon usage in Chlamydomonas does not preclude expression.

Agrobacterium tumefaciens↗

Nucleotide sequence of cDNA encoding the small subunit of ribulose-1,5-bisphosphate carboxylase from maize.

We have cloned a full length cDNA for the small subunit of ribulose-1,5-bisphosphate carboxylase from C4 monocot maize, determined the complete nucleotide sequence of this cDNA and deduced its amino acid sequence. The cDNA insert included 513 bp of the coding region, and 65 and 252 nucleotides of the 5' and 3' untranslated regions, respectively. The transit and mature peptides have, respectively, 47 and 123 amino acids. Comparison with the small subunit genes from other plants revealed that the maize small subunit is similar to the wheat one, there being 73% homology between the transit peptides and 64% between the mature proteins. This indicates that there is no noteworthy difference between the C3 and C4 small subunit structures. Extreme codon bias was observed for this gene, and similar codon preferences are observed for other proteins highly expressed in maize leaf, light harvesting chlorophyll binding protein and phosphoenolpyruvate carboxylase. The results indicate that preferential codon usage for highly expressed genes occurs in maize leaf.

Amino Acid Sequence↗

Sequence, transcription and translation of a late gene of the Autographa californica nuclear polyhedrosis virus encoding a 34.8K polypeptide.

A 1.4 kb region downstream of the DNA polymerase gene of Autographa californica nuclear polyhedrosis virus was sequenced. Two open reading frames (ORFs) were identified of 927 and 474 bases in length. The 927 base ORF encodes a 34.8K protein as determined by in vitro translation of both hybrid-selected RNA and RNA synthesized in vitro from a 927 base ORF template. The predicted amino acid sequence of the 34.8K polypeptide (p34.8) reveals a hydrophobic N terminus, two potential N-glycosylation sites, and potential sites for phosphorylation by casein kinase I and protein kinase C. The p34.8 gene has a strong codon usage bias which is strikingly different from that of the polyhedrin gene. The two 5' ends of the 927 base ORF transcripts initiate from an ATAAG sequence and a GTAAG sequence 11 and 87 bases upstream of the ATG codon respectively. A short upstream reading frame is present in the leader sequence of the longer RNA. The transcripts have multiple 3' ends; the most proximal endpoint correlates with a polyadenylation signal overlapping the translational termination codon of the 927 base ORF. Transcripts of the latter were not observed early in the infection cycle but appeared 6 h after infection and were maximally expressed at 12 to 24 h post-infection. The late nature of these transcripts was confirmed by their sensitivity to aphidicolin and cycloheximide, inhibitors of DNA replication and protein synthesis respectively. Attempts to construct viral mutants carrying a deletion of the p34.8 gene and fusion with the beta-galactosidase gene suggest that the former gene is essential for viral replication.

Amino Acid Sequence↗

The mitochondrial genome of the mosquito Anopheles gambiae: DNA sequence, genome organization, and comparisons with mitochondrial sequences of other insects.

The entire 15,363 bp mitochondrial genome was cloned and sequenced from the mosquito Anopheles gambiae. With respect to the protein-coding genes, rRNA genes and the control region, the gene order was identical to that reported for other insects. There were significant differences, however, in the position and orientation of specific tRNA loci. The overall nucleotide composition was heavily biased towards adenine and thymine, which accounted for 77.6% of all nucleotides. Comparisons were made with the mitochondrial genomes of other insects on the basis genome size and organization, DNA and putative amino acid sequence data, nucleotide substitutions, codon usage and bias, and patterns of AT enrichment.

Amino Acid Sequence↗

Evidence for a coding pattern on the non-coding strand of the E. coli genome.

Analysis of codon usage frequency for the combined coding sequences of 52 E. coli genes, taken from the European Molecular Biology Laboratory Nucleotide Sequence Data Library, Release 2, shows that there is a significant positive correlation between the frequency with which a given codon appears on the coding strand and the frequency with which it appears, in phase, on the non-coding strand.

Amino Acid Sequence↗

The intein of the Thermoplasma A-ATPase A subunit: structure, evolution and expression in E. coli.

BACKGROUND: Inteins are selfish genetic elements that excise themselves from the host protein during post translational processing, and religate the host protein with a peptide bond. In addition to this splicing activity, most reported inteins also contain an endonuclease domain that is important in intein propagation. RESULTS: The gene encoding the Thermoplasma acidophilum A-ATPase catalytic subunit A is the only one in the entire T. acidophilum genome that has been identified to contain an intein. This intein is inserted in the same position as the inteins found in the ATPase A-subunits encoding gene in Pyrococcus abyssi, P. furiosus and P. horikoshii and is found 20 amino acids upstream of the intein in the homologous vma-1 gene in Saccharomyces cerevisiae. In contrast to the other inteins in catalytic ATPase subunits, the T. acidophilum intein does not contain an endonuclease domain.T. acidophilum has different codon usage frequencies as compared to Escherichia coli. Initially, the low abundance of rare tRNAs prevented expression of the T. acidophilum A-ATPase A subunit in E. coli. Using a strain of E. coli that expresses additional tRNAs for rare codons, the T. acidophilum A-ATPase A subunit was successfully expressed in E. coli. CONCLUSIONS: Despite differences in pH and temperature between the E. coli and the T. acidophilum cytoplasms, the T. acidophilum intein retains efficient self-splicing activity when expressed in E. coli. The small intein in the Thermoplasma A-ATPase is closely related to the endonuclease containing intein in the Pyrococcus A-ATPase. Phylogenetic analyses suggest that this intein was horizontally transferred between Pyrococcus and Thermoplasma, and that the small intein has persisted in Thermoplasma apparently without homing.

Adenosine Triphosphatases↗

A new putative gene in the mitochondrial genome of Saccharomyces cerevisiae.

The 2200-bp ori2-ori7 region of the mitochondrial (mt) genome of Saccharomyces cerevisiae has been sequenced on the genome of a petite, b7, excised at those ori sequences from wild-type strain B. The region contains an open reading frame, ORF5, which is transcribed into a 900-nucleotide (nt) RNA in both the parental wild-type strain and its derived petite, b7. This RNA uses as a template the strand used by most mt transcripts. Its start point is located 337 nt upstream of ORF5; and a messenger termination site has been found 900 nt downstream of the initiation site. These data suggest that ORF5 is a new mitochondrial gene. The G + C content of ORF5 is only 15.7%; 90% of the G + C base pairs of ORF5 are comprised in a palindromic G + C cluster similar to that present in the varl gene. The coding capacity of ORF5 is 46 amino acids (aa), mainly represented by methionine, phenylalanine, arginine, valine, asparagine, isoleucine and tyrosine. The aa composition and the codon usage of ORF5 are reminiscent of those of varl and of other intergenic ORFs.

Amino Acid Sequence↗

Nucleotide sequence of the invasion plasmid antigen B and C genes (ipaB and ipaC) of Shigella flexneri.

The nucleotide sequence of a 4.8 kilobase (kb) HindIII fragment from pWR100, the virulence plasmid of Shigella flexneri 5, was determined and analysed. This fragment encodes polypeptides b (62 kilodalton, kD) and c (43 kD) which have already been described as two of the four immunogenic polypeptides of Shigellae. The nucleotide sequence revealed that in addition to the ipaB and ipaC genes encoding polypeptides b and c, a third complete open reading frame was found within the fragment. The gene, named ippI, encoded a 17 kD polypeptide. The deduced amino acids sequence of polypeptides b and c showed no signal peptide but presence of highly hydrophobic domains compatible with a transmembraneous location. The surprising A and T richness of the three genes as compared with the Escherichia coli and Shigella genomes, resulted in a biased codon usage, and raises the question of the origin of the sequences.

Amino Acid Sequence↗

A muscle-specific actin gene from the Mediterranean fruit fly, Ceratitis capitata.

A characterization of an actin gene isolated from the genome of the Mediterranean fruit fly, Ceratitis capitata, including the complete sequencing of the coding, 3' and 5' flanking regions of this gene and a partial cDNA was carried out. The partial cDNA was derived from the 3' untranslated region of the actin gene described here, and has been used to identify this gene uniquely. The DNA sequence data presented here, together with the pattern of expression exhibited by this gene during development, strongly support the interpretation that this is a muscle-specific actin gene. Peaks of expression are seen in tissues and during temporal phases of development where muscle differentiation is occurring. The derived protein sequence of the Medfly acting gene shows the highest degrees of similarity, 98.4 and 96.6% respectively, with the two muscle-specific actin genes 79B and 88F from D. melanogaster. The Medfly actin gene also has a single intervening sequence, and an intron is found at the same position in the 79B and 88F actin genes. In the coding region at the DNA level, 17.2 and 16.4% nucleotide differences, respectively, are observed between the Medfly actin gene and these same two D. melanogaster actin genes. The disparity between the amino acid and nucleotide comparisons can be explained, in part, by a high level of synonymous changes in the DNA sequence. In addition, despite the many similarities, codon usage appears to be very different between the actin genes of these species.

Actins↗

The structure of a plasmid of Chlamydia trachomatis believed to be required for growth within mammalian cells.

Sequence analysis of a 7.5 kb DNA plasmid isolated from Chlamydia trachomatis shows 8 open reading frames (ORFs) regularly spaced along most of the sequence. One of these ORFs encodes a 451-amino-acid polypeptide highly homologous to the DnaB protein of Escherichia coli. A region between ORFs 6 and 7 contains a cluster of alternating ATs and a 22 bp sequence tandemly repeated 4 times, suggesting a replication control region. Several ORFs correspond to plasmid-specific polypeptides that have been described. Codons ending with A or T are more frequent, as might be expected from the high A/T content (64%) of the plasmid, and codon usage is similar to that of the C. trachomatis chromosomal gene, omp1L2.

Amino Acid Sequence↗

Redesign, synthesis and functional expression of the 6-deoxyerythronolide B polyketide synthase gene cluster.

A generic design of Type I polyketide synthase genes has been reported in which modules, and domains within modules, are flanked by sets of unique restriction sites that are repeated in every module [1]. Using the universal design, we synthesized the six-module DEBS gene cluster optimized for codon usage in E. coli, and cloned the three open reading frames into three compatible expression vectors. With one correctable exception, the amino acid substitutions required for restriction site placements were compatible with polyketide production. When expressed in E. coli the codon-optimized synthetic gene cluster produced significantly more protein than did the wild-type sequence. Indeed, for optimal polyketide production, PKS expression had to be down-regulated by promoter attenuation to achieve balance with expression of the accessory proteins needed to support polyketide biosynthesis.

Escherichia coli↗

Codon optimization markedly improves doxycycline regulated gene expression in the mouse heart.

Tetracycline regulated gene expression in transgenic animals is potentially a very powerful technique (Furth et al., 1994; Gossen & Bujard 1992). We have utilized this system in an attempt to overcome the perinatal lethality resulting from constitutive transgenic expression in the heart (Valencik & McDonald, Am J Physiol Heart Circ Physiol 280: H361-H367). We found that compound hemizygous animals created by mating selected reverse tetracycline transactivator (rtTA) and transresponder (TR) lines display tightly regulated TR expression in the heart. However, we identified two fundamental problems. First, codon usage bias appeared to severely limit the expression of the rtTA driven by the cardiac alpha-myosin heavy chain promoter. Second, co-injection of rtTA and TR transgenes led to compound hemizygous animals that exhibited unregulated TR gene expression. Codon optimization of the rtTA construct leads to marked improvement (increasing the average induction from 20-fold to 832-fold) in cardiac myocyte expression. The resulting opt-rtTA lines can be bred to homozygosity, facilitating rapid screening of F0 TR animals for doxycycline regulated transgene expression.

Animals↗

Characterization of a halobacterial gene affecting bacterio-opsin gene expression.

A substantial number of spontaneous bacterio-opsin mutants of Halobacterium halobium are the result of insertion elements up to 1400 bp upstream of the bacterio-opsin (bop) gene. The nucleotide sequence of 1800 bp upstream of the bop gene has been determined. There is a 1118 bp open reading frame (ORF) located within this region which is transcribed and which coincides with the distribution of insertion elements upstream of the bop gene in Bop mutants. Therefore, we propose that there is a gene (brp gene) 526 bp upstream of the bop gene. This putative gene is transcribed in the opposite direction as the bop gene and could encode a protein of 37,500 D (359 amino acids) with a codon usage similar to bacterio-opsin. The 5' terminus of the brp transcript has been determined. The brp transcript and the bop mRNA are complementary for 13 residues near their 5' termini and both transcripts start at or near the initiating codon of the gene. Both transcripts could form similar hairpin loop structures at their 5' termini which contain possible ribosomal binding sites. The DNA sequences immediately upstream of the bop and the brp genes have significant homologies and there is a short complementary sequence. The role of the brp gene in bacterio-opsin gene expression is unclear.

Amino Acid Sequence↗

Cloning and sequencing of the LEU2 homologue gene of Schwanniomyces occidentalis.

A gene that complements the leu2 mutation of Saccharomyces cerevisiae has been cloned from Schwanniomyces occidentalis. The gene codes for a protein of 379 amino acids. As expected for a Schwanniomyces gene, it has a high AT content, which is also reflected in the codon usage. The sequence homology with other known leu2 complementing genes is low.

3-Isopropylmalate Dehydrogenase↗

Nucleotide composition of genes and hydrophobicity of the encoded proteins.

We find that true proteins are generally more hydrophobic than the corresponding hypothetical proteins encoded by the randomized gene nucleotide sequences. Furthermore, the protein hydrophobicity but not its gene nucleotide composition is conserved within evolutionary families of functionally related proteins. These two findings indicate that there is a general drift to modify gene nucleotide composition in the course of evolution. An inspection of codon usage in genes shows that the drift mainly increases the content of adenine at the expense of thymine.

Adenosine↗

Evolutionary rate variation in eukaryotic lineage specific human intronless proteins.

The present study examines 783 human-mouse orthologous gene pairs for their pattern of sequence evolution, contrasting mammalia, eukaryota, coelomata, and bilateria specific human intronless genes. Such comparisons may be of use in understanding the general evolution of human genome. Evolutionary rate analyses indicate that mammalia specific human intronless genes are evolving faster as compared to other intronless genes specific to eukaryotic lineage, indicating towards their rapid evolution. The observations indicates that the genes conserved in eukaryota, coelomata, and bilateria, that is, proteins that arose earlier in evolution as compared to mammalia specific genes evolve slowly and are subjected to negative selection. The cause underlying rate variations was also explored. Although mutational bias might slightly fasten the nonsynonymous rates in mammalia specific genes, it is unlikely to be major cause of rate difference between the various categories. Furthermore, rate of divergence of mammalia specific intronless genes has been related to functional classification using the protein family annotation. Protein function was found in some cases to have larger impact on the rate of evolution of genes. Also, the codon usage pattern of mammalia specific intronless genes do not seem to differ much from those of other intronless genes conserved solely in eukaryotic lineage.

Animals↗

Silent mutations affect in vivo protein folding in Escherichia coli.

As an approach to investigate the molecular mechanism of in vivo protein folding and the role of translation kinetics on specific folding pathways, we made codon substitutions in the EgFABP1 (Echinococcus granulosus fatty acid binding protein1) gene that replaced five minor codons with their synonymous major ones. The altered region corresponds to a turn between two short alpha helices. One of the silent mutations of EgFABP1 markedly decreased the solubility of the protein when expressed in Escherichia coli. Expression of this protein also caused strong activation of a reporter gene designed to detect misfolded proteins, suggesting that the turn region seems to have special translation kinetic requirements that ensure proper folding of the protein. Our results highlight the importance of codon usage in the in vivo protein folding.

Amino Acid Substitution↗

The amber codon in the gene encoding the monomethylamine methyltransferase isolated from Methanosarcina barkeri is translated as a sense codon.

Each of the genes encoding the methyltransferases initiating methanogenesis from trimethylamine, dimethylamine, or monomethylamine by various Methanosarcina species possesses one naturally occurring in-frame amber codon that does not appear to act as a translation stop during synthesis of the biochemically characterized methyltransferase. To investigate the means by which suppression of the amber codon within these genes occurs, MtmB, a methyltransferase initiating metabolism of monomethylamine, was examined. The C-terminal sequence of MtmB indicated that synthesis of this mtmB1 gene product did not cease at the internal amber codon, but at the following ochre codon. Antibody raised against MtmB revealed that Escherichia coli transformed with mtmB1 produced the amber termination product. The same antibody detected primarily a 50-kDa protein in Methanosarcina barkeri, which is the mass predicted for the amber readthrough product of the mtmB1 gene. Sequencing of peptide fragments from MtmB by Edman degradation and mass spectrometry revealed no change in the reading frame during mtmB1 expression. The amber codon position corresponded to a lysyl residue using either sequencing technique. The amber codon is thus read through during translation at apparently high efficiency and corresponds to lysine in tryptic fragments of MtmB even though canonical lysine codon usage is encountered in other Methanosarcina genes.

Amino Acid Sequence↗