Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

High-level production of yeast (Schwanniomyces occidentalis) phytase in transgenic rice plants by a combination of signal sequence and codon modification of the phytase gene.

This study was designed to produce yeast (Schwanniomyces occidentalis) phytase in rice with a view to future applications in the animal feed industry. To achieve high-level production, chimeric genes with the secretory signal sequence of the rice chitinase-3 gene were constructed using either the original full-length or N-truncated yeast phytase gene, or a modified gene whose codon usage was changed to be more similar to that of rice, and then introduced into rice (Oryza sativa L.). When the original phytase genes were used, the phytase activity in the leaves of transgenic rice was of the same level as in wild-type plants, whose mean value was 0.039 U/g fresh weight (g-FW) (1 U of activity was defined as 1 micromol P released per min at 37 degrees C). In contrast, the enzyme activity was increased markedly when codon-modified phytase genes were introduced: up to 4.6 U/g-FW of leaves for full-length codon-modified phytase, and 10.6 U/g-FW for truncated codon-modified phytase. A decrease in the optimum temperature and thermal stability was observed in the truncated heterologous enzyme, suggesting that the N-terminal region plays an important role in enzymatic properties. In contrast, the optimum temperature and pH of full-length heterologous phytase were indistinguishable from those of the benchmark yeast phytase, although the heterologous enzyme was less glycosylated. Full-length heterologous phytase in leaf extract showed extreme stability. These results indicate that codon modification, combined with the use of a secretory signal sequence, can be used to produce substantial amounts of yeast phytase, and possibly any phytases from various organisms, in an active and stable form.

Journal Article↗

Structure of the Caulobacter crescentus trpFBA operon.

The DNA sequences of the Caulobacter crescentus trpF, trpB, and trpA genes were determined, along with 500 base pairs (bp) of 5'-flanking sequence and 320 bp of 3'-flanking sequence. An open reading frame, designated usg, occurs upstream of trpF and encodes a polypeptide of 89 amino acids which seems to be expressed in a coupled transcription-translation system. Interestingly, the usg polypeptide is not homologous to any known tryptophan biosynthetic enzyme. S1 nuclease mapping of in vivo transcripts indicated that usg, trpF, trpB, and trpA are arranged into a single operon, with the transcription initiation site located 30 bp upstream from the start of usg. Sequences centered at -30 and -6 bp upstream from the transcription initiation site are somewhat homologous to the Escherichia coli promoter consensus sequence and are homologous to sequences found upstream of genes from several organisms which are evolutionarily related to C. crescentus. Furthermore, the trpFBA operon promoter sequence lacks homology to promoter sequences identified for certain developmentally regulated C. crescentus genes. The structures of the C. crescentus usg, trpF, trpB, and trpA genes were further analyzed in terms of codon usage, G+C content, and genetic signals and were related to genetic signals previously identified in C. crescentus and other bacteria. Taken together, these results are relevant to the analysis of gene expression in C. crescentus and the study of trp gene structure and regulation.

Amino Acid Sequence↗

Gene structure, organization, and expression in archaebacteria.

Major advances have recently been made in understanding the molecular biology of the archaebacteria. In this review, we compare the structure of protein and stable RNA-encoding genes cloned and sequenced from each of the major classes of archaebacteria: the methanogens, extreme halophiles, and acid thermophiles. Protein-encoding genes, including some encoding proteins directly involved in methanogenesis and photoautotrophy, are analyzed on the basis of gene organization and structure, transcriptional control signals, codon usage, and evolutionary conservation. Stable RNA-encoding genes are compared for gene organization and structure, transcriptional signals, and processing events involved in RNA maturation, including intron removal. Comparisons of archaebacterial structures and regulatory systems are made with their eubacterial and eukaryotic homologs.

Archaea↗

Conservation of the basic pattern of cellular amino acid composition of archaeobacteria during biological evolution and the putative amino acid composition of primitive life forms.

Previous studies showed that the cellular amino acid composition obtained by amino acid analysis of whole cells, differs such as eubacteria, protozoa, fungi and mammalian cells. These results suggest that the difference in the cellular amino acid composition reflects biological changes as the result of evolution. However, the basic pattern of cellular amino acid composition was relatively constant in all organisms examined. In the present study, we examined archaeobacteria, because they are considered important in understanding the relationship between biological evolution and cellular amino acid composition. The cellular amino acid compositions of Archaeoglobus fulgidus, Pyrococcus horikoshii, Methanobacterium thermoautotrophicum and Methanococcus jannaschii differed slightly from each other, but were similar to those determined from codon usage data, based on the complete genomes. Thus, the cellular amino acid composition reflects biological evolution. We suggest that primitive forms of life appearing on earth at the end of prebiotic evolution had a similar-cellular amino acid composition.

Amino Acids↗

Sequencing of coronavirus IBV genomic RNA: three open reading frames in the 5' 'unique' region of mRNA D.

The nucleotide sequence of a genomic cDNA clone corresponding to the 5' terminal domain of mRNA D of the Beaudette strain of infectious bronchitis virus (IBV) has been determined. This region contains three open reading frames which predict polypeptides of molecular weights 6700 (6.7K), 7.4K and 12.4K. The predicted 12.4K polypeptide has a codon usage very similar to that predicted for the products of the IBV nucleocapsid, membrane and spike genes. The sequence also predicts a hydrophobic, potentially membrane-anchoring, region in the N terminal half of the 12.4K polypeptide, and a hydrophilic C terminus.

Amino Acid Sequence↗

Nucleotide sequence of the 26 S mRNA of the virulent Trinidad donkey strain of Venezuelan equine encephalitis virus and deduced sequence of the encoded structural proteins.

A cDNA clone containing all of the 26 S mRNA coding region of the RNA genome of Venezuelan equine encephalitis (VEE) virus, virulent strain Trinidad donkey (TRD), has been constructed and sequenced. The nucleotide and deduced amino acid sequences of the 26 S RNA of VEE virus conform to the general organization of the alphavirus subgenomic mRNA. Excluding the poly(A) tail, the VEE 26 S RNA is 3913 nucleotides long with a protein coding region of 3762 nucleotides. Codon usage in the translated region is nonrandom and correlates well with that reported for Sindbis (SIN), Semliki Forest (SF), and Ross River (RR) alphaviruses. Highly conserved sequences of 19 to 22 nucleotides representing putative replicase recognition sites occur at the 26 S RNA junction region of the 42 S genomic RNA and at the 3' terminus immediately preceding the poly(A) tail. The conserved sequence at the 26 S/42 S junction region of VEE virus differs from that of other alphaviruses in that an ochre termination codon (UAA) is substituted for a GGU (Gly) codon present in the other viruses. The 5' and 3' noncoding regions (30 and 121 nucleotides, respectively) of the VEE 26 S RNA are shorter than has been reported for several other alphaviruses. The approximate transmembrane domains of the VEE E1 and E2 envelope glycoproteins have been identified. VEE E1 contains a single asparagine-linked glycosylation site, whereas E2 has three such sites, all of which are apparently glycosylated. The deduced amino acid sequence of the VEE polyprotein shows an overall homology of 44 to 46% with the precursor polyproteins of SIN, SF, and RR viruses. VEE virus capsid, E1, and E2 structural proteins show 43 to 46%, 50 to 53%, and 36 to 41% homology, respectively, with the cognate proteins of SIN, SF, and RR viruses.

Amino Acid Sequence↗

Nucleotide sequence and deduced amino acid sequence of Escherichia coli pyruvate oxidase, a lipid-activated flavoprotein.

The entire nucleotide sequence of the poxB (pyruvate oxidase) gene of Escherichia coli K-12 has been determined by the dideoxynucleotide (Sanger) sequencing of fragments of the gene cloned into a phage M13 vector. The gene is 1716 nucleotides in length and has an open reading frame which encodes a protein of Mr 62,018. This open reading frame was shown to encode pyruvate oxidase by alignment of the amino acid sequences deduced for the amino and carboxy termini and several internal segments of the mature protein with sequences obtained by amino acid sequence analysis. The deduced amino acid sequence of the oxidase was not unusually rich in hydrophobic sequences despite the peripheral membrane location and lipid binding properties of the protein. The codon usage of the oxidase gene was typical of a moderately expressed protein. The deduced amino acid sequence shares homology with the large subunits of the acetohydroxy acid synthase isozymes I, II, and III, encoded by the ilvB, ilvG, and ilvI genes of E. coli.

Acetolactate Synthase↗

Complete mitochondrial genome sequences for Crown-of-thorns starfish Acanthaster planci and Acanthaster brevispinus.

BACKGROUND: The crown-of-thorns starfish, Acanthaster planci (L.), has been blamed for coral mortality in a large number of coral reef systems situated in the Indo-Pacific region. Because of its high fecundity and the long duration of the pelagic larval stage, the mechanism of outbreaks may be related to its meta-population dynamics, which should be examined by larval sampling and population genetic analysis. However, A. planci larvae have undistinguished morphological features compared with other asteroid larvae, hence it has been difficult to discriminate A. planci larvae in plankton samples without species-specific markers. Also, no tools are available to reveal the dispersal pathway of A. planci larvae. Therefore the development of highly polymorphic genetic markers has the potential to overcome these difficulties. To obtain genomic information for these purposes, the complete nucleotide sequences of the mitochondrial genome of A. planci and its putative sibling species, A. brevispinus were determined and their characteristics discussed. RESULTS: The complete mtDNA of A. planci and A. brevispinus are 16,234 bp and 16,254 bp in size, respectively. These values fall within the length variation range reported for other metazoan mitochondrial genomes. They contain 13 proteins, 2 rRNA, and 22 tRNA genes and the putative control region in the same order as the asteroid, Asterina pectinifera. The A + T contents of A. planci and A. brevispinus on their L strands that encode the majority of protein-coding genes are 56.3% and 56.4% respectively and are lower than that of A. pectinifera (61.2%). The percent similarity of nucleotide sequences between A. planci and A. brevispinus is found to be highest in the CO2 and CO3 regions (both 90.6%) and lowest in ND2 gene (84.2%) among the 13 protein-coding genes. In the deduced putative amino acid sequences, CO1 is highly conserved (99.2%), and ATP8 apparently evolves faster any of the other protein-coding gene (85.2%). CONCLUSION: The gene arrangement, base composition, codon usage and tRNA structure of A. planci are similar to those of A. brevispinus. However, there are significant variations between A. planci and A. brevispinus. Complete mtDNA sequences are useful for the study of phylogeny, larval detection and population genetics.

Animals↗

A genomic basis for the evolution of vertebrate transcription factors containing amino Acid runs.

We have previously shown that polyAla (A) tract-containing proteins frequently present runs of glycine (G), proline (P), and histidine (H) and that, in their ORFs, GC content at all codon positions is higher than that in the rest of the genome. In this study, we present new analyses of these human proteins/ORFs. We detected striking differences in codon usage for A, G, and P in and out of runs. After dividing the ORFs, we found that 5' halves were richer in runs than 3' halves. Afterward, when removing the runs, we observed that the run-rich halves (grouped irrespectively of their 5' or 3' position) had a marked statistical tendency to have more homo- and hetero-dicodons for A, G, P, and H than the run-poor halves. This suggests that, in addition to the necessary GC-rich genomic background, a specific codon organization is probably required to generate these coding repeats. Homo-dicodons may indeed provide primers for run formation through polymerase slippage. The compositional analysis of human HOX genes, the most polyAla-rich family, and their comparison with their zebrafish homologs, support these hypotheses and suggest possible effects of genomic environment on ORF evolution and organismal diversification.

Amino Acid Sequence↗

Aminoacylation of tRNAs encoded by Chlorella virus CVK2.

Viruses that infect certain strains of the unicellular green alga, Chlorella, have a large, linear dsDNA genome that is 330-380 kb in size; this genomic size is the largest known among viruses and is equivalent to approximately 60% of the smallest prokaryotic genome of Mycoplasma genitalium (580 kb). Besides many putative protein-coding genes, a cluster of 10-15 tRNA genes is present in these viral genomes. Some of these tRNA genes contain peculiar insertions. In infected host cells, the viral tRNAs of CVK2, a Chlorella virus isolate, have been demonstrated to be cotranscribed as a large precursor, approximately 1.0 kb in size, that is precisely processed into individual mature tRNA species. Acidic Northern blot analysis of eight of these tRNAs has revealed that they are actually aminoacylated in vivo, indicating their involvement in viral protein synthesis. They may help the virus reach maximal replication potential by overcoming codon usage barriers that exist between the virus and its host. These results provide evidence that some components of the host protein synthesis machinery can be replaced by viral gene products. This is the first report of tRNA aminoacylation encoded by viruses of eukaryotes.

Acylation↗

Nucleotide sequence of the gene coding for yeast cytoplasmic aspartyl-tRNA synthetase (APS); mapping of the 5' and 3' termini of AspRS mRNA.

A 3.8 Kb DNA fragment, which contains the structural gene of aspartyl-tRNA synthetase (AspRS) and its flanking regions, has been fully sequenced by the combined M13/dideoxy chain terminator method. From the single open reading frame of correct length (1671 bp) we deduced an amino acid sequence consistent with that of several peptides of AspRS. No significant internal sequence repeats were observed in the primary structure of the protein. The AspRS gene (APS) has a codon usage pattern typical of non abundant proteins. S1 nuclease analysis of APS mRNA showed a major start 17 bases downstream from a "TATA box" and stops near an RNA polymerase terminator sequence.

Amino Acid Sequence↗

The complete DNA sequence and analysis of the large virulence plasmid of Escherichia coli O157:H7.

The complete DNA sequence of pO157, the large virulence plasmid of EHEC strain O157:H7 EDL 933, is presented. The 92 kb F-like plasmid is composed of segments of putative virulence genes in a framework of replication and maintenance regions, with seven insertion sequence elements, located mostly at the boundaries of the virulence segments. One hundred open reading frames (ORFs) were identified, of which 19 were previously sequenced potential virulence genes. Forty-two ORFs were sufficiently similar to known proteins for suggested functions to be assigned, and 22 had no convincing similarity with any known proteins. Of the newly identified genes, an unusually large ORF of 3169 amino acids has a putative cytotoxin active site shared with the large clostridial toxin (LCT) family and proteins such as ToxA and B of Clostridium difficile . A conserved motif was detected that links the large ORF and the LCT proteins with the OCH1 family of glycosyltransferases. In the complete sequence, the mosaic form can be observed at the levels of base composition, codon usage and gene organization. Insights were obtained from patterns of DNA composition as well as the pathogenic and 'housekeeping' gene segments. Evolutionary trees built from shared plasmid maintenance genes show that even these genes have heterogeneous origins.

Amino Acid Sequence↗

A Rev-independent human immunodeficiency virus type 1 (HIV-1)-based vector that exploits a codon-optimized HIV-1 gag-pol gene.

The human immunodeficiency virus (HIV) genome is AU rich, and this imparts a codon bias that is quite different from the one used by human genes. The codon usage is particularly marked for the gag, pol, and env genes. Interestingly, the expression of these genes is dependent on the presence of the Rev/Rev-responsive element (RRE) regulatory system, even in contexts other than the HIV genome. The Rev dependency has been explained in part by the presence of RNA instability sequences residing in these coding regions. The requirement for Rev also places a limitation on the development of HIV-based vectors, because of the requirement to provide an accessory factor. We have now synthesized a complete codon-optimized HIV-1 gag-pol gene. We show that expression levels are high and that expression is Rev independent. This effect is due to an increase in the amount of gag-pol mRNA. Provision of the RRE in cis did not lower protein or RNA levels or stimulate a Rev response. Furthermore we have used this synthetic gag-pol gene to produce HIV vectors that now lack all of the accessory proteins. These vectors should now be safer than murine leukemia virus-based vectors.

Base Sequence↗

Cloning and characterization of the gene (rfc) encoding O-antigen polymerase of Pseudomonas aeruginosa PAO1.

The lipopolysaccharide (LPS) O-antigen polymerase is the product of the rfc gene. Loss of O-antigen polymerase activity due to mutation in rfc gives rise to a characteristic LPS phenotype known as core-plus-one or semi-rough, wherein the LPS core is capped with a single oligosaccharide unit. Pseudomonas aeruginosa (Pa) AK1401, a derivative of strain PAO1 (serogroup O5), expresses a semi-rough LPS; this mutant phenotype was complemented by a 2.2-kb NsiI-SacI fragment of Pa PAO1 DNA. Sequence analysis of this fragment revealed a 1317-bp open reading frame (ORF) potentially encoding a 438-amino-acid (aa) protein of 48,849 Da. This DNA sequence and the inferred aa sequence contain many of the features of other O-antigen polymerases, including an aberrantly low G + C content (particularly apparent in the high-G + C background of Pa), an unusual codon usage pattern, and a hydrophobicity profile indicative of a membrane protein. A 345-bp fragment internal to the ORF hybridized to genomic DNA from two of ten Pa serogroup strains examined by Southern blot; these two strains express O antigens structurally related to that of strain PAO1.

Base Composition↗

Mitochondrial cytochrome C oxidase subunit I of Manduca sexta and a comparison with other invertebrate genes.

A cDNA encoding mitochondrial cytochrome c oxidase subunit I (mt COI) from Manduca sexta (Lepidoptera: Sphingidae) was cloned and sequenced. AT (adenine-thymine) content is high and codon usage is biased and likely reflects the role of mt COI in electron transport. The encoded protein is 514 amino acids long, contains seven invariant His residues observed in COIs in all organisms and would be predicted to be composed of 12 transmembrane regions.

Amino Acid Sequence↗

mRNA secondary structure at start AUG codon is a key limiting factor for human protein expression in Escherichia coli.

Codon usage and thermodynamic optimization of the 5'-end of mRNA have been applied to improve the efficiency of human protein production in Escherichia coli. However, high level expression of human protein in E. coli is still a challenge that virtually depends upon each individual target genes. Using human interleukin 10 (huIL-10) and interferon alpha (huIFN-alpha) coding sequences, we systematically analyzed the influence of several major factors on expression of human protein in E. coli. The results from huIL-10 and reinforced by huIFN-alpha showed that exposing AUG initiator codon from base-paired structure within mRNA itself significantly improved the translation of target protein, which resulted in a 10-fold higher protein expression than the wild-type genes. It was also noted that translation process was not affected by the retained short-range stem-loop structure at Shine-Dalgarno (SD) sequences. On the other hand, codon-optimized constructs of huIL-10 showed unimproved levels of protein expression, on the contrary, led to a remarkable RNA degradation. Our study demonstrates that exposure of AUG initiator codon from long-range intra-strand secondary structure at 5'-end of mRNA may be used as a general strategy for human protein production in E. coli.

Base Sequence↗

D-Xylose (D-glucose) isomerase from Arthrobacter strain N.R.R.L. B3728. Gene cloning, sequence and expression.

Arthrobacter strain N.R.R.L. B3728 superproduces a D-xylose isomerase that is also a useful industrial D-glucose isomerase. The gene (xylA) that encodes it has been cloned by complementing a xylA mutant of the ancestral strain, with the use of a shuttle vector. The 5' region shows strong sequence similarity to Escherichia coli consensus promoters and ribosome-binding sequences and allows high levels of expression in E. coli. The coding sequence shows similarity to those for other D-xylose isomerases and is followed by 22 nucleotide residues with stop codons in each reading frame, a good 'consensus' ribosome-binding site and an open reading frame showing similarity to those of known D-xylulokinases (xylB). Studies on the expression of the cloned gene in Arthrobacter and in E. coli suggest that the two genes are part of a xyl operon regulated by a repressor that is defective in strain B3728. Codon usage in these two genes, and in another open reading frame (nxi) that was adventitiously isolated during early cloning attempts, shows some characteristic omissions and a strong G + C preference in redundant positions.

Aldose-Ketose Isomerases↗

A convenient and adaptable package of DNA sequence analysis programs for microcomputers.

We describe a package of DNA data handling and analysis programs designed for microcomputers. The package is convenient for immediate use by persons with little or no computer experience, and has been optimized by trial in our group for a year. By typing a single command, the user enters a system which asks questions or gives instructions in English. The system will enter, alter, and manage sequence files or a restriction enzyme library. It generates the reverse complement, translates, calculates codon usage, finds restriction sites, finds homologies with various degrees of mismatch, and graphs amino acid composition or base frequencies. A number of options for data handling and printing can be used to produce figures for publication. The package will be available in ANSI Standard FORTRAN for use with virtually any FORTRAN compiler.

Amino Acid Sequence↗