Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Nucleotide sequence of the Rickettsia prowazekii citrate synthase gene.

The Rickettsia prowazekii citrate synthase (gltA) gene, previously cloned in Escherichia coli, was localized to a 2.0-kilobase chromosomal fragment. DNA sequence analysis of a portion of this fragment revealed an open reading frame of 1,308 base pairs that encodes a protein of 435 amino acids with a molecular weight of 49,171. This translation product is comparable in size to both the E. coli and pig heart citrate synthase monomers and to the protein synthesized in E. coli minicells containing the rickettsial gene. Comparisons between the deduced amino acid sequence of R. prowazekii citrate synthase and those of the E. coli and pig heart enzymes revealed extensive homology (59%) between the two bacterial proteins. In contrast, only 20% of the rickettsial enzyme residues were shared with the functionally similar pig heart enzyme residues. Upstream from the open reading frame and in close proximity to one another, sequences with homology to E. coli consensus sequences for RNA polymerase and ribosome binding were identified. S1 nuclease mapping experiments demonstrated that the start of transcription for this gene in E. coli was located in the upstream region. Codon usage in the rickettsial gltA gene was found to be very biased and differed from the pattern observed in E. coli. Adenine and uracil were used preferentially in the third base position of rickettsial codons.

Amino Acid Sequence↗

Strategies for achieving high-level expression of genes in Escherichia coli.

Progress in our understanding of several biological processes promises to broaden the usefulness of Escherichia coli as a tool for gene expression. There is an expanding choice of tightly regulated prokaryotic promoters suitable for achieving high-level gene expression. New host strains facilitate the formation of disulfide bonds in the reducing environment of the cytoplasm and offer higher protein yields by minimizing proteolytic degradation. Insights into the process of protein translocation across the bacterial membranes may eventually make it possible to achieve robust secretion of specific proteins into the culture medium. Studies involving molecular chaperones have shown that in specific cases, chaperones can be very effective for improved protein folding, solubility, and membrane transport. Negative results derived from such studies are also instructive in formulating different strategies. The remarkable increase in the availability of fusion partners offers a wide range of tools for improved protein folding, solubility, protection from proteases, yield, and secretion into the culture medium, as well as for detection and purification of recombinant proteins. Codon usage is known to present a potential impediment to high-level gene expression in E. coli. Although we still do not understand all the rules governing this phenomenon, it is apparent that "rare" codons, depending on their frequency and context, can have an adverse effect on protein levels. Usually, this problem can be alleviated by modification of the relevant codons or by coexpression of the cognate tRNA genes. Finally, the elucidation of specific determinants of protein degradation, a plethora of protease-deficient host strains, and methods to stabilize proteins afford new strategies to minimize proteolytic susceptibility of recombinant proteins in E. coli.

Biotechnology↗

The adhesive protein cDNA of Mytilus galloprovincialis encodes decapeptide repeats but no hexapeptide motif.

A mussel is attached to hard surfaces by its byssus, which consists of a bundle of threads, each with a fibrous collagenous core coated with adhesive proteins. We constructed a cDNA library from RNA isolated from the foot of the mussel Mytilus galloprovincialis sampled in Japan. The library was probed with a nucleotide sequence corresponding to a part of the decapeptide repeat motif in the major adhesive protein of the closely related species M. edulis, and a clone including the whole coding region of the same adhesive protein of M. galloprovincialis was isolated. The sequences of the signal and nonrepetitive regions of the protein of M. galloprovincialis were homologous to those of M. edulis, despite several substitutions and a deletion of 18 amino acids. The repetitive region included a tetradecapeptide sequence and 62 repeats of the same decapeptide motif as in M. edulis, but hexapeptide sequences present in M. edulis were absent in the protein of M. galloprovincialis. In the decapeptide motif, two tyrosine residues, two lysine residues, and one of the two proline residues were highly conserved, but other residues were frequently substituted. In some residues in the decapeptide motif, specific codon usages were observed, suggesting that the nucleotide sequence itself has a function.

Amino Acid Sequence↗

Clustered bottlenecks in mRNA translation and protein synthesis.

Using a model based on the totally asymmetric exclusion process, we investigate the effects of slow codons along messenger RNA. Ribosome density profiles near neighboring clusters of slow codons interact, enhancing suppression of ribosome throughput when such bottlenecks are closely spaced. Increasing the slow codon cluster size beyond approximately 3-4 codons does not significantly reduce the ribosome current. Our results are verified by both extensive Monte Carlo simulations and numerical calculation, and provide a biologically motivated explanation for the experimentally observed clustering of low-usage codons.

Algorithms↗

Determinants of translational initiation efficiency in the atp operon of Escherichia coli.

Transcription and translation of the atp genes encoding the subunits b, delta, alpha, gamma and epsilon of the Escherichia coli H+-ATPase were studied. The nature and quantities of the respective transcripts initiated from different promoters were compared with overall expression rates thus yielding accurate information about relative translational efficiency and its coupling to mRNA levels. Part of the highly efficient subunit c gene translational initiation region (TIR) was used as a tool in manipulating the TIRs of the other genes. Rate control of atp cistron translation occurs at the initiation level and is determined locally by each gene's TIR. In this way, individual subunit synthesis rates are set to match the requirements for H+-ATPase assembly. There is no (or very restricted) translational coupling between the cistrons. Translational initiation rates of the normally weakly expressed atp genes could be increased by up to a factor of 27 by manipulating the sequences upstream of the start codons, despite biased codon usages. In the presence of an improved upstream sequence, the N-terminal sequence of the subunit gamma gene exerted a limiting effect. This could be relieved by altering the sequence of the first seven codons. The levels of subunit gamma mRNA were more sensitive to changes in translational efficiency than the concentrations of the other atp mRNAs. The relationships between initiation efficiency and primary and secondary structure in the natural and manipulated atp TIRs are discussed in detail.

Base Sequence↗

[Non-fused expression of HAb18GEF by reducing stability of translational initiation region in mRNA].

To express the extracellular fragement of hepatoma associated antigen HAbl8G(HAb18GEF) in E. coli efficiently in a non-fusing way, the cDNA of HAb18GEF gene was inserted into prokaryotic expression vector pET21a + . The secondary structure and codon adaptation of translational initiation region (TIR, from-30 to + 39) in mRNA of recombinant vector HAb18GEF/ pET21a + was predicted simultaneously by computer-aided design. Stable Stem-Loop structures and many low-usage codons were detected in mRNA-TIR of non-optimized recombinant HAb18GEF/pET21a + vector. The stability of mRNA-TIR in recombinant HAb18GEF/pET21a + vector was reduced with following methods: (1) optimization of secondary structure (2) optimization of codon adaptation. These optimization were realized by non-continual site-directed mutagenesis without changing any amino acid sequence in TIR. After being checked through restriction endonuclease digestion and confirmed through nucleotide sequencing, the pre-optimized and post-optimized recombinant vectors were transformed into competent E. coli JM109-DE3. The resulted recombinant clones were selected randomly and induced by IPTG at 37 degrees C. The induced production of these recombinants was analyzed by SDS-PAGE, indirect ELISA, Western blot, and cell fractionation assay. The amount of HAb18GEF mRNA was also detected by RNA dot blot between pre-optimized recombinant and post-optimized recombinant. The results revealed that recombinant non-fused vectors HAb18GEF/pET21a + were successfully constructed and optimized in the secondary structure and codon adaptation of TIR respectively. The HAb18GEF was expressed efficiently in a non-fusing way in recombinant E. coli by secondary structure optimization or codon adaptation optimization. Whereas, no expression of HAb18GEF was detected in pre-optimized recombinants. The non-fused expression products-HAb18GEF, mainly as inclusion body in E. coli, yielded highly above 29.3%. A trait of expression HAb18GEF was also detected both in intermembrane space and in culture medium due to over-expression and cell leakage. Difference in non-fused expression level of HAb18GEF between secondary structure optimization and codon adaptation optimization was negligible. No difference in amount of transcribed mRNA of HAb18GEF between the pre-optimized and the post-optimized recombinants was detected. To sum up, it's feasible to express hepatoma associated antigen HAb18GEF in a non-fusing way by reducing the stability of TIR in mRNA.

Antigens, Neoplasm↗

Strand compositional asymmetries of nuclear DNA in eukaryotes.

Both DNA replication and transcription are structurally asymmetric processes. An asymmetric nucleotide substitution pattern has been observed between the leading and the lagging strand, and between the coding and the noncoding strand, in eubacterial, viral, and organelle genomes. Similar studies in eukaryotes have been rare, because the origins of replication in nuclear genomes are mostly unknown and the replicons are much shorter than those of prokaryotes. To circumvent these predicaments, all possible pairs of neighboring genes that are located on different strands of nuclear DNA were selected from the complete genomes of Saccharomyces cerevisiae, Schizosaccharomyces pombe, Plasmodium falciparum, Encephalitozoon cuniculi, Arabidopsis thaliana, Caenorhabditis elegans, Drosophila melanogaster, Anopheles gambiae, Mus musculus, and Homo sapiens. For such a pair of genes, one is likely coded from the leading strand and the other from the lagging strand. By examining the introns and the fourfold degenerate sites of codons in the genes of each pair, we found that the relative frequencies of T vs. A and of G vs. C are significantly skewed in most eukaryotes studied. In a gene pair, the potential effects of replication- and transcription-associated mutation bias on strand asymmetry are in the same direction for one gene where leading strand synthesis shares the same template with transcription, while they tend to be canceled out in the other gene. Our study demonstrates that DNA replication-associated and transcription-associated mutation bias and/or selective codon usage bias may affect the strand nucleotide composition asymmetrically in eukaryotic genomes.

Animals↗

Expression of simian virus 40 large T antigen in Escherichia coli using vectors based on the regulatable rac promoter.

Simian Virus 40 large T antigen is a multi-functional protein that is involved in the initiation of viral DNA replication, regulation of viral transcription and cell transformation. Bacterial expression vectors, pER23-1 and pER23-2, that are based on the regulatable rac promoter were used to produce T antigen either as a free protein or as a fusion protein. We have observed efficient transcription of the cloned T antigen gene in most of the recombinants. However, expression of the T antigen protein was inefficient and most of the expressed protein was truncated. This may be due to differences in codon usage in E. coli or to rapid protein degradation.

Antigens, Polyomavirus Transforming↗

Selective charging of tRNA isoacceptors induced by amino-acid starvation.

Aminoacylated (charged) transfer RNA isoacceptors read different messenger RNA codons for the same amino acid. The concentration of an isoacceptor and its charged fraction are principal determinants of the translation rate of its codons. A recent theoretical model predicts that amino-acid starvation results in 'selective charging' where the charging levels of some tRNA isoacceptors will be low and those of others will remain high. Here, we developed a microarray for the analysis of charged fractions of tRNAs and measured charging for all Escherichia coli tRNAs before and during leucine, threonine or arginine starvation. Before starvation, most tRNAs were fully charged. During starvation, the isoacceptors in the leucine, threonine or arginine families showed selective charging when cells were starved for their cognate amino acid, directly confirming the theoretical prediction. Codons read by isoacceptors that retain high charging can be used for efficient translation of genes that are essential during amino-acid starvation. Selective charging can explain anomalous patterns of codon usage in the genes for different families of proteins.

Amino Acids↗

The isochore organization and the compositional distribution of homologous coding sequences in the nuclear genome of plants.

The isochore structure of the nuclear genome of angiosperms described by Salinas et al. (1) was confirmed by using a different experimental approach, namely by showing that the levels of coding sequences from both dicots and Gramineae are linearly correlated with GC levels of the corresponding flanking sequences. The compositional distribution of homologous coding sequences from several orders of dicots and from Gramineae were also studied and shown to mimick the compositional distributions previously seen (1) for coding sequences in general, most coding sequences from Gramineae being much higher than those of the dicots explored. These differences were even stronger for third codon positions and led to striking codon usages for many coding sequences especially in the case of Gramineae.

Cell Nucleus↗

Soluble, highly fluorescent variants of green fluorescent protein (GFP) for use in higher plants.

Green fluorescent protein (GFP) from Aequorea victoria has rapidly become a standard reporter in many biological systems. However, the use of GFP in higher plants has been limited by aberrant splicing of the corresponding mRNA and by protein insolubility. It has been shown that GFP can be expressed in Arabidopsis thaliana after altering the codon usage in the region that is incorrectly spliced, but the fluorescence signal is weak, possibly due to aggregation of the encoded protein. Through site-directed mutagenesis, we have generated a more soluble version of the codon-modified GFP called soluble-modified GFP (smGFP). The excitation and emission spectra for this protein are nearly identical to wild-type GFP. When introduced into A. thaliana, greater fluorescence was observed compared to the codon-modified GFP, implying that smGFP is 'brighter' because more of it is present in a soluble and functional form. Using the smGFP template, two spectral variants were created, a soluble-modified red-shifted GFP (smRS-GFP) and a soluble-modified blue-fluorescent protein (smBFP). The increased fluorescence output of smGFP will further the use of this reporter in higher plants. In addition, the distinct spectral characters of smRS-GFP and smBFP should allow for dual monitoring of gene expression, protein localization, and detection of in vivo protein-protein interactions.

Arabidopsis↗

Codon and amino acid usage in retroviral genomes is consistent with virus-specific nucleotide pressure.

Retroviral RNA genomes are known to have a biased nucleotide composition. For instance, the plus-strand RNA of human immunodeficiency virus (HIV) is A-rich, and the genome of human T cell leukemia virus (HTLV) is C-rich, and other retroviruses have a U-rich or G-rich genome. The biased composition of these genomes is most likely caused by directional mutational pressure of the respective reverse transcriptase enzymes. Using a set of retroviral genomes with a distinct nucleotide composition, we performed skew analyses of the nucleotide bias along the complete viral genome. Distinct nucleotide signatures were apparent, and these typical patterns were generally conserved across the viral genome. Furthermore, it is demonstrated that this typical nucleotide bias, combined with a profound discrimination against the CpG dinucleotide sequence, strongly influences the codon usage of the retroviruses in a direct manner, and their amino acid usage in an indirect manner. The fact that both codon usage and amino acid usage are so closely entwined with the genome composition has important practical implications. For instance, the typical trends in nucleotide usage could influence the molecular phylogenetic reconstruction of the family Retroviridae.

Amino Acids↗

The structure of the Escherichia coli hemB gene.

The Escherichia coli hemB gene, which encodes 5-aminolevulinic acid dehydratase, and was cloned into pTZ18U, a multicopy plasmid, was sequenced. The hemB insert was double-digested with restriction enzymes and recloned back into pTZ18U and pTZ19U to allow for sequencing in two directions. In a second procedure, used to fill in gaps and to confirm the sequence derived from the first procedure, the whole insert was cloned into M13 phages. A nested set of deletions was constructed and recloned into M13. Both the double-digested fragments cloned into plasmids pTZ18U and pTZ19U and the overlapping fragments contained in M13 phages were sequenced using the dideoxy procedure with [35S]dATP. Computer software was used to identify coding regions and the correct reading frame. Two promoter regions, two Shine-Dalgarno sequences and two possible start sites were identified. Extensive homologies with yeast (36%), human liver (40%) and rat liver (40%) amino-acid (aa) sequences were observed, especially in the 16-aa Zn-binding region (75%) and the 4 aa surrounding the essential lysine at the active site (100% for rat and human proteins). Computer analysis of promoter strength and two independent analyses of codon usage indicated that the hemB gene is moderately expressed.

Amino Acid Sequence↗

Structure of the Bacillus sphaericus R modification methylase gene.

A 2.5 X 10(3) base-pair segment of Bacillus sphaericus R DNA cloned in Escherichia coli has previously been shown to carry the functional BspRI modification methylase gene. The approximate location of the gene on this DNA segment and its direction of transcription were established by subcloning experiments. The nucleotide sequence of the relevant region was determined by the Maxam-Gilbert procedure. An open reading frame that can code for a 424 amino acid protein was found. The calculated molecular weight (48,264) of this protein is in fair agreement with previous estimates (50,000 to 52,000). The synthesis of this protein was demonstrated in E. coli minicells. The initiation point of transcription by E. coli RNA polymerase was localized by in vitro transcription experiments. The open reading frame starts 29 base-pairs downstream from the transcription initiation site and it is preceded by a sequence showing extensive Shine-Dalgarno complementarity. Subcloning experiments and translation in minicells suggest that after removal of this translational initiation site, a secondary start site 29 amino acids downstream can also start translation in E. coli, and this shorter protein retains the methylase activity. The overall base composition of the gene and the codon usage indicate a strong preference for A.T base-pairs.

Bacillus↗

Coding sequence evolution.

Dramatic progress has been made in the past ten years in the development of statistical and experimental techniques for investigating features of molecular evolution. Applied to coding regions, these techniques have produced remarkable advances in our understanding of selection for codon usage but, ironically, have had little impact on our understanding of protein evolution. That may be about to change.

Animals↗

The nucleotide sequence of the PRI1 gene related to DNA primase in Saccharomyces cerevisiae.

The PRI1 gene of Saccharomyces cerevisiae encodes for the p48 polypeptide of DNA primase. We have determined the nucleotide sequence of a 1,965 bp DNA fragment containing the PRI1 locus. The entire coding sequence of the gene lies within an open reading frame, and there are 409 amino acids in the single polypeptide protein if translation is assumed to start at the first ATG in this frame. The 5' and 3' end-points of PRI1 mRNA have been determined by S1 mapping and primer extension analysis. The primary structure and the codon usage of PRI1 suggest that this essential gene is poorly expressed in yeast cells.

Amino Acid Sequence↗

TransTerm: a database of translational signals.

The TransTerm database of sequence contexts of stop and start codons has been expanded to include approximately 50% more species than last year's release. It now contains 148 organisms and >39 500 coding sequences; it is now available on the World Wide Web. The database includes: (i) initiation and termination sequence contexts organized by species; (ii) summary parameters about the individual sequences (sequence length, GC%, GC3, Nc, CAI) in addition to tables of base frequencies for each species' stop and start codon sequence context; (iii) species codon usage tables; and (iv) summary tables of stop signal frequency.

Animals↗

Analysis of distribution of bases in the coding sequences by a diagrammatic technique.

The frequencies of occurrence of four bases in the first, second and third codon positions and in the total coding sequences have been calculated by the codon usage table published in 1990 by Ikemura et al. The distribution of frequencies are further analysed in detail by a graphic technique presented recently by us. Formulas expressing the frequencies of four bases in the first and second codon positions in terms of frequencies of amino acids have been given. It is shown by the graphic analysis that for 90 species, in the first codon position the purine bases are dominant and in most cases G is the most dominant base. In the second codon position A is the most dominant base, while G is the least dominant base. In the third codon position the G + C content varies from 0.1 to 0.9, keeping the A + C content equal to 1/2 and G content equal to that of C, approximately. If the frequencies for bases A, C, G and U in the total coding sequences are denoted by a, c, g and u, respectively, it is found that the unequal formula: a2 + c2 + g2 + u2 less than 1/3, is valid for each of the 90 species including the human and E.coli etc.

Base Sequence↗