Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Antisense overlapping open reading frames in genes from bacteria to humans.

Long Open Reading Frames (ORFs) in antisense DNA strands have been reported in the literature as being rare events. However, an extensive analysis of the GenBank database revealed that a substantial number of genes from several species contain an in-phase ORF in the antisense strand, that overlaps entirely the coding sequence of the sense strand, or even extends beyond. The findings described in this paper show that this is a frequent, non-random phenomenon, which is primarily dependent on codon usage, and to a lesser extent on gene size and GC content. Examination of the sequence database for several prokaryotic and eukaryotic organisms, demonstrates that coding sequences with in-phase, 100% overlapping antisense ORFs are present in every genome studied so far.

Amino Acid Sequence↗

[Construction and immune potency of recombinant adenovirus containing codon-modified HIV-1 gp120].

BACKGROUND: To construct replication-deficient recombinant adenovirus expressing wild and codon-modified HIV-1 gp120. METHODS: The viral codons were changed to the codon usage of highly expressed mammal gene, the resulting modified gp120 gene was synthesized. The wild and modified gp120 genes were cloned into shuttle vector pShuttle-CMV respectively, and then the constructed plasmids containing gp120 gene was cotransformed with the backbone vector pADeasy-1 into E.coli BJ5183. Transfection of the recombinant AdEasy plasmid into 293 cells was performed to obtain recombinant adenoviruses. The mice were immunized with the recombinant adenoviruses. Their immunogenicity was evaluated by testing antibody and CTL levels of immunized mice. RESULTS: Two strains of recombinant adenovirus expressing wild and codon-modified HIV-1 gp120 were obtained. The protein expressing level of the recombinant adenoviruses containing modified genes was much higher than that containing wild genes. The mice immunized with recombinant adenoviruses elicited HIV-1 specific antibody and CTL response. The rAd-mod gp120 group was better than the rAd-wt gp120 group. CONCLUSION: Replication-deficient recombinant adenovirus expressing HIV-1 gp120 can elicit HIV-1 specific humoral and cellular response, the codon-modified recombinant virus was more efficient than the native.

AIDS Vaccines↗

Translational pauses during the synthesis of proteins and mRNA structure.

Translational pauses are observed during a spider fibroin synthesis (1,2). The spider major ampullate (dragline) silk of the spider Nephila clavipes is composed of multiple proteins. The amino acid sequences of the partial cDNA clones for the two major dragline silk fibroin components (Spidroin 1 and 2) exhibit repetitive motifs (3,4). Our detailed inspection of the nucleotide sequences of the repetitive motifs revealed highly selective site-specific codon usage patterns within a motif, suggesting that the secondary structure of the spider fibroin mRNA is optimized by the nucleotide sequence of the fibroin gene. The results, combined with our preceding results on silk fibroin from Bombyx mori (5) suggest that translational pauses of spider silk are interpreted in terms of the mRNA secondary structure.

Amino Acid Sequence↗

Nucleotide sequences encoding and promoting expression of three antibiotic resistance genes indigenous to Streptomyces.

Promoter-probe plasmid vectors were used to isolate putative promoter-containing DNA fragments of three Streptomyces antibiotic resistance genes, the rRNA methylase (tsr) gene of S. azureus, the aminoglycoside phosphotransferase (aph) gene of S. fradiae, and the viomycin phosphotransferase (vph) gene of S. vinaceus. DNA sequence analysis was carried out for all three of the fragments and for the protein-coding regions of the tsr and vph genes. No sequences resembling typical E. coli promoters or Bacillus vegetatively-expressed promoters were identified. Furthermore, none of the three DNA fragments found to be transcriptionally active in Streptomyces could initiate transcription when introduced into E. coli. An extremely biased codon usage pattern that reflects the high G + C composition of Streptomyces DNA was observed for the protein-coding regions of the tsr and vph genes, and of the previously sequenced aph gene. This pattern enabled delineation of the protein-coding region and identification of the coding strand of the genes.

Base Sequence↗

The nucleotide sequences of the tail fiber gene 36 of bacteriophage T2 and of genes 36 of the T-even type Escherichia coli phages K3 and Ox2.

Genes 36 have been cloned from phage T2 and the T-even type phages K3 and Ox2. The products of these genes are part of the long tail fibers of the phages, they form the proximal moiety of the distal half fiber. The genes have been sequenced, the nucleotide sequence of gene 36 of phage T4 is known (Oliver, D.B. & Crowther, R.A. (1981) J.Mol.Biol. 153, 545-568). Comparison of the deduced amino acid sequences of the four proteins revealed a surprising pattern. These sequences can be divided into two highly conserved and one very variable region. The former consist of about 60 NH2-terminal and 70 CO2H-terminal residues flanking the variable middle region of about 100 residues. Thus, an identical and unique morphology can be formed by a number of different primary structures. It is proposed that the conserved areas are involved in binding of the proteins to the neighboring products of genes 35 and 37 and that this function has put constraints on the variability of the primary protein structure. The overall amino acid composition of the proteins is rather similar; the codon usage is that known for phage T4. The intercistronic region between genes 35 and 36 consisting of 62 base pairs and containing a presumed terminator for g35 transcription and the 'late type' promoter for transcription of genes 36, 37, and 38, is almost completely identical in the four phages.

Amino Acid Sequence↗

Root nodule Bradyrhizobium spp. harbor tfdAalpha and cadA, homologous with genes encoding 2,4-dichlorophenoxyacetic acid-degrading proteins.

The distribution of tfdAalpha and cadA, genes encoding 2,4-dichlorophenoxyacetate (2,4-D)-degrading proteins which are characteristic of the 2,4-D-degrading Bradyrhizobium sp. isolated from pristine environments, was examined by PCR and Southern hybridization in several Bradyrhizobium strains including type strains of Bradyrhizobium japonicum USDA110 and Bradyrhizobium elkanii USDA94, in phylogenetically closely related Agromonas oligotrophica and Rhodopseudomonas palustris, and in 2,4-D-degrading Sphingomonas strains. All strains showed positive signals for tfdAalpha, and its phylogenetic tree was congruent with that of 16S rRNA genes in alpha-Proteobacteria, indicating evolution of tfdAalpha without horizontal gene transfer. The nucleotide sequence identities between tfdAalpha and canonical tfdA in beta- and gamma-Proteobacteria were 46 to 57%, and the deduced amino acid sequence of TfdAalpha revealed conserved residues characteristic of the active site of alpha-ketoglutarate-dependent dioxygenases. On the other hand, cadA showed limited distribution in 2,4-D-degrading Bradyrhizobium sp. and Sphingomonas sp. and some strains of non-2,4-D-degrading B. elkanii. The cadA genes were phylogenetically separated between 2,4-D-degrading and nondegrading strains, and the cadA genes of 2,4-D degrading strains were further separated between Bradyrhizobium sp. and Sphingomonas sp., indicating the incongruency of cadA with 16S rRNA genes. The nucleotide sequence identities between cadA and tftA of 2,4,5-trichlorophenoxyacetate-degrading Burkholderia cepacia AC1100 were 46 to 53%. Although all root nodule Bradyrhizobium strains were unable to degrade 2,4-D, three strains carrying cadA homologs degraded 4-chlorophenoxyacetate with the accumulation of 4-chlorophenol as an intermediate, suggesting the involvement of cadA homologs in the cleavage of the aryl ether linkage. Based on codon usage patterns and GC content, it was suggested that the cadA genes of 2,4-D-degrading and nondegrading Bradyrhizobium spp. have different origins and that the genes would be obtained in the former through horizontal gene transfer.

2,4-Dichlorophenoxyacetic Acid↗

Plasmid mutagenesis by PCR for high-level expression of para-hydroxybenzoate hydroxylase.

We report a PCR deletion mutagenesis method for the exact positioning of a foreign gene (pobA) in the lac operon of an expression plasmid in place of the lacZ protein code. This method requires the synthesis of four oligonucleotides and three PCR reactions to delete unwanted bases and retain the nucleotide sequence naturally found between the lac promoter and the protein code. The engineered plasmid results in the production of at least 40% of the cellular protein as the foreign polypeptide. In the example presented the expression of the protein is high even with a substantial difference in codon usage between the host (Escherichia coli) and a foreign gene from Pseudomonas aeruginosa. Some of the polypeptide produced has the ame properties as native protein and is easily purified. The remainder is present as insoluble inclusion bodies. This method of plasmid refinement may be applicable to the expression of many proteins.

4-Hydroxybenzoate-3-Monooxygenase↗

Nucleotide sequence of the Bacillus anthracis edema factor gene (cya): a calmodulin-dependent adenylate cyclase.

The nucleotide sequence of the Bacillus anthracis edema factor (EF) gene (cya), which encodes a calmodulin-dependent adenylate cyclase, has been determined. EF is part of the tripartite protein exotoxin of B. anthracis. An ATG start codon, immediately upstream from codons which specify the first 15 amino acids (aa) of EF, was preceded by an AAAGGAGGT sequence which is its probable ribosome-binding site. Starting at this ATG codon, there was a continuous 2400-bp open reading frame which encodes the 800-aa EF-precursor protein with a Mr of 92,464. The mature, secreted protein (767 aa; Mr 88,808) was preceded by a 33-aa signal peptide which has characteristics in common with leader peptides for other secreted proteins of the Bacillus species. A consensus amino acid sequence (Gly-X-X-X-X-Gly-Lys-Ser,X = any aa), which was part of the presumed ATP binding site for EF, was also present. The codon usage of the EF gene reflected the high A + T (71%) base composition for its DNA. B. anthracis EF was not related to the Escherichia coli or yeast adenylate cyclases, but was related to the Bordetella pertussis calmodulin-dependent adenylate cyclase.

Amino Acid Sequence↗

Quantitative assessment of peptide sequence diversity in M13 combinatorial peptide phage display libraries.

Novel statistical methods have been developed and used to quantitate and annotate the sequence diversity within combinatorial peptide libraries on the basis of small numbers (1-200) of sequences selected at random from commercially available M13 p3-based phage display libraries. These libraries behave statistically as though they correspond to populations containing roughly 4.0+/-1.6% of the random dodecapeptides and 7.9+/-2.6% of the random constrained heptapeptides that are theoretically possible within the phage populations. Analysis of amino acid residue occurrence patterns shows no demonstrable influence on sequence censorship by Escherichia coli tRNA isoacceptor profiles or either overall codon or Class II codon usage patterns, suggesting no metabolic constraints on recombinant p3 synthesis. There is an overall depression in the occurrence of cysteine, arginine and glycine residues and an overabundance of proline, threonine and histidine residues. The majority of position-dependent amino acid sequence bias is clustered at three positions within the inserted peptides of the dodecapeptide library, +1, +3 and +12 downstream from the signal peptidase cleavage site. Conformational tendency measures of the peptides indicate a significant preference for inserts favoring a beta-turn conformation. The observed protein sequence limitations can primarily be attributed to genetic codon degeneracy and signal peptidase cleavage preferences. These data suggest that for applications in which maximal sequence diversity is essential, such as epitope mapping or novel receptor identification, combinatorial peptide libraries should be constructed using codon-corrected trinucleotide cassettes within vector-host systems designed to minimize morphogenesis-related censorship.

Amino Acids↗

Genetic variation between Helicobacter pylori strains: gene acquisition or loss?

Previously identified strain-specific genes of Helicobacter pylori were analysed for GC content and preference in codon usage. The results indicate that in H. pylori strain specificity is mainly driven by gene uptake. Incoming strains of Helicobacter or other species can occasionally donate genes but the identification of the donor species is hampered by ongoing evolutionary processes and the lack of an adequate number, or indeed a total absence, of gene homologues.

Codon↗

Variable rates of evolution among Drosophila opsin genes.

DNA sequences and chromosomal locations of four Drosophila pseudoobscura opsin genes were compared with those from Drosophila melanogaster, to determine factors that influence the evolution of multigene families. Although the opsin proteins perform the same primary functions, the comparisons reveal a wide range of evolutionary rates. Amino acid identities for the opsins range from 90% for Rh2 to more than 95% for Rh1 and Rh4. Variation in the rate of synonymous site substitution is especially striking: the major opsin, encoded by the Rh1 locus, differs at only 26.1% of synonymous sites between D. pseudoobscura and D. melanogaster, while the other opsin loci differ by as much as 39.2% at synonymous sites. Rh3 and Rh4 have similar levels of synonymous nucleotide substitution but significantly different amounts of amino acid replacement. This decoupling of nucleotide substitution and amino acid replacement suggests that different selective pressures are acting on these similar genes. There is significant heterogeneity in base composition and codon usage bias among the opsin genes in both species, but there are no consistent relationships between these factors and the rate of evolution of the opsins. In addition to exhibiting variation in evolutionary rates, the opsin loci in these species reveal rearrangements of chromosome elements.

Amino Acid Sequence↗

Interspecific and intraspecific comparisons of the period locus in the Drosophila willistoni sibling species.

The period (per) locus has received much attention in molecular evolution studies because it is one of the best studied "behavioral genes" and because it offers insight into the evolution of repetitive sequences. We studied most of the coding region of per in Drosophila willistoni and confirmed previously observed patterns of conservation and divergence among distantly related species. Five regions are so highly diverged that they cannot be aligned, whereas a region encompassing the PAS domain is very conserved. Structural and nucleotide polymorphism patterns in the willistoni group are not the same as those observed in previously studied species. We sequenced the region homologous to the highly polymorphic threonine-glycine repeat of D. melanogaster in multiple strains of D. willistoni, as well as in other members of willistoni group, and found an unusual amount of conservation in this region. However, the next nonconserved region downstream in the sequence is quite variable and polymorphic for the number of repeated glycines. The glycine codon usage is significantly different in this glycine repeat as compared to other parts of the gene. We were able to plot the directionality of change in the glycine repeat region onto a phylogeny and find that the addition of glycines is the general trend with the diversification of the willistoni group.

Amino Acid Sequence↗

Guanine-adenine bias: a general property of retroid viruses that is unrelated to host-induced hypermutation.

The recently discovered mammalian enzymes, APOBEC3G and 3F, induce guanine-to-adenine hypermutation in retroviruses. However, the preference of adenine over guanine in retroviral codon usage is not correlated with the presence or absence of APOBEC3G or its viral inhibitor (Vif), and its pattern does not reflect the biochemical properties of APOBEC3G action. The guanine-adenine bias of retroviruses is thus probably not a result of host-induced mutational pressure, but rather reflects a general predisposition associated with reverse transcription.

APOBEC-3G Deaminase↗

Nucleotide sequence of the trpD and trpC genes of Salmonella typhimurium.

We have completed the nucleotide sequence determination of trpD and trpC, the second and third genes of the trp operon of Salmonella typhimurium. These genes encode two bifunctional proteins thought to have arisen by gene fusions: the trpD polypeptide contains the glutamine amido transferase and the phosphoribosyl anthranilate transferase activities, and the trpC protein possesses the N-(5'-phosphoribosyl)-anthranilic acid isomerase and the indole-3-glycerol phosphate synthetase activities. The trpD gene consists of 1593 nucleotides encoding 531 amino acids, and possesses an internal promoter (p2) located within a region from about 1400 to 1441 of the nucleotide sequence. The trpC gene contains 1356 nucleotides encoding 452 amino acids. In this paper we compare the trpD and trpC genes of S. typhimurium to those of Escherichia coli with respect to codon usage, nucleotide and amino acid conservation, p2 promoter characteristics and intercistronic regions. The sequence of the two genes we present here completes the sequence determination of the trp operon of S. typhimurium and should prove useful in comparisons with the E. coli trp operon and in future studies of operon structure in S. typhimurium.

Amino Acid Sequence↗

Cloning, molecular characterization and chromosome localization of the inorganic pyrophosphatase (PPA) gene from S. cerevisiae.

The gene for Saccharomyces cerevisiae inorganic pyrophosphatase, PPA, has been cloned by hybridization of "long" oligonucleotide probes with both cDNA and genomic S. cerevisiae libraries. The nucleotide sequence of 1612 bp from a genomic subclone that includes the entire coding region gives a deduced amino acid sequence that has nine differences (out of a total of 286 residues) from the previously published amino acid sequence that was determined directly. The codon usage in PPA is as expected for a "highly expressed" yeast gene. The upstream region contains a poly dA/dT sequence that might comprise a constitutive promoter. The PPA gene appears to be present in a single copy within the S. cerevisiae genome and has been localized to chromosome II.

Amino Acid Sequence↗

Selection of mutations that increase alpha 1-antitrypsin gene expression in Escherichia coli.

The gene encoding human alpha-1-antitrypsin (A1AT), when cloned and expressed as a full-length, non-fusion gene product in Escherichia coli, accumulates to levels up to 0.1% of total cellular protein. Truncation of the gene at its 5' end or synthesis as a fusion protein increases expression up to 200-fold. Extensive mutagenesis in vitro within this same 5'-terminal region aimed at improving codon usage and disrupting potential secondary structure increased expression only 10 to 20-fold. We have developed a translational fusion system for selecting mutations and applied it to the study of A1AT expression in E. coli. With this methodology, we have obtained single base-pair mutations having up to a 20-fold effect on A1AT expression. When we combined these multiple single base-pair mutations, we achieve up to a 200-fold increase in A1AT expression. The resulting gene product is of authentic size (394 amino acid residues) and contains two amino acid substitutions (Asn in place of Asp) in codons 2 and 6. This protein is primarily in the soluble fraction of the E. coli lysate and has identical activity to A1AT purified from human sera. The methodology used to generate these mutations may be generally applicable to the study of genes that do not express well in E. coli initially, and provides an alternative to secondary structure analysis in the redesign of such genes.

Amino Acid Sequence↗

Total synthesis and expression in Escherichia coli of a gene encoding human tropoelastin.

To elucidate the structural features and interactions of tropoelastin (TEL) molecules which assist in giving the elastic fibre its physical properties, a 2210-bp synthetic human TEL-encoding gene (SHEL) was constructed for expression in Escherichia coli. To this end, a model of codon adjustment was tested which better suits the polypeptide biosynthetic needs of E. coli than the human sequence, where over one-third of this natural sequence contains expression-limiting rare codons and 4 amino acids alone account for 75% of the resulting polypeptide. This large synthetic TEL gene was expressed at a high level as the recombinant counterpart of human TEL and as a C-terminal fusion with glutathione S-transferase. This demonstrates that a synthetic approach based upon matching codon usage to that of the host organism can support significant expression of recombinant sequences. The synthetic gene incorporates the facility for simple cassette replacement in future insertion, deletion and mutagenesis experiments, including the introduction and removal of exon homologues. The resulting soluble polypeptide is easily purified and displays properties expected for this protein.

Amino Acid Sequence↗

Translational selection and yeast proteome evolution.

The primary structures of peptides may be adapted for efficient synthesis as well as proper function. Here, the Saccharomyces cerevisiae genome sequence, DNA microarray expression data, tRNA gene numbers, and functional categorizations of proteins are employed to determine whether the amino acid composition of peptides reflects natural selection to optimize the speed and accuracy of translation. Strong relationships between synonymous codon usage bias and estimates of transcript abundance suggest that DNA array data serve as adequate predictors of translation rates. Amino acid usage also shows striking relationships with expression levels. Stronger correlations between tRNA concentrations and amino acid abundances among highly expressed proteins than among less abundant proteins support adaptation of both tRNA abundances and amino acid usage to enhance the speed and accuracy of protein synthesis. Natural selection for efficient synthesis appears to also favor shorter proteins as a function of their expression levels. Comparisons restricted to proteins within functional classes are employed to control for differences in amino acid composition and protein size that reflect differences in the functional requirements of proteins expressed at different levels.

Adaptation, Physiological↗