Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Evidence for use of rare codons in the dnaG gene and other regulatory genes of Escherichia coli.

Amino acid sequence and composition data of Escherichia coli dnaG primase protein and its tryptic peptides have confirmed that the dnaG gene contains an unusually high number of codons that are not frequently used in most E. coli genes. In 25 E. coli proteins analyzed the codons AUA, UCG, CCU, CCC, ACG, CAA, AAT, and AGG are infrequently used, occurring as 4% of the total codons in the reading frame and 11% and 10% in the nonreading frames. In dnaG they occur as 11% in the reading frame and 12% in the nonreading frames. The rpsU and rpoD genes, which flank the dnaG gene [Smiley, B. L., Lupski, J. R., Svec, P. S., McMacken, R. & Godson, G. N. (1982) Proc. Natl. Acad. Sci. USA 79, 4550-4554], however, have normal codon usage. Translational modulation using isoaccepting tRNA availability may therefore be part of the mechanism of keeping the dnaG gene expression low, while expression of the adjacent rpsU and rpoD genes on the same mRNA transcript is high.

Amino Acid Sequence↗

Structure of a protein superfiber: spider dragline silk.

Spider major ampullate (dragline) silk is an extracellular fibrous protein with unique characteristics of strength and elasticity. The silk fiber has been proposed to consist of pseudocrystalline regions of antiparallel beta-sheet interspersed with elastic amorphous segments. The repetitive sequence of a fibroin protein from major ampullate silk of the spider Nephila clavipes was determined from a partial cDNA clone. The repeating unit is a maximum of 34 amino acids long and is not rigidly conserved. The repeat unit is composed of three different segments: (i) a 6 amino acid segment that is conserved in sequence but has deletions of 3 or 6 amino acids in many of the repeats; (ii) a 13 amino acid segment dominated by a polyalanine sequence of 5-7 residues; (iii) a 15 amino acid, highly conserved segment. The latter is predominantly a Gly-Gly-Xaa repeat with Xaa being alanine, tyrosine, leucine, or glutamine. The codon usage for this DNA is highly selective, avoiding the use of cytosine or guanine in the third position. A model for the physical properties of fiber formation, strength, and elasticity, based on this repetitive protein sequence, is presented.

Amino Acid Sequence↗

Fission yeast gene structure and recognition.

A database of 210 Schizosaccharomyces pombe DNA sequences (524,794 bp) was extracted from GenBank (release number 81.0) and examined by a number of methods in order to characterize statistical features of these sequences that might serve as signals or constraints for messenger RNA splicing. The statistical information compiled includes splicing signal (donor, acceptor and branch site) profiles, translational initiation start profile, exon/intron length distributions, ORF distribution, CDS size distribution, codon usage table, and 6-tuple distribution. The information content of the various signals are also presented. A rule-based interactive computer program for finding introns called INTRON.PLOT has been developed and was used to successfully analyze 7 newly sequenced genes.

Base Sequence↗

Sequence analysis of expressed sequence tags from an ABA-treated cDNA library identifies stress response genes in the moss Physcomitrella patens.

Partial cDNA sequencing was used to obtain 169 expressed sequence tags (ESTs) in the moss, Physcomitrella patens. The source of ESTs was a random cDNA library constructed from 7 day-old protonemata following treatment with 10(-4) M abscisic acid (ABA). Analysis of the ESTs identified 69% with homology to known sequences, 61% of which had significant homology to sequences of plant origin. More importantly, at least 11 ESTs had significant similarities to genes which are implicated in plant stress-responses, including responses which may involve ABA. These included a cDNA associated with desiccation tolerance, two heat shock protein genes, one cold acclimation protein cDNA and five others that may be involved in either oxidative or chemical stress or both, i.e., Zn/Cu-superoxide dismutase, NADPH protochlorophyllide oxidoreductase (PorB), selenium binding protein, glutathione peroxidase and glutathione S transferase. Analysis of codon usage between P. patens and seed plants indicated that although mosses and higher plants are to a large extent similar, minor variations also exists that may represent the distinctiveness of each group.

Abscisic Acid↗

The complete nucleotide sequence of coxsackievirus A21.

We have determined the complete nucleotide sequence of coxsackievirus A21 (CAV-21), the first member of this enterovirus subgroup to be analysed in molecular detail. The sequence, which is 7401 nucleotides long, encodes an open reading frame of 2206 codons, preceded by a 5' non-coding region of 711 nucleotides and followed by a 3' non-coding region of 72 nucleotides plus a poly(A) tract. The most striking feature is the remarkable homology to the poliovirus (greater than 90% at the amino acid level) in the 3' part of the genome. The rest of the genome is much less homologous, suggesting that CAV-21 is a recombinant virus. Rhinovirus-like characteristics, including the length of the 5' non-coding region and a slight --U/--A imbalance in codon usage, may be related to the fact that CAV-21, like rhinoviruses, infects the upper respiratory tract. However, the sequence sheds little light on the molecular basis of the shared receptor specificity.

Amino Acid Sequence↗

Nucleotide and amino acid polymorphisms at drug resistance sites in non-B-subtype variants of human immunodeficiency virus type 1.

We have compared nucleotide substitutions and polymorphisms at codons known to confer drug resistance in subtype B strains of human immunodeficiency virus type 1 (HIV-1) with similar substitutions in viruses of other subtypes. Genotypic analysis was performed on viruses from untreated individuals. Nucleotide and amino acid diversity at resistance sites was compared with a consensus subtype B reference virus. Among patients with non-subtype B infections, polymorphisms relative to subtype B were observed at codon 10 in protease (PR). These included silent substitutions (CTC-->CTT, CTA, TTA) and an amino acid mutation, L10I. Subtype A viruses possessed a V179I substitution in reverse transcriptase (RT). Subtype G viruses were identified by silent substitutions at codon 181 in RT (TAT-->TAC). Similarly, subtype A/G viruses were identified by a substitution at position 67 in RT (GAC-->GAT). Subtype C was distinguished by silent substitutions at codons 106 (GTA-->GTG) and 219 (AAA-->AAG) in RT and codon 48 (GGG-->GGA) in PR. Variations relative to subtype B were seen at RT position 215 (ACC-->ACT) for subtypes A and A/E. These substitutions and polymorphisms reflect different patterns of codon usage among viruses of different subtypes. However, the existence of different subtypes may only rarely affect patterns of drug resistance-associated mutations.

Amino Acid Substitution↗

Cloning and sequencing of the gene for apocytochrome b of the yeast Kluyveromyces lactis strains WM27 (NRRL Y-17066) and WM37 (NRRL Y-1140).

The apocytochrome b genes from two strains of the yeast Kluyveromyces lactis, have been isolated and sequenced. The coding sequences in strains WM27 (NRRL Y-17066) and WM37 (NRRL Y-1140) were identical but the upstream noncoding regions were slightly different. The sequences demonstrated the presence of a continuous open reading frame with no introns. The amino acid sequence, derived from the coding strand, showed 82% homology to the apocytochrome b of Saccharomyces cerevisiae strain D273-10B and only 58% homology to the protein from Schizosaccharomyces pombe strain 50. CUN and CGN codon families were absent from the K. lactis gene. Codon usage was very similar to that of other mitochondrial genomes with mostly U or A in the third position. There were two unusual features. All threonines were coded by ACA(U) and all arginines by AGA.

Amino Acid Sequence↗

Isolation by genetic complementation of two differentially expressed genes for beta-isopropylmalate dehydrogenase from Aspergillus niger.

We have constructed an Aspergillus niger cDNA library with a yeast expression vector. The library DNA complemented a leucine auxotroph of Saccharomyces cerevisiae (strain BWG1-7a) at a frequency of 4x10(-4). Plasmids rescued from the yeast prototrophs also complemented Escherichia coli (strain MC1066) deficient in leucine biosynthesis. Sequence determination of the rescued plasmids revealed two genes for beta-isopropylmalate dehydrogenase, which we called leu2A and leu2B. Genomic-blot analysis suggested that both leu2A and leu2B were derived from single-copy genes. Northern-blot hybridization showed that in nutrient-rich medium a leu2A transcript accumulated during germination and log-phase growth while the leu2B transcript appeared late in the growth phase. In minimal medium, only leu2A expression was greatly stimulated. We examined the codon preference of these two genes. Whereas leu2A shows a bias in codon usage typical of A. niger genes, leu2B does not. These results indicate the presence in A. niger of two highly divergent, differentially regulated, isozymes for beta-isopropylmalate dehydrogenase.

3-Isopropylmalate Dehydrogenase↗

The complete nucleotide sequence of RNA beta from the type strain of barley stripe mosaic virus.

The complete nucleotide sequence of RNA beta from the type strain of barley stripe mosaic virus (BSMV) has been determined. The sequence is 3289 nucleotides in length and contains four open reading frames (ORFs) which code for proteins of Mr 22,147 (ORF1), Mr 58,098 (ORF2), Mr 17,378 (ORF3), and Mr 14,119 (ORF4). The predicted N-terminal amino acid sequence of the polypeptide encoded by the ORF nearest the 5'-end of the RNA (ORF1) is identical (after the initiator methionine) to the published N-terminal amino acid sequence of BSMV coat protein for 29 of the first 30 amino acids. ORF2 occupies the central portion of the coding region of RNA beta and ORF3 is located at the 3'-end. The ORF4 sequence overlaps the 3'-region of ORF2 and the 5'-region of ORF3 and differs in codon usage from the other three RNA beta ORFs. The coding region of RNA beta is followed by a poly(A) tract and a 238 nucleotide tRNA-like structure which are common to all three BSMV genomic RNAs.

Amino Acid Sequence↗

The Enhanced Microbial Genomes Library.

Since the obtention of the complete sequence of Haemophilus influenzae Rd in 1995, the number of bacterial genomes entirely sequenced has regularly increased. A problem is that the quality of the annotations of these very large sequences is usually lower than those of the shorter entries encountered in the repository collections. Moreover, classical sequence database management systems have difficulties in handling entries of that size. In this context, we have decided to build the Enhanced Microbial Genomes Library (EMGLib) in which these two problems are alleviated. This library contains all the complete genomes from bacteria already sequenced and the yeast genome in GenBank format. The annotations are improved by the introduction of data on codon usage, gene orientation on the chromosome and gene families. It is possible to access EMGLib through two database systems set up on World Wide Web servers: the PBIL server at http://pbil.univ-lyon1.fr/emglib/emglib. html and the MICADO server at http://locus.jouy.inra.fr/micado

Base Sequence↗

Conserved Ser residues, the shutter region, and speciation in serpin evolution.

The suicide inhibitory mechanism of serine protease inhibitors of the serpin superfamily depends heavily on their structural flexibility, which is controlled in large part by the breach and shutter regions of the central Abeta-sheet. We examined codon usage by the highly conserved residues, Ser-53 and Ser-56, of the shutter region and found a TCN-AGY usage dichotomy for Ser-56 that remarkably is linked to the protostome-deuterostome split. Our results suggest that serpin evolution was driven by phylogenetic speciation and not pressure to fulfill new physiologic functions mitigating against coevolution with the family of serine proteases they inhibit.

Codon↗

Cloning and characterization of the rad4 gene of Schizosaccharomyces pombe; a gene showing short regions of sequence similarity to the human XRCC1 gene.

The rad4.116 mutant of the fission yeast Schizosaccharomyces pombe is temperature-sensitive for growth, as well as being sensitive to the killing actions of both ultraviolet light and ionizing radiation. We have cloned the rad4 gene by complementation of the temperature sensitive phenotype of the rad4.116 mutant with a S. pombe gene bank. The rad4 gene fully complemented the UV sensitivity of the rad4.116 mutant. The gene is predicted to encode a protein of 579 amino acids with a basic tail, a possible zinc finger and a nuclear location signal. The amino terminal part of the predicted rad4 ORF contains two short regions of similarity to the C-terminal part of the human XRCC1 gene. Codon usage suggests that the gene is very poorly expressed, and this was confirmed by RNA studies. Gene disruption showed that the rad4 gene was essential for the mitotic growth of S. pombe.

Base Sequence↗

The mitochondrial genome of the bristletail Petrobius brevistylis (Archaeognatha: Machilidae).

The complete mitochondrial genome of the bristletail Petrobius brevistylis has been determined. The genome is 15,698 bp long and bears the standard set of genes common to all arthropods as well as a major non-coding A+T-rich region, the putative mitochondrial control region. A unique gene order was revealed as it differs from other hexapod and crustacean mitochondrial genomes in the position of tRNA-Tyr. Genome features like nucleotide composition and codon usage are compared with that of other insect taxa. A+T content is similar in species of Archaeognatha and Zygentoma, but obviously lower than in Collembola and Pterygota. This A+T bias significantly affects also amino acid frequencies and may be a problem for phylogenetic analyses.

Animals↗

The nucleotide sequence of the Mr = 28,500 flagellin gene of Caulobacter crescentus.

The DNA sequences which encode the Mr = 28,500 flagellin polypeptide of Caulobacter crescentus CB15 have been determined. The size of the protein, deduced from its DNA sequence (276 amino acids), is in agreement with its apparent molecular weight as measured by sodium dodecyl sulfate-polyacrylamide gel electrophoresis. The distribution of arginine residues within the protein sequence encoded by the gene correlates with their relative location as predicted by peptide alignment analysis (Gill, P.R., and Agabian, N. (1982) J. Bacteriol. 150, 925-933). DNA sequences 5' and 3' to the coding sequence were also determined. In the 5' region, DNA sequences homologous to consensus sequences associated with RNA polymerase recognition and transcription initiation sites in Escherichia coli (Pribnow box) are found. These are centered around 60, 90, and 120 base pairs upstream from the ATG codon at the beginning of the structural gene. Sequences 3' to the coding region were identified which might signal transcription termination. A typical E. coli 16 S ribosomal binding site (Shine-Dalgarno sequence) is located just 5' to the coding sequence, and for most of the amino acids there is a strong codon usage preference. Although this protein is exported from the cell (Gill, P.R., and Agabian, N. (1982) J. Bacteriol. 150, 925-933), the encoded NH2-terminal amino acid sequence is not different from the mature product.

Amino Acid Sequence↗

Constant and variable parts in the Balbiani ring 2 repeat unit and the translation termination region.

The large Balbiani rings of the chironomids produce giant internally repeated transcripts that are translated into silk-like proteins used for protective tubes. We have cloned fragments of the Balbiani ring 2 (BR2) gene of Chironomus pallidivittatus, normally the most prominent BR, for sequencing and restriction analysis. The results indicate a basic, tandemly arranged repeat unit of 198 bp, consisting of an invariant region of 102 bp followed by a variable region of 96 bp, the latter containing short internal tandem repeats. In the coding strand of both regions, there is a tendency to maximize adenine ( 40%) and minimize cytosine + thymidine (32-33%). Stringency in codon usage and absence of preference for third position nucleotide substitutions between homologous sequences suggest selection at the nucleotide level to be important. Both regions code for peptides rich in basic amino acids, but show distinct differences in other coding properties. The invariant region probably codes for a crystalline domain, and the variable region for a proline-rich, amorphous domain. One of the clones includes the 3' end of the translated region. Surrounding and following two stop codons is a sequence of four short palindromes. Furthermore, the first stop codon is part of a sequence reminiscent of a "TATA box". The possibilities are discussed that this area of the gene might be a target for regulatory molecules controlling translation termination and/or the expression of an overlapping cistron.

Journal Article↗

Codon optimization and mRNA amplification effectively enhances the immunogenicity of the hepatitis C virus nonstructural 3/4A gene.

We have recently shown that the NS3-based genetic immunogens should contain also hepatitis C virus (HCV) nonstructural (NS) 4A to utilize fully the immunogenicity of NS3. The next step was to try to enhance immunogenicity by modifying translation or mRNA synthesis. To enhance translation efficiency, a synthetic NS3/4A-based DNA (coNS3/4A-DNA) vaccine was generated in which the codon usage was optimized (co) for human cells. In a second approach, expression of the wild-type (wt) NS3/4A gene was enhanced by mRNA amplification using the Semliki forest virus (SFV) replicon (wtNS3/4A-SFV). Transient tranfections of human HepG2 cells showed that the coNS3/4A gene gave 11-fold higher levels of NS3 as compared to the wtNS3/4A gene when using the CMV promoter. We have previously shown that the presence of NS4A enhances the expression by SFV. Both codon optimization and mRNA amplification resulted in an improved immunogenicity as evidenced by higher levels of NS3-specific antibodies. This improved immunogenicity also resulted in a more rapid priming of cytotoxic T lymphocytes (CTLs). Since HCV is a noncytolytic virus, the functionality of the primed CTL responses was evaluated by an in vivo challenge with NS3/4A-expressing syngeneic tumor cells. The priming of a tumor protective immunity required an endogenous production of the immunogen and CD8+ CTLs, but was independent of B and CD4+ T cells. This model confirmed the more rapid in vivo activation of an NS3/4A-specific tumor-inhibiting immunity by codon optimization and mRNA amplification. Finally, therapeutic vaccination with the coNS3/4A gene using gene gun 6-12 days after injection of tumors significantly reduced the tumor growth in vivo. Codon optimization and mRNA amplification effectively enhances the overall immunogenicity of NS3/4A. Thus, either, or both, of these approaches should be utilized in an NS3/4A-based HCV genetic vaccine.

Animals↗

Factors regulating cryIVB expression in the cyanobacterium--Synechococcus PCC 7942.

The expression of the larvicidal Bacillus thuringiensis subsp. israelensis cryIVB gene in cyanobacteria has been suggested to be an effective means of controlling mosquito populations. Using a variety of cryIVB constructs, in this study we have examined the effect of Synechococcus PCC 7942 culture age on intracellular toxin levels and have attempted to determine the mechanisms by which cryIVB gene expression is regulated. The data suggest that specific degradation of the cryIVB mRNA limits toxin production; however, the addition of cyanobacterial 3' untranslated DNA sequences to the cryIVB gene did not improve mRNA stability or toxin levels. An analysis of the cryIVB sequence and comparison of codon usage patterns with highly expressed cyanobacterial genes suggest that inefficient translation and intragenic ribosomal binding sites impede protein synthesis and result in rapid turnover of the toxin mRNA.

Bacillus thuringiensis↗

Truncation of limonene synthase preprotein provides a fully active 'pseudomature' form of this monoterpene cyclase and reveals the function of the amino-terminal arginine pair.

The monoterpene cyclase limonene synthase transforms geranyl diphosphate to a monocyclic olefin and constitutes the simplest model for terpenoid cyclase catalysis. (-)-4S-Limonene synthase preprotein from spearmint bears a long plastidial targeting sequence. Difficulty expressing the full-length preprotein in Escherichia coli is encountered because of host codon usage, inclusion body formation, and the tight association of bacterial chaperones with the transit peptide. The purified preprotein is also kinetically impaired relative to the mixture of N-blocked native proteins produced in vivo by proteolytic processing in plastids. Therefore, the targeting sequence, that precedes a tandem pair of arginines (R58R59) which is highly conserved in the monoterpene synthases, was removed. Expression of this truncated protein, from a vector that encodes a tRNA for two rare arginine codons (pSBET), affords a soluble, tractable 'pseudomature' form of the enzyme that is catalytically more efficient than the native species. Truncation up to and including R58, or substitution of R59, yields enzymes that are incapable of converting the natural substrate geranyl diphosphate, via the enzymatically formed tertiary allylic isomer 3S-linalyl diphosphate, to (-)-limonene. However, these enzymes are able to cyclize exogenously supplied 3S-linalyl diphosphate to the olefinic product. This result indicates a role for the tandem arginines in the unique diphosphate migration step accompanying formation of the intermediate 3S-linalyl diphosphate and preceding the final cyclization reaction catalyzed by the monoterpene synthases. The structural basis for this coupled isomerization-cyclization reaction sequence can be inferred by homology modeling of (-)-4S-limonene synthase based on the three-dimensional structure of the sesquiterpene cyclase epi-aristolochene synthase [Starks, C. M., Back, K., Chappell, J., and Noel, J. P. (1997) Science 277, 1815-1820].

Amino Acid Sequence↗