Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Codon optimization improves heterologous expression of a Schistosoma mansoni cDNA in HEK293 cells.

Differences in codon usage can seriously hamper the expression of cloned cDNAs in heterologous systems. In this study, we show that the expression of a cloned Schistosoma mansoni cDNA in cultured HEK293 cells was dramatically increased by rewriting a portion of the cDNA according to human preferred codon usage, suggesting that codon optimization is a valuable strategy for improving the heterologous expression of helminth sequences. We further describe a simple modification of a recursive PCR-based method, which allows the rewriting of long stretches of DNA sequence in a single PCR reaction. This method can be used to optimize the codon usage of virtually any DNA from helminths and other parasites.

Animals↗

Influence of the codon following the initiation codon on the expression of the lacZ gene in Saccharomyces cerevisiae.

A set of 32 different codons were introduced in a lacZ expression vector (pPTK400) immediately 3' from the AUG initiation codon. Expression of the lacZ gene was determined in Saccharomyces cerevisiae by measuring the amount of beta-galactosidase fusion protein using immuno-gel electrophoresis. A 5.3-fold difference in expression was found among the various constructs. It was found that there was no preference for a certain nucleotide in any position of the second codon and there was no distinct correlation between the level of tRNA corresponding to any particular second codon and expression. No correlation could be found between the local secondary structure and expression. When the overall codon usage in yeast and the codon usage in the second position of the mRNA is compared, there is no obvious significant difference in preference. This indicates that in yeast, in contrast to Escherichia coli, the codon choice at the beginning of the mRNA does not deviate from the one further downstream and is determined by the requirements for optimal translation elongation. Important determinants of the optimal context for an initiation codon in yeast therefore must be located mainly 5' from this codon.

Amino Acid Sequence↗

Genetic plasticity of V genes under somatic hypermutation: statistical analyses using a new resampling-based methodology.

Evidence for somatic hypermutation of immunoglobulin genes has been observed in all of the species in which immunoglobulins have been found. Previous studies have suggested that codon usage in immunoglobulin variable (V) region genes is such that the sequence-specificity of somatic hypermutation results in greater mutability in complementarity-determining regions of the gene than in the framework regions. We have developed a new resampling-based methodology to explore genetic plasticity in individual V genes and in V gene families in a statistically meaningful way. We determine what factors contribute to this mutability difference and characterize the strength of selection for this effect. We find that although the codon usage in immunoglobulin V genes renders them distinct among translationally equivalent sequences with random codon usage, they are nevertheless not optimal in this regard. We find that the mutability patterns in a number of species are similar to those we find for human sequences. Interestingly, sheep sequences show extremely strong mutability differences, consistent with the role of somatic hypermutation in the diversification of primary antibody repertoire in these animals. Human TCR V(beta) sequences resemble immunoglobulin in mutability pattern, suggesting one of several alternatives, that hypermutation is functionally operating in TCR, that it was once operating in TCR or in the common precursor of TCR and immunoglobulin, or that the hypermutation mechanism has evolved to exploit the codon usage in immunoglobulin (and fortuitously, TCR) rather than vice-versa. Our findings provide support to the hypothesis that somatic hypermutation appeared very early in the phylogeny of immune systems, that it is, to a large extent, shared between species, and that it makes an essential contribution to the generation of the antibody repertoire.

Base Sequence↗

A nucleotide polymorphism in ERCC1 in human ovarian cancer cell lines and tumor tissues.

We studied the DNA sequence of the entire coding region of ERCC1 gene, in five cell lines established from human ovarian cancer (A2780, A2780/CP70, MCAS, OVCAR-3, SK-OV-3), 29 human ovarian cancer tumor tissue specimens, one human T-lymphocyte cell line (H9), and non-malignant human ovary tissue (NHO). Samples were assayed by PCR-SSCP and DNA sequence analyses. A silent mutation at codon 118 (site for restriction endonuclease MaeII) in exon 4 of the gene was detected in MCAS, OVCAR-3 and SK-OV-3 cells, and NHO. This mutation was a C-->T transition, that codes for the same amino acid: asparagine. This transition converts a common codon usage (AAC) to an infrequent codon usage (AAT), whereas frequency of use is reduced two-fold. This base change was associated with a detectable band shift on SSCP analysis. For the 29 ovarian cancer specimens, the same base change was observed in 15 tumor samples and was associated with the same band shift in exon 4. Cells and tumor tissue specimens that did not contain the C-->T transition, did not show the band shift in exon 4. Our data suggest that this alteration at codon 118 within the ERCC1 gene, may exist in platinum-sensitive and platinum-resistant ovarian cancer tissues.

Antineoplastic Agents↗

Interaction of silent and replacement changes in eukaryotic coding sequences.

We examined the codon usages in well-conserved and less-well-conserved regions of vertebrate protein genes and found them to be similar. Despite this similarity, there is a statistically significant decrease in codon bias in the less-well-conserved regions. Our analysis suggests that although those codon changes initially fixed under amino acid replacements tend to follow the overall codon usage pattern, they also reduce the bias in codon usage. This decrease in codon bias leads one to predict that the rate of change of synonymous codons should be greater in those regions that are less well conserved at the amino acid level than in the better-conserved regions. Our analysis supports this prediction. Furthermore, we demonstrate a significantly elevated rate of change of synonymous codons among the adjacent codons 5' to amino acid replacement positions. This provides further support for the idea that there are contextual constraints on the choice of synonymous codons in eukaryotes.

Cell Physiological Phenomena↗

On the origin of Ser/Thr kinases in a prokaryote.

The family of Ser/Thr and/or Tyr kinases and that of His kinases play essential roles in signal transduction. For a long time, the former has been found in eukaryotes, the latter in prokaryotes. Studies in the last decade have shown, however, that most bacteria possess from one to more than 10 genes encoding Ser/Thr kinases. This observation raises an important question concerning the evolutionary origin of Ser/Thr kinases found in bacteria. To answer this question, we have analyzed a family of 11 genes encoding Ser/Thr kinases in the cyanobacterium Synechocystis sp. PCC 6803. This bacterium contains the largest number of Ser/Thr kinases among all bacteria whose genomic sequences have been released so far. In this study, we have developed a user-friendly computer program for statistical analysis of codon usages and GC content. The results demonstrate that Ser/Thr kinases have similar codon usages and GC contents as the average of all possible open reading frames (ORFs) deduced from the genome. In contrast, ORFs encoding transposases, as a control in our analysis, display a disparity in both codon usage and GC content, confirming their multiple origin and genetic promiscuity. In light of our results, we propose that Ser/Thr kinases existed before the divergence between prokaryotes and eukaryotes during evolution, or were laterally transferred into prokaryotes at the early stages of bacterial evolution. If Ser/Thr kinases have persisted ever since in prokaryotes under evolutionary pressure, it is then expected that they play important, possibly even essential roles in regulating bacterial activities as do their counterparts in eukaryotes.

Base Composition↗

Poly(3-hydroxyvalerate) depolymerase of Pseudomonas lemoignei.

Pseudomonas lemoignei is equipped with at least five polyhydroxyalkanoate (PHA) depolymerase structural genes (phaZ1 to phaZ5) which enable the bacterium to utilize extracellular poly(3-hydroxybutyrate) (PHB), poly(3-hydroxyvalerate) (PHV), and related polyesters consisting of short-chain-length hxdroxyalkanoates (PHA(SCL)) as the sole sources of carbon and energy. Four genes (phaZ1, phaZ2, phaZ3, and phaZ5) encode PHB depolymerases C, B, D, and A, respectively. It was speculated that the remaining gene, phaZ4, encodes the PHV depolymerase (D. Jendrossek, A. Frisse, A. Behrends, M. Andermann, H. D. Kratzin, T. Stanislawski, and H. G. Schlegel, J. Bacteriol. 177:596-607, 1995). However, in this study, we show that phaZ4 codes for another PHB depolymeraes (i) by disagreement of 5 out of 41 amino acids that had been determined by Edman degradation of the PHV depolymerase and of four endoproteinase GluC-generated internal peptides with the DNA-deduced sequence of phaZ4, (ii) by the lack of immunological reaction of purified recombinant PhaZ4 with PHV depolymerase-specific antibodies, and (iii) by the low activity of the PhaZ4 depolymerase with PHV as a substrate. The true PHV depolymerase-encoding structural gene, phaZ6, was identified by screening a genomic library of P. lemoignei in Escherichia coli for clearing zone formation on PHV agar. The DNA sequence of phaZ6 contained all 41 amino acids of the GluC-generated peptide fragments of the PHV depolymerase. PhaZ6 was expressed and purified from recombinant E. coli and showed immunological identity to the wild-type PHV depolymerase and had high specific activities with PHB and PHV as substrates. To our knowledge, this is the first report on a PHA(SCL) depolymerase gene that is expressed during growth on PHV or odd-numbered carbon sources and that encodes a protein with high PHV depolymerase activity. Amino acid analysis revealed that PhaZ6 (relative molecular mass [M(r)], 43,610 Da) resembles precursors of other extracellular PHA(SCL) depolymerases (28 to 50% identical amino acids). The mature protein (M(r), 41,048) is composed of (i) a large catalytic domain including a catalytic triad of S(136), D(211), and H(269) similar to serine hydrolases; (ii) a linker region highly enriched in threonine residues and other amino acids with hydroxylated or small side chains (Thr-rich region); and (iii) a C-terminal domain similar in sequence to the substrate-binding domain of PHA(SCL) depolymerases. Differences in the codon usage of phaZ6 for some codons from the average codon usage of P. lemoignei indicated that phaZ6 might be derived from other organisms by gene transfer. Multialignment of separate domains of bacterial PHA(SCL) depolymerases suggested that not only complete depolymerase genes but also individual domains might have been exchanged between bacteria during evolution of PHA(SCL) depolymerases.

Acyltransferases↗

Codon catalog usage and the genome hypothesis.

Frequencies for each of the 61 amino acid codons have been determined in every published mRNA sequence of 50 or more codons. The frequencies are shown for each kind of genome and for each individual gene. A surprising consistency of choices exists among genes of the same or similar genomes. Thus each genome, or kind of genome, appears to possess a "system" for choosing between codons. Frameshift genes, however, have widely different choice strategies from normal genes. Our work indicates that the main factors distinguishing between mRNA sequences relate to choices among degenerate bases. These systematic third base choices can therefore be used to establish a new kind of genetic distance, which reflects differences in coding strategy. The choice patterns we find seem compatible with the idea that the genome and not the individual gene is the unit of selection. Each gene in a genome tends to conform to its species' usage of the codon catalog; this is our genome hypothesis.

Animals↗

Origin and evolution of genes specifying resistance to macrolide, lincosamide and streptogramin antibiotics: data and hypotheses.

Resistance to macrolide, lincosamide and streptogramin antibiotics is due to alteration of the target site or detoxification of the antibiotic. Postranscriptional methylation of 23S ribosomal rRNA confers resistance to macrolide (M), lincosamide (L) and streptogramin (S) B-type antibiotics, the so-called MLSB phenotype. Several classes of rRNA methylases conferring resistance to MLSB antibiotics have been characterized in Gram-positive cocci, in Bacillus spp, and in strains of actinomycetes producing erythromycin. The enzymes catalyze N6-dimethylation of an adenine residue situated in a highly conserved region of prokaryotic 23S rRNA. In this review, we compare the amino acid sequences of the rRNA methylases and analyze the codon usage in the corresponding erm (erythromycin resistance methylase) genes. The homology detected at the protein level is consistent with the notion that an ancestor of the erm genes was implicated in erythromycin resistance in a producing strain. However, the rRNA methylases of producers and non-producers present substantial sequence diversity. In Gram-positive bacteria the preferential codon usage in the erm genes reflects the guanosine plus cytosine content of the chromosome of the host. These observations suggest that the presence of erm genes in these micro-organisms is ancient. By contrast, it would appear that enterobacteria have acquired only recently an rRNA methylase gene of the ermB class from a Gram-positive coccus since the genes isolated in Escherichia coli and in Gram-positive cocci are highly homologous (homology greater than 98%) and present a codon usage typical of the latter micro-organisms. As opposed to the MLSB phenotype which results from a single biochemical mechanism, inactivation of structurally related antibiotics of the MLS group involves synthesis of various other enzymes. In enterobacteria, resistance to erythromycin and oleandomycin is due to production of erythromycin esterases which hydrolyze the lactone ring of the 14-membered macrolides. We recently reported the nucleotide sequence of ereA and ereB (erythromycin resistance esterase) genes which encode erythromycin esterases type I and II, respectively. The amino acid sequences of the two isozymes do not exhibit statistically significant homology. Analysis of codon usage in both genes suggests that esterase type I is indigenous to E. coli, whereas the type II enzyme was acquired by E. coli from a phylogenetically remote micro-organism. Inactivation of lincosamides, first reported in staphylococci and lactobacilli of animal origin, was also recently detected in Gram-positive cocci isolated from humans.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence↗

The correlation between synonymous and nonsynonymous substitutions in Drosophila: mutation, selection or relaxed constraints?

Codon usage bias, the preferential use of particular codons within each codon family, is characteristic of synonymous base composition in many species, including Drosophila, yeast, and many bacteria. Preferential usage of particular codons in these species is maintained by natural selection acting largely at the level of translation. In Drosophila, as in bacteria, the rate of synonymous substitution per site is negatively correlated with the degree of codon usage bias, indicating stronger selection on codon usage in genes with high codon bias than in genes with low codon bias. Surprisingly, in these organisms, as well as in mammals, the rate of synonymous substitution is also positively correlated with the rate of nonsynonymous substitution. To investigate this correlation, we carried out a phylogenetic analysis of substitutions in 22 genes between two species of Drosophila, Drosophila pseudoobscura and D. subobscura, in codons that differ by one replacement and one synonymous change. We provide evidence for a relative excess of double substitutions in the same species lineage that cannot be explained by the simultaneous mutation of two adjacent bases. The synonymous changes in these codons also cannot be explained by a shift to a more preferred codon following a replacement substitution. We, therefore, interpret the excess of double codon substitutions within a lineage as being the result of relaxed constraints on both kinds of substitutions in particular codons.

Animals↗

Statistical method for predicting protein coding regions in nucleic acid sequences.

Protein coding regions of a genome fragment can be mathematically predicted by studying variations in the statistical properties or by searching the signals characteristic of the junctions between the coding and non-coding regions. We propose here a new statistical method using correspondence analysis. This method does not use any reference codon set but takes into account the codon usage homogeneity along the studied genome fragment. Comparison with previously published methods especially the 'codon usage method' of Staden has been made, and two examples are presented here. Applications to analysis of prokaryotic operon and eukaryotic split genes are also discussed. Use of the method has also shown two structures not previously described: i) in the human prt gene, a strong triplet structure exists in a non-coding region; ii) in the human tp-a codon usage is not uniform between the different exons.

Algorithms↗

Directional mutation pressure and transfer RNA in choice of the third nucleotide of synonymous two-codon sets.

Bacterial species have diverged into a series of families, some with high G + C content in their DNA, and other with high A + T content, resulting, respectively, from G.C- and A.T-directional mutation pressures. Such mutation pressure (G.C/A.T pressure) may be an important determinant for codon usage. It has also been suggested that tRNA acts as a selective constraint for determining codon usage. We have studied the relation between G.C/A.T pressure and tRNA constraints in determining choice of the third nucleotide of eight two-codon sets, using codon usage data obtained from protein genes in four bacterial species, Mycoplasma capricolum, Bacillus subtilis, Escherichia coli, and Micrococcus luteus, and in liverwort (Marchantia polymorpha) chloroplasts. The genomic G + C contents of these range from 25% to 74%. The results demonstrate that tRNA levels act additively to A.T and G.C pressure in affecting contents of A (pairing with *UNN anticodons, in which *U indicates a 2-thiouridine derivative) and C (pairing with GNN anticodons) or G (pairing with CNN anticodons), respectively, in third nucleotide positions of codons.

Chloroplasts↗

Compositional properties of nuclear genes from Plasmodium falciparum.

We have analyzed the compositional distributions of coding sequences and their different codon positions, as well as the codon usage of the nuclear genes of Plasmodium falciparum, a parasite characterized by an extremely GC-poor genome. As expected, coding sequences are AT-rich, codon usage is strongly biased towards A or T in third codon positions, and some particular amino acids (aa) are especially abundant in the encoded proteins. Remarkably, however, no difference was detected between housekeeping (HK) and antigen (Ag) genes, in spite of differences in expression level and evolutionary constraints. Moreover, all the features found in P. falciparum are very similar to those found in a bacterium characterized by a very GC-poor genome, Staphylococcus aureus. These findings stress the importance of compositional constraints in determining codon usage and aa utilisation.

Amino Acids↗

Comparison of three actin-coding sequences in the mouse; evolutionary relationships between the actin genes of warm-blooded vertebrates.

We have determined the sequences of three recombinant cDNAs complementary to different mouse actin mRNAs that contain more than 90% of the coding sequences and complete or partial 3' untranslated regions (3'UTRs): pAM 91, complementary to the actin mRNA expressed in adult skeletal muscle (alpha sk actin); pAF 81, complementary to an actin mRNA that is accumulated in fetal skeletal muscle and is the major transcript in adult cardiac muscle (alpha c actin); and pAL 41, identified as complementary to a beta nonmuscle actin mRNA on the basis of its 3'UTR sequence. As in other species, the protein sequences of these isoforms are highly (greater than 93%) conserved, but the three mRNAs show significant divergence (13.8-16.5%) at silent nucleotide positions in their coding regions. A nucleotide region located toward the 5' end shows significantly less divergence (5.6-8.7%) among the three mouse actin mRNAs; a second region, near the 3' end, also shows less divergence (6.9%), in this case between the mouse beta and alpha sk actin mRNAs. We propose that recombinational events between actin sequences may have homogenized these regions. Such events distort the calculated evolutionary distances between sequences within a species. Codon usage in the three actin mRNAs is clearly different, and indicates that there is no strict relation between the tissue type, and hence the tRNA precursor pool, and codon usage in these and other muscle mRNAs examined. Analysis of codon usage in these coding sequences in different vertebrate species indicates two tendencies: increases in bias toward the use of G and C in the third codon position in paralogous comparisons (in the order alpha c less than beta less than alpha sk), and in orthologous comparisons (in the order chicken less than rodent less than man). Comparison of actin-coding sequences between species was carried out using the Perler method of analysis. As one moves backward in time, changes at silent sites first accumulate rapidly, then begin to saturate after -(30-40) million years (MY), and actually decrease between -400 and -500 MY. Replacements or silent substitutions therefore cannot be used as evolutionary clocks for these sequences over long periods. Other phenomena, such as gene conversion or isochore compartmentalization, probably distort the estimated divergence time.

Actins↗

Nucleotide sequence of the structural gene for tryptophanase of Escherichia coli K-12.

The tryptophanase structural gene, tnaA, of Escherichia coli K-12 was cloned and sequenced. The size, amino acid composition, and sequence of the protein predicted from the nucleotide sequence agree with protein structure data previously acquired by others for the tryptophanase of E. coli B. Physiological data indicated that the region controlling expression of tnaA was present in the cloned segment. Sequence data suggested that a second structural gene of unknown function was located distal to tnaA and may be in the same operon. The pattern of codon usage in tnaA was intermediate between codon usage in four of the ribosomal protein structural genes and the structural genes for three of the tryptophan biosynthetic proteins.

Amino Acid Sequence↗

Selective differences among translation termination codons.

The frequency of use of the three alternative translation termination codons has been examined in 165 Escherichia coli, 52 Bacillus subtilis and 106 Saccharomyces cerevisiae genes. Genes were first categorised according to their degree of bias in sense codon usage. In each species there is a very strong bias in favour of UAA (over UAG and UGA) in genes where sense codon usage is highly biased. This bias declines, principally with an increase in the use of UGA, in genes with lower sense codon bias. It appears that selection operating during translation may maintain the bias in stop codon usage. Such selection could result from the greater availability of UAA-cognate release factor(s), or from a lower frequency of translational readthrough at UAA.

Bacillus subtilis↗

Comparison and evolutionary analysis of the glycosomal glyceraldehyde-3-phosphate dehydrogenase from different Kinetoplastida.

In this work, we present the sequences and a comparison of the glycosomal GAPDHs from a number of Kinetoplastida. The complete gene sequences have been determined for some species (Crithidia fasciculata, Herpetomonas samuelpessoai, Leptomonas seymouri, and Phytomonas sp), whereas for other species (Trypanosoma brucei gambiense, Trypanosoma congolense, Trypanosoma vivax, and Leishmania major), only partial sequences have been obtained by PCR amplification. The structure of all available glycosomal GAPDH genes was analyzed in detail. Considerable variations were observed in both their nucleotide composition and their codon usage. The GC content varies between 64.4% in L. seymouri and 49.5% in the previously sequenced GAPDH gene from Trypanoplasma borreli. A highly biased codon usage was found in C. fasciculata, with only 34 triplets used, whereas in T. borreli 57 codons were employed. No obvious correlation could be observed between the codon usage and either the nucleotide composition or the level of gene expression. The glycosomal GAPDH is a very well-conserved enzyme. The maximal overall difference observed in the amino acid sequences is only 25%. Specific insertions and extensions are retained in all sequences. The residues involved in catalysis, substrate, and inorganic phosphate binding are fully conserved, whereas some variability is observed in the cofactor-binding pocket. The implications of these data for the design of new trypanocidal drugs targeted against GAPDH are discussed. All available gene and amino acid sequences of glycosomal GAPDHs were used for a phylogenetic analysis. The division of the Kinetoplastida into two suborders, Bodonina and Trypanosomatina, was well supported. Within the letter group, the Trypanosoma species appeared to be monophyletic, whereas the other trypanosomatids form a second clade.

Amino Acid Sequence↗

Expression of the green fluorescent protein in Paramecium tetraurelia.

In this paper we describe the expression of green fluorescent protein (GFP) as a reporter in vivo to monitor transformation in Paramecium cells. This is not trivial because of the limited number of strong promoters available for heterologous expression and the very high AT content of the genomic DNA, the consequence of which is a very aberrant codon usage. Taking into account differences in codon usage we selected and modified the original GFP open reading frame (ORF) from Aequorea victoria and placed the altered ORF into the Paramecium expression vector pPXV. Injection of the linearized plasmid into the macronucleus resulted in a cytoplasmic fluorescence signal in the clonal descendants, which was proportional to the number of copies injected. Southern hybridization indicated the establishment and replication of the plasmid during vegetative growth. Expression was also monitored by Northern and Western analysis. The results indicate that the modified GFP can be used in Paramecium as a reporter for transformation as an alternative to selection with antibiotics and that it may also be used to construct and localize fusion proteins.

Animals↗