Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

A plasmid system for optimization of Fab' production in Escherichia coli: importance of balance of heavy chain and light chain synthesis.

We demonstrate the importance of optimizing the balance of light chain (LC) and heavy chain (HC) expression to achieve high level production of Fab' fragments in the Escherichia coli periplasm. The LC:HC balance has been controlled by varying the codon usage of the signal peptide (SP) and 5' mature domain coding regions. Different SP coding regions have been identified from a codon wobble-based library using alkaline phosphatase (AP) as a reporter gene. A plasmid system that enables random combination of these variant SP coding regions is used to construct optimized Fab' expression plasmids. These small plasmid libraries facilitated selection of optimal Fab' expression plasmids and resulted in increases of periplasmic yield, up to 580 mgL(-1) from E. coli fermentations and will enable rapid variable region subcloning and selection of future Fab(') expression plasmids.

Base Sequence↗

A cluster of vitellogenin genes in the Mediterranean fruit fly Ceratitis capitata: sequence and structural conservation in dipteran yolk proteins and their genes.

Four genes encoding the major egg yolk polypeptides of the Mediterranean fruit fly Ceratitis capitata, vitellogenins 1 and 2 (VG1 and VG2), were cloned, characterized and partially sequenced. The genes are located on the same region of chromosome 5 and are organized in pairs, each encoding the two polypeptides on opposite DNA strands. Restriction and nucleotide sequence analysis indicate that the gene pairs have arisen from an ancestral pair by a relatively recent duplication event. The transcribed part is very similar to that of the Drosophila melanogaster yolk protein genes Yp1, Yp2 and Yp3. The Vg1 genes have two introns at the same positions as those in D. melanogaster Yp3; the Vg2 genes have only one of the introns, as do D. melanogaster Yp1 and Yp2. Comparison of the five polypeptide sequences shows extensive homology, with 27% of the residues being invariable. The sequence similarity of the processed proteins extends in two regions separated by a nonconserved region of varying size. Secondary structure predictions suggest a highly conserved secondary structure pattern in the two regions, which probably correspond to structural and functional domains. The carboxy-end domain of the C. capitata proteins shows the same sequence similarities with triacyglycerol lipases that have been reported previously for the D. melanogaster yolk proteins. Analysis of codon usage shows significant differences between D. melanogaster and C. capitata vitellogenins with the latter exhibiting a less biased representation of synonymous codons.

Alleles↗

Crystallin gene expression during rat lens development.

The analysis of the developmental pattern of the alpha A-, alpha B-, beta B1-, beta B2-, beta B3-, beta A3/A1-, and beta s-crystallin genes during fetal and postnatal development of the rat shows that the differential regulation of crystallin synthesis relies on differential gene shutdown rather than differential gene activation; that is, all crystallin genes are active during early development but turn off at different stages. The only two exceptions to this rule are the alpha B- and beta s-crystallin genes. The alpha B-crystallin gene transcript becomes first detectable at 18 days of fetal development, while the beta s-crystallin gene appears to be active only in the postnatal period. We also determined the absolute numbers of the alpha A-, alpha B-, beta B1-, beta B2-, beta B3-, beta A3/A1-, beta s-, and gamma-crystallin gene transcripts present in the lens at various times after birth. Comparison of these RNA data with the published protein data shows that the alpha B- and beta B2-crystallin RNAs are relatively overrepresented, suggesting the possibility that these two RNA species are not used as efficiently as other crystallin mRNAs. Examination of the known (hamster) alpha B-crystallin sequence and elucidation of the (rat) beta B2-crystallin sequence yielded no evidence for aberrant codon usage. These two RNAs have one sequence motif in common: they are the only crystallin mRNAs in which the translation initiation codon is preceded by CCACC.

Age Factors↗

Glutamyl-tRNA synthetase from Thermus thermophilus HB8. Molecular cloning of the gltX gene and crystallization of the overproduced protein.

The gene for the Glu-tRNA synthetase from an extreme thermophile, Thermus thermophilus HB8, was isolated using a synthetic oligonucleotide probe coding for the N-terminal amino acid sequence of Glu-tRNA synthetase. Nucleotide-sequence analysis revealed an open reading frame coding for a protein composed of 468 amino acid residues (Mr 53,901). Codon usage in the T. thermophilus Glu-tRNA synthetase gene was in fact similar to the characteristic usages in the genes for proteins from bacteria of genus Thermus: the G + C content in the third position of the codons was as high as 94%. In contrast, the amino acid sequence of T. thermophilus Glu-tRNA synthetase showed high similarity with bacterial Glu-tRNA synthetases (35-45% identity); the sequences of the binding sites for ATP and for the 3' terminus of tRNA(Glu) are highly conserved. The Glu-tRNA synthetase gene was efficiently expressed in Escherichia coli under the control of the tac promoter. The recombinant T. thermophilus Glu-tRNA synthetase was extremely thermostable and was purified to homogeneity by heat treatment and three-step column chromatography. Single crystals of T. thermophilus Glu-tRNA synthetase were obtained from poly(ethylene glycol) 6000 solution by a vapor-diffusion technique. The crystals diffract X-rays beyond 0.35 nm. The crystal belongs to the orthorhombic space group P2(1)2(1)2(1), with unit-cell parameters of a = 8.64 nm, b = 8.86 nm and c = 8.49 nm.

Adenosine Triphosphate↗

Effects of a minor isoleucyl tRNA on heterologous protein translation in Escherichia coli.

In Escherichia coli, the isoleucine codon AUA occurs at a frequency of about 0.4% and is the fifth rarest codon in E. coli mRNA. Since there is a correlation between the frequency of codon usage and the level of its cognate tRNA, translational problems might be expected when the mRNA contains high levels of AUA codons. When a hemagglutinin from the influenza virus, a 304-amino-acid protein with 12 (3.9%) AUA codons and 1 tandem codon, and a mupirocin-resistant isoleucyl tRNA synthetase, a 1,024-amino-acid protein, with 33 (3.2%) AUA codons and 2 tandem codons, were expressed in E. coli, product accumulation was highly variable and dependent to some degree on the growth medium. In rich medium, the flu antigen represented about 16% of total cell protein, whereas in minimal medium, it was only 2 to 3% of total cell protein. In the presence of the cloned ileX, which encodes the cognate tRNA for AUA, however, the antigen was 25 to 30% of total cell protein in cells grown in minimal medium. Alternatively, the isoleucyl tRNA synthetase did not accumulate to detectable levels in cells grown in Luria broth unless the ileX tRNA was coexpressed when it accounted for 7 to 9% of total cell protein. These results indicate that the rare isoleucine AUA codon, like the rare arginine codons AGG and AGA, can interfere with the efficient expression of cloned proteins.

Bacterial Proteins↗

Cloning and analysis of the DNA polymerase-encoding gene from Thermus caldophilus GK24.

The gene encoding Thermus caldophilus GK24 (Tca) DNA polymerase was cloned into Escherichia coli using the structural gene coding for Thermus aquaticus YT-1 (Taq) DNA polymerase as a hybridization probe. The nucleotide sequence of the cloned DNA was determined. The primary structure of the Tca DNA polymerase was deduced from the nucleotide sequence. The Tca DNA polymerase comprised 834 amino acid residues and its molecular mass was determined to be 93,810. On alignment of the whole amino acid sequence, Tca DNA polymerase showed a high sequence homology with the E. coli DNA polymerase I-like DNA polymerases, and 86% identity with Taq DNA polymerase, 38% with E. coli and Streptococcus pneumoniae (Spn) DNA polymerase I. An extremely high sequence identity was observed in the region containing the polymerase activity. The codon usage in the Tca DNA polymerase gene was in fact similar to the characteristic usages in the genes for proteins from bacteria of genus Thermus: the G+C content in the third position of the codons was as high as 93%. The Tca DNA polymerase gene was expressed under the control of tac promoter on a high copy plasmid, pTCA in E. coli.

Amino Acid Sequence↗

Nucleotide sequence of regions homologous to nifH (nitrogenase Fe protein) from the nitrogen-fixing archaebacteria Methanococcus thermolithotrophicus and Methanobacterium ivanovii: evolutionary implications.

DNA fragments bearing sequence similarity to eubacterial nif H probes were cloned from two nitrogen-fixing archaebacteria, a thermophilic methanogen, Methanococcus (Mc.) thermolithotrophicus, and a mesophilic methanogen, Methanobacterium (Mb.) ivanovii. Regions carrying similarities with the probes were sequenced. They contained several open reading frames (ORF), separated by A + T-rich regions. The largest ORFs in both regions, an 876-bp sequence in Mc. thermolithotrophicus and a 789-bp sequence in Mb. ivanovii, were assumed to be ORFsnif H. They code for polypeptides of mol. wt. 32,025 and 28,347, respectively. Both ORFsnifH were preceded by potential ribosome binding sites and followed by potential hairpin structures and by oligo-T sequences, which may act as transcription termination signals. The codon usage was similar in both ORFsnifH and was analogous to that used in the Clostridium pasteurianum nifH gene, with a preference for codons ending with A or U. The ORFnifH deduced polypeptides contained 30% sequence matches with all eubacterial nifH products already sequenced. Four cysteine residues were found at the same position in all sequences, and regions surrounding the cysteine residues are highly conserved. Comparison of all pairs of methanogenic and eubacterial nifH sequences is in agreement with a distant phylogenetic position of archaebacteria and with a very ancient origin of nif genes. However, sequence similarity between Methanobacteriales and Methanococcales is low (around 50%) as compared to that found among eubacteria, suggesting a profound divergence between the two orders of methanogens. From comparison of amino acid sequences, C. pasteurianum groups with the other eubacteria, whereas comparison of nucleotide sequences seems to bring C. pasteurianum closer to methanogens. The latter result may be due to the high A + T content of both C. pasteurianum and methanogens ORFsnif H or may come from an ancient lateral transfer between Clostridium and methanogens.

Amino Acid Sequence↗

A comparison of homologous developmental genes from Drosophila and Tribolium reveals major differences in length and trinucleotide repeat content.

The flour beetle Tribolium castaneum has become an important model organism for comparative studies of insect development. Many developmentally important genes have now been cloned from both Tribolium and Drosophila and their expression characteristics were studied. We analyze here the complete coding sequences of 17 homologous gene pairs from D. melanogaster and T. castaneum, most of which encode transcription factors. We find that the Tribolium genes are on average 30% shorter than their Drosophila homologues. This appears to be due largely to the almost-complete absence of trinucleotide repeats in the coding sequences of Tribolium as well as the generally lower degree of internal repetitiveness. Clusters of polar and other amino acids such as glutamine, proline, and serine, which are often considered to be important for transcriptional activation domains in Drosophila, are almost completely absent in Tribolium. Codon usage is generally less biased in Tribolium, although we find a similar tendency for the preference of G- or C-ending codons and a higher bias in conserved subregions of the proteins as in Drosophila. Most of the aminoacid substitutions in the DNA-binding domains of the transcription factors occur at residues that do not make a specific contact to DNA, suggesting that the recognition sequences are likely to be conserved between the two species.

Amino Acid Sequence↗

Codon bias evolution in Drosophila. Population genetics of mutation-selection drift.

Although non-random patterns of synonymous codon usage are a prominent feature in the genomes of many organisms, the relatives roles of mutational biases and natural selection in maintaining codon bias remain a contentious issue. In some species, patterns of codon bias and empirical findings on the biology of translation suggest 'major codon preference', a balance among mutation pressure, genetic drift, and weak selection in favor of translationally superior codons. Population genetics theory makes testable predictions to distinguish such a model from a strictly mutational model of codon bias. Major codon preference predicts two fitness classes of synonymous DNA changes: 'preferred' mutations from non-major to major codons and 'unpreferred' changes in the opposite direction. An extension of current statistical methods is employed to reveal differences in the within and between species dynamics of preferred and unpreferred silent mutations in Drosophila simulans. In this lineage, codon bias appears to be maintained under roughly equal magnitudes of natural selection and genetic drift. In the sibling species, D. melanogaster, however, a reduction in N(e)s, the product of effective population size and selection coefficient, appears to have allowed a genome-wide reduction in codon bias.

Animals↗

DNAskew: statistical analysis of base compositional asymmetry and prediction of replication boundaries in the genome sequences.

Sueoka and Lobry declared respectively that, in the absence of bias between the two DNA strands for mutation and selection, the base composition within each strand should be A=T and C=G (this state is called Parity Rule type 2, PR2). However, the genome sequences of many bacteria, vertebrates and viruses showed asymmetries in base composition and gene direction. To determine the relationship of base composition skews with replication orientation, gene function, codon usage biases and phylogenetic evolution, in this paper a program called DNAskew was developed for the statistical analysis of strand asymmetry and codon composition bias in the DNA sequence. In addition, the program can also be used to predict the replication boundaries of genome sequences. The method builds on the fact that there are compositional asymmetries between the leading and the lagging strand for replication. DNAskew was written in Perl script language and implemented on the LINUX operating system. It works quickly with annotated or unannotated sequences in GBFF (GenBank flatfile) or fasta format. The source code is freely available for academic use at http://www.epizooty.com/pub/stat/DNAskew.

Algorithms↗

Codon adaptation index as a measure of dominating codon bias.

UNLABELLED: We propose a simple algorithm to detect dominating synonymous codon usage bias in genomes. The algorithm is based on a precise mathematical formulation of the problem that lead us to use the Codon Adaptation Index (CAI) as a 'universal' measure of codon bias. This measure has been previously employed in the specific context of translational bias. With the set of coding sequences as a sole source of biological information, the algorithm provides a reference set of genes which is highly representative of the bias. This set can be used to compute the CAI of genes of prokaryotic and eukaryotic organisms, including those whose functional annotation is not yet available. An important application concerns the detection of a reference set characterizing translational bias which is known to correlate to expression levels; in this case, the algorithm becomes a key tool to predict gene expression levels, to guide regulatory circuit reconstruction, and to compare species. The algorithm detects also leading-lagging strands bias, GC-content bias, GC3 bias, and horizontal gene transfer. The approach is validated on 12 slow-growing and fast-growing bacteria, Saccharomyces cerevisiae, Caenorhabditis elegans and Drosophila melanogaster. AVAILABILITY: http://www.ihes.fr/~materials.

Adaptation, Physiological↗

Coding strategy differences between constant and variable segments of immunoglobulin genes.

Vertebrate immunoglobulin (Ig) mRNAs reveal intraspecies variation in codon usage distinct from that seen with yeast or bacterial genes. Comparison of all available Ig gene sequences shows that %(G + C) in codon position III is consistently lower in variable (V) segments than in constant (C) segments. I find an even lower %(G + C) in the hypervariable domains of V segments. This analysis suggests that base substitution in Ig genes correlates positively with local A + T content.

Alligators and Crocodiles↗

Hierarchy of sequence-dependent features associated with prokaryotic translation.

Protein expression in the cell is affected by various sequence-dependent features. Several such sequence-dependent features have been individually studied,yet they have not been compared quantitatively in terms of their relative influence on protein expression,and a hierarchy of these elements has not been determined. Here we present a quantitative analysis examining sequence-dependent features involved in prokaryotic translation,namely,the base-pairing potential between the mRNA Shine-Dalgarno sequence and the ribosomal RNA,codon bias,and the identity of the stop codon. We analyzed these features both at intra- and intergenomic levels using the Escherichia coli and Haemophilus influenzae genomes. Within each genome,we examined the relationship between each feature and protein expression levels determined by 2D-gel analyses. At the intergenomic level,comparative genomic principles were applied to study the relative preservation of the different sequence-dependent properties between orthologs. From these analyses,we determined that biased codon usage is the property that is most highly associated with protein expression and that is most conserved. The identity of the stop codon and the base-pairing potential of the mRNA Shine-Dalgarno sequence and the rRNA seem to have less of an effect on protein expression.

3' Untranslated Regions↗

The "universal" leucine codon CTG in the secreted aspartyl proteinase 1 (SAP1) gene of Candida albicans encodes a serine in vivo.

A number of Candida species possess a tRNA(Ser)-like species that recognizes CTG codons that normally specify leucine (Leu) in the universal code of codon usage. Mass spectrometry and Edman sequencing of peptides from the secreted aspartyl proteinase isoenzyme (Sap1) demonstrate that positions specified by the CTG codon contain a nonmodified serine (Ser) in Candida albicans.

Aspartic Acid Endopeptidases↗

Internal structure of the silk fibroin gene of Bombyx mori. I The fibroin gene consists of a homogeneous alternating array of repetitious crystalline and amorphous coding sequences.

The DNA sequence orgainzation of the protein encoding region of the gene for silk fibroin has been analyzed. The accompanying paper (Manningm R. F., and Gage, L. P. (1980) J. Biol. Chem. 255, 9451-9457) shows that the total length of the gene, and its protein, as well as the pattern of restriction sites in the gene is highly polymorphic among inbred stocks of Bombyx mori, In this paper, those features of fibroin gene structure which are invariant among these alleles are presented. Fibroin is composed primarily of relatively short "crystalline" and "amorphous" peptides of known sequence whose arrangement in the protein is unknown. Knowledge of the codons most commonly used in fibroin mRNA allowed utilization of particular restriction inzymes as a means for determing the nature and organization of crystalline and amorphous coding sequences in the fibroin gene. Three restriction endonucleases were identified that cleve sequences coding for amorphous region peptides. Their cleavage pattern revelaed that the repetitive coding sequence of the gene core (approximately 15 kilobases) is divided into at least 10 large crystalline coding domains interrupted by smaller amorphous coding domains. Many restriction endoncleases do not cleave the fibroin core at all, three of them with four gase recognition sequences. Specific deductions as to codon usage and repetitive sequence homogeneity in the gene follow from these results. One novel finding is the rigorous exclusion of the glycine codon GGA prior to serine codons even though this glycine codon is used frequently prior to alanine codons. The sequence homogeneity and the regularly alternating arrangement of crystalline and amorphous coding sequences of the gene are discussed in terms of the function of fibroin protein and the evolution of highly repetitive DNA.

Alleles↗

Doublet frequencies and codon weighting in the DNA of Escherichia coli and its phages.

A compilation of nucleic acid sequences from E. coli and its phages has been analysed for the frequency of occurrence of nearest neighbour base doublets and codons. Several statistically significant deviations from random are found in both doublet and codon frequencies. The deviations in E. coli also appear to occur in lambda and in the coat protein gene of MS2, whereas T4 and other parts of the MS2 genome show different sequence properties. These and other findings are discussed in relation to the hypothesis that rapidity of translation of mRNAs in the E. coli system is dependent on doublet frequency and codon usage patterns.

Base Sequence↗

Universal replication biases in bacteria.

Analysis of 15 complete bacterial chromosomes revealed important biases in gene organization. Strong compositional asymmetries between the genes lying on the leading versus lagging strands were observed at the level of nucleotides, codons and, surprisingly, amino acids. For some species, the bias is so high that the sole knowledge of a protein sequence allows one to predict with almost no errors whether the gene is transcribed from one strand or the other. Furthermore, we show that these biases are not species specific but appear to be universal. These findings may have important consequences in our understanding of fundamental biological processes in bacteria, such as replication fidelity, codon usage in genes and even amino acid usage in proteins.

Amino Acids↗

Codon-based mutagenesis using dimer-phosphoramidites.

A new approach for the synthesis of randomized DNA sequences containing the 20 codons corresponding to all natural amino acids is described. The strategy is based on the use of dinucleotide phosphoramidite building blocks within a resin-splitting procedure. Through this protocol, a minimal number of seven dimers is sufficient to encode all 20 natural amino acids. This synthesis procedure is extremely flexible and allows codon usage from different hosts to be accommodated.

Base Sequence↗