Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Structure of the gene encoding the exoglucanase of Cellulomonas fimi.

In Cellulomonas fimi the cex gene encodes an exoglucanase (Exg) involved in the degradation of cellulose. The gene now has been sequenced as part of a 2.58-kb fragment of C. fimi DNA. The cex coding region of 1452 bp (484 codons) was identified by comparison of the DNA sequence to the N-terminal amino acid (aa) sequence of the Exg purified from C. fimi. The Exg sequence is preceded by a putative signal peptide of 41 aa, a translational initiation codon, and a sequence resembling a ribosome-binding site five nucleotides (nt) before the initiation codon. The nt sequence immediately following the translational stop codon contains four inverted repeats, two of which overlap, and which can be arranged in stable secondary structures. The codon usage in C. fimi appears to be quite different from that of Escherichia coli. A dramatic (98.5%) bias occurs for G or C in the third position for the 35 codons utilized in the cex gene.

Amino Acid Sequence↗

Malaria parasites contain two identical copies of an elongation factor 1 alpha gene.

Elongation factor 1alpha (EF-1alpha) is an abundant protein in eukaryotic cells, involved chiefly in translation of mRNA on the ribosomes, and is frequently encoded by more than one gene. Here we show the presence of two identical copies of the EF-1alpha gene in the genome of three malaria parasites, Plasmodium knowlesi, P. berghei and P. falciparum. They are organized in a head-to-head orientation and both genes are expressed in a stage specific manner at a high level, indicating that the small intergenic region contains either two strong promoters or a single bidirectional one. Both genes are expressed at the same time during erythrocytic development of the parasite. This expression pattern and the 100% similarity of the two genes excludes the possibility that the duplicated genes developed in accordance to the different types of ribosomes in Plasmodium. It is more likely that the duplication reflects a gene dosage effect. Comparison of codon usage in the Cdc2-related kinase genes (CRK2) of Plasmodium, which are expressed at a very low level, with the EF-1alpha genes indicates the existence of a codon bias for highly expressed genes, as has been shown in other organisms.

Amino Acid Sequence↗

Phylogenetic affinities of Diplonema within the Euglenozoa as inferred from the SSU rRNA gene and partial COI protein sequences.

In order to shed light on the phylogenetic position of diplonemids within the phylum Euglenozoa, we have sequenced small subunit rRNA (SSU rRNA) genes from Diplonema (syn. Isonema) papillatum and Diplonema sp. We have also analyzed a partial sequence of the mitochondrial gene for cytochrome c oxidase subunit I from D. papillatum. With both markers, the maximum likelihood method favored a closer grouping of diplonemids with kinetoplastids, while the parsimony and distance suggested a closer relationship of diplonemids with euglenoids. In each case, the differences between the best tree and the alternative trees were small. The frequency of codon usage in the partial D. papillatum COI was different from both related groups; however, as is the case in kinetoplastids but not in Euglena, both the non-canonical UGA codon and the canonical UGG codon were used to encode tryptophan in Diplonema.

Amino Acid Sequence↗

The guanine and cytosine content of genomic DNA and bacterial evolution.

The genomic guanine and cytosine (G + C) content of eubacteria is related to their phylogeny. The G + C content of various parts of the genome (protein genes, stable RNA genes, and spacers) reveals a positive linear correlation with the G + C content of their genomic DNA. However, the plotted correlation slopes differ among various parts of the genome or among the first, second, and third positions of the codons depending on their functional importance. Facts suggest that biased mutation pressure, called A X T/G X C pressure, has affected whole DNA during evolution so as to determine the genomic G + C content in a given bacterium. The role of A X T/G X C pressure in diversification of bacterial DNA sequences and codon usage patterns is discussed in the perspective of the neutral theory of molecular evolution.

Biological Evolution↗

Mitochondrial genomes of Clymenella torquata (Maldanidae) and Riftia pachyptila (Siboglinidae): evidence for conserved gene order in annelida.

Mitochondrial genomes are useful tools for inferring evolutionary history. However, many taxa are poorly represented by available data. Thus, to further understand the phylogenetic potential of complete mitochondrial genome sequence data in Annelida (segmented worms), we examined the complete mitochondrial sequence for Clymenella torquata (Maldanidae) and an estimated 80% of the sequence of Riftia pachyptila (Siboglinidae). These genomes have remarkably similar gene orders to previously published annelid genomes, suggesting that gene order is conserved across annelids. This result is interesting, given the high variation seen in the closely related Mollusca and Brachiopoda. Phylogenetic analyses of DNA sequence, amino acid sequence, and gene order all support the recent hypothesis that Sipuncula and Annelida are closely related. Our findings suggest that gene order data is of limited utility in annelids but that sequence data holds promise. Additionally, these genomes show AT bias (approximately 66%) and codon usage biases but have a typical gene complement for bilaterian mitochondrial genomes.

Animals↗

Cloning, sequencing, and expression of the P-protein gene (pheA) of Pseudomonas stutzeri in Escherichia coli: implications for evolutionary relationships in phenylalanine biosynthesis.

The pheA gene encoding the bifunctional P-protein (chorismate mutase:prephenate dehydratase) was cloned from Pseudomonas stutzeri and sequenced. This is the first gene of phenylalanine biosynthesis to be cloned and sequenced from Pseudomonas. The pheA gene was expressed in Escherichia coli, allowing complementation of an E. coli pheA auxotroph. The enzymic and physical properties of the P-protein from a recombinant E. coli auxotroph expressing the pheA gene were identical to those of the native enzyme from P. stutzeri. The nucleotide sequence of the P. stutzeri pheA gene was 1095 base pairs in length, predicting a 365-residue protein product with an Mr of 40,844. Codon usage in the P. stutzeri pheA gene was similar to that of Pseudomonas aeruginosa but unusual in that cytosine and guanine were used at nearly equal frequencies in the third codon position. The deduced P-protein product showed sequence homology with peptide sequences of the E. coli P-protein, the N-terminal portion of the E. coli T-protein (chorismate mutase:prephenate dehydrogenase), and the monofunctional prephenate dehydratases of Bacillus subtilis and Corynebacterium glutamicum. A narrow range of values (26-35%) for amino acid matches revealed by pairwise alignments of monofunctional and bifunctional proteins possessing activity for prephenate dehydratase suggests that extensive divergence has occurred between even the nearest phylogenetic lineages.

Amino Acid Sequence↗

Sequence and analysis of the DNA encoding protective antigen of Bacillus anthracis.

The nucleotide sequence of the protective antigen (PA) gene from Bacillus anthracis and the 5' and 3' flanking sequences were determined. PA is one of three proteins comprising anthrax toxin; and its nucleotide sequence is the first to be reported from B. anthracis. The open reading frame (ORF) is 2319 bp long, of which 2205 bp encode the 735 amino acids of the secreted protein. This region is preceded by 29 codons, which appear to encode a signal peptide having characteristics in common with those of other secreted proteins. A consensus TATAAT sequence was located at the putative -10 promoter site. A Shine-Dalgarno site similar to that found in genes of other Bacillus sp. was located 7 bp upstream from the ATG start codon. The codon usage for the PA gene reflected its high A + T (69%) base composition and differed from those of genes for bacterial proteins from most other sequences examined. The TAA translation stop codon was followed by an inverted repeat forming a potential termination signal. In addition, a 192-codon ORF of unknown significance, theoretically encoding a 21.6-kDa protein, preceded the 5' end of the PA gene.

Amino Acid Sequence↗

Codon optimization, expression, and characterization of an internalizing anti-ErbB2 single-chain antibody in Pichia pastoris.

Anti-ErbB2 antibodies are used as convenient tools in exploration of ErbB2 functional mechanisms and in treatment of ErbB2-overexpressing tumors. When we employed the yeast Pichia pastoris to express an anti-ErbB2 single-chain antibody (scFv) derived from the tumor-inhibitory monoclonal antibody A21, the yield did not exceed 1-2 mg/L in shake flask cultures. As we considered that the poor codon usage bias may be one limiting factor leading to the inefficient translation and scFv production, we designed and synthesized the full-length scFv gene by choosing the P. pastoris preferred codons while keeping the G+C content at relatively low level. Codon optimization increased the scFv expression level 3- to 5-fold and up to 6-10 mg/L. Northern blotting further confirmed that the increase of scFv expression was mainly due to the enhancement of translation efficiency. Investigation of culture conditions revealed that the maximal cell growth and scFv expression were achieved at pH 6.5-7.0 with 2% casamino acids after 72 h methanol induction. Secreted scFv was easily purified (>95% homogeneous product) from culture supernatants in one step by using Ni2+ chelating affinity chromatography. The yield was approximately 10-15 mg/L. Functional studies showed that the A21 scFv could be internalized with high efficiency after binding to the ErbB2-overexpressing cells, suggesting this regent may prove especially useful for ErbB2-targeted immunotherapy.

Amino Acid Sequence↗

A test of translational selection at 'silent' sites in the human genome: base composition comparisons in alternatively spliced genes.

Natural selection appears to discriminate among synonymous codons to enhance translational efficiency in a wide range of prokaryotes and eukaryotes. Codon bias is strongly related to gene expression levels in these species. In addition, between-gene variation in silent DNA divergence is inversely correlated with codon bias. However, in mammals, between-gene comparisons are complicated by distinctive nucleotide-content bias (isochores) throughout the genome. In this study, we attempted to identify translational selection by analyzing the DNA sequences of alternatively spliced genes in humans and in Drosophila melanogaster. Among codons in an alternatively spliced gene, those in constitutively expressed exons are translated more often than those in alternatively spliced exons. Thus, translational selection should act more strongly to bias codon usage and reduce silent divergence in constitutive than in alternative exons. By controlling for regional forces affecting base-composition evolution, this within-gene comparison makes it possible to detect codon selection at synonymous sites in mammals. We found that GC-ending codons are more abundant in constitutive than alternatively spliced exons in both Drosophila and humans. Contrary to our expectation, however, silent DNA divergence between mammalian species is higher in constitutive than in alternative exons.

Alternative Splicing↗

The complete sequence of the zebrafish (Danio rerio) mitochondrial genome and evolutionary patterns in vertebrate mitochondrial DNA.

We describe the complete sequence of the 16,596-nucleotide mitochondrial genome of the zebrafish (Danio rerio); contained are 13 protein genes, 22 tRNAs, 2 rRNAs, and a noncoding control region. Codon usage in protein genes is generally biased toward the available tRNA species but also reflects strand-specific nucleotide frequencies. For 19 of the 20 amino acids, the most frequently used codon ends in either A or C, with A preferred over C for fourfold degenerate codons (the lone exception was AUG: methionine). We show that rates of sequence evolution vary nearly as much within vertebrate classes as between them, yet nucleotide and amino acid composition show directional evolutionary trends, including marked differences between mammals and all other taxa. Birds showed similar compositional characteristics to the other nonmammalian taxa, indicating that the evolutionary trend in mammals is not solely due to metabolic rate and thermoregulatory factors. Complete mitochondrial genomes provide a large character base for phylogenetic analysis and may provide for robust estimates of phylogeny. Phylogenetic analysis of zebrafish and 35 other taxa based on all protein-coding genes produced trees largely, but not completely, consistent with conventional views of vertebrate evolution. It appears that even with such a large number of nucleotide characters (11,592), limited taxon sampling can lead to problems associated with extensive evolution on long phyletic branches.

Animals↗

Massive overproduction of dihydrofolate reductase in bacteria as a response to the use of trimethoprim.

Among several observations of greatly increased levels of chromosomal dihydrofolate reductase as a cause of resistance to high concentrations of the antifolate drug trimethoprim, in clinically isolated bacteria, one is described here of a strain of Escherichia coli overproducing dihydrofolate reductase several hundredfold. The chromosomally located resistance gene of this strain was isolated, inserted into a plasmid vector, and analyzed for its nucleotide sequence. The structural gene for the overproduced dihydrofolate reductase was found to be identical to that of E. coli K12, with nine exceptions, of which seven resulted in synonymous codon usage. Two transversions resulted in a substitution of Gly or Trp at amino acid position 30, and of Gln for Glu at position 154. Six of the nine base changes resulted in codons more frequently used. The Gly substitution which leads to a less commonly used codon, was thought to relate to the observed threefold increase in Ki for trimethoprim. Furthermore, a C----T transition was found in the -35 region of the promoter, increasing its homology with the E. coli consensus promoter sequence. In the ribosome-binding area of the resistant strain, finally, seven base changes were observed, two of which resulted in a five-base sequence of complementarity with the 3'-end of ribosomal 16S RNA. The distance between the -10 site of the promoter and the start codon for translation was finally increased one base pair by the insertion of an A at position +9 in the resistant strain. These genetic changes towards more efficient transcriptional and translational start sequences and towards increased mRNA expressivity are interpreted to reflect an evolutionary adaptation to the presence of antifolates.

Base Sequence↗

Recombination and selection in the evolution of picornaviruses and other Mammalian positive-stranded RNA viruses.

Picornaviridae are a large virus family causing widespread, often pathogenic infections in humans and other mammals. Picornaviruses are genetically and antigenically highly diverse, with evidence for complex evolutionary histories in which recombination plays a major part. To investigate the nature of recombination and selection processes underlying the evolution of serotypes within different picornavirus genera, large-scale analysis of recombination frequencies and sites, segregation by serotype within each genus, and sequence selection and composition was performed, and results were compared with those for other nonenveloped positive-stranded viruses (astroviruses and human noroviruses) and with flavivirus and alphavirus control groups. Enteroviruses, aphthoviruses, and teschoviruses showed phylogenetic segregation by serotype only in the structural region; lack of segregation elsewhere was attributable to extensive interserotype recombination. Nonsegregating viruses also showed several characteristic sequence divergence and composition differences between genome regions that were absent from segregating virus control groups, such as much greater amino acid sequence divergence in the structural region, markedly elevated ratios of nonsynonymous-to-synonymous substitutions, and differences in codon usage. These properties were shared with other picornavirus genera, such as the parechoviruses and erboviruses. The nonenveloped astroviruses and noroviruses similarly showed high frequencies of recombination, evidence for positive selection, and differential codon use in the capsid region, implying similar underlying evolutionary mechanisms and pressures driving serotype differentiation. This process was distinct from more-recent sequence evolution generating diversity within picornavirus serotypes, in which neutral or purifying selection was prominent. Overall, this study identifies common themes in the diversification process generating picornavirus serotypes that contribute to understanding of their evolution and pathogenicity.

Evolution, Molecular↗

Biosynthetic thiolase from Zoogloea ramigera. III. Isolation and characterization of the structural gene.

The gene coding for the biosynthetic thiolase from Zoogloea ramigera has been isolated by using antibody screening methods to detect its expression in Escherichia coli under the transcriptional control of the lac promoter. We have located and determined the nucleotide sequence of the gene. The structural gene is 1173 nucleotides long and codes for a polypeptide of 391 amino acids; 282 nucleotides 5' and 58 nucleotides 3' to the coding sequence are also reported. By comparing the amino acid sequence data predicted from the gene with data determined experimentally, we have derived the complete primary structure of thiolase. A catalytically essential cysteine is located at residue 89. The DNA sequence presented has a very high G/C content, 66.2%, typical of the Z. ramigera genome. In the coding region, this increases to 68.2% and is strongly reflected in the codon usage which demonstrates a strong preference for G or C in the third position. Examination of the 5'-flanking sequence establishes that the NH2-terminal methionine is specified by an ATG codon, 7 nucleotides downstream from a Shine-Dalgarno sequence.

Acetyl-CoA C-Acetyltransferase↗

The translational termination signal database.

The Translational Termination Database (TransTerm) consists of the immediate context sequences around the natural termination codons from 45 organisms, and summary tables. The influence of termination codon context on their effectivness as stop signals has been widely documented. The SPECIES--TRI.DAT table shows trinucleotide stop codon usage in each organism and for comparison the occurrence of these sequences in the noncoding region. The SPECIES--TETRA.DAT table contains is a similar table of tetranucleotide stop signal usage. The database is available from EMBL.

Animals↗

Molecular cloning, enhancement of expression efficiency and site-directed mutagenesis of rat epidermal cystatin A.

A rat cystatin A cDNA clone was isolated from a lambda ZAP library representing newborn rat skin mRNA by screening with a synthetic oligonucleotide designed from amino acid sequence 15-23 of the cysteine proteinase inhibitor. The obtained clone contained a partial coding region of the inhibitor, lacking the 5'-untranslated region and coding sequence for the NH(2)-terminal 13 residues. The amino acid sequence deduced from the base sequence, Glu14-Phe103, coincided with that determined at the amino acid level. To obtain the recombinant cystatin A protein, the DNA was fused with a synthetic linker encoding its missing N-terminal 17 residues and introduced into an expression vector, pMK2. In Escherichia coli, however, the expression level of the semi-synthetic gene was low, 0. 5 mg of the purified recombinant protein per 1 liter culture being produced. Changing of the codon usage of the N-terminal region in a pET-15b expression system led to an increase in the yield depending on the instability of the putative secondary structure around an initiation codon of the mRNA. The expressed cystatin A showed identical characteristics with the authentic form except for the absence of the N-terminal acetyl blocking group. Using the expression system, two kinds of point mutation, the conservative Val54 in the first loop QxVxG region being changed to Lys and Glu, were introduced, but there was almost no effect on the inhibitory activity toward papain. This suggests that the conserved Val in the reactive site is not restricted and that the hydrophobicity of the position is not essential for the activity of rat cystatin A.

Amino Acid Sequence↗

Sequence of the ebgA gene of Escherichia coli: comparison with the lacZ gene.

We have sequenced the ebgA (evolved beta-galactosidase) gene of Escherichia coli K12. The sequence shows 50% nucleotide identity with the E. coli lacZ gene, demonstrating that the two genes are related by descent from a common ancestral gene. Comparison of the two sequences suggests that the ebgA gene has recently been under selection. A significant excess of identical, rather than synonymous, codons used to encode identical amino acids at the same positions in the aligned sequences implies that some form of selection is operating directly at the DNA level. This selection is independent of, and in addition to, selection based on codon usage or on function of the gene products.

Amino Acid Sequence↗

DNA sequence and comparative analyses of the equine herpesvirus type 1 immediate early gene.

The immediate early (IE) proteins of herpesviruses are important regulatory factors which control the expression of genes at the transcriptional level. We report the DNA sequence of the immediate early gene of the alphaherpesvirus equine herpesvirus type 1 (EHV-1). This sequence is shown to be extremely rich in guanine and cytosine, resulting in a highly biased codon usage. The IE gene region possesses 38 open reading frames (ORFs) greater than 300 bp in length, 11 of which have coding regions of at least 100 amino acids (aa) following potential translation initiator codons. The largest ORF consists of 1487 codons (4461 bp) starting with the first ATG and would encode a protein of MW 155,000. TATA and CCAAT sequences as well as several potential cis-acting elements lie upstream to the major ORF. The deduced amino acid sequence for the 155,000 protein has a high degree of homology to the herpes simplex virus type 1 (HSV-1) ICP4 protein and its varicella-zoster virus (VZV) homolog. The regions of the EHV-1 IE protein that are homologous with these proteins correspond to the previously determined pattern of homology between the HSV and VZV IE polypeptides. However, there are are a number of differences within these broadly defined regions. It is therefore expected that this comparative study will facilitate the identification of functionally important residues within the amino acid sequence of IE proteins.

Amino Acid Sequence↗

Functional expression of the gene encoding cytidine triphosphate synthetase from Plasmodium falciparum which contains two novel sequences that are potential antimalarial targets.

CTP synthetase (E C 6.3.4.2 UTP: ammonia ligase (ADP-forming)) catalyses the formation of CTP from UTP and, in the human parasite Plasmodium falciparum, is the sole source of cytidine nucleotides. It is thus a potential chemotherapeutic target, especially as the gene sequence indicated that the encoded GAT-domain of the enzyme contains two extended peptide segments (42aa and 223aa as compared to the host enzyme). Here, we circumvent the codon usage problems associated with the high A/T content of the P. falciparum sequence, especially evident in sequences encoding the extra peptides, to successfully express active recombinant P. falciparum CTP synthetase using preferred E. coli codons. This partially synthetic gene produced recombinant enzyme, containing the additional segments, which was functionally assayed for activity in vitro. We also show the native enzyme contains the additional peptides using immunoblots with antibodies derived from the recombinant protein. Confocal microscopy, using antibodies to the recombinant protein, provided evidence that the enzyme is expressed in vivo. This establishes for the first time that P. falciparum contain active CTP synthetase and that this enzyme contains two novel insert sequences in the functional enzyme.

Animals↗