Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Design, total synthesis, and functional overexpression of the Candida rugosa lip1 gene coding for a major industrial lipase.

The dimorphic yeast Candida rugosa has an unusual codon usage that hampers the functional expression of genes derived from this yeast in a conventional heterologous host. Commercial samples of C. rugosa lipase (CRL) are widely used in industry, but contain several different isoforms encoded by the lip gene family, among which the isoform encoded by the gene lip1 is the most prominent. In a first laborious attempt, the lip1 gene was systematically modified by site-directed mutagenesis to gain functional expression in Saccharomyces cerevisiae. As alternative approach, the gene (1647 bp) was completely synthesized with an optimized nucleotide sequence in terms of heterologous expression in yeast and simplified genetic manipulation. The synthetic gene was functionally expressed in both hosts S. cerevisiae and Pichia pastoris, and the effect of heterologous leader sequences on expression and secretion was investigated. In particular, using P. pastoris cells, the synthetic gene was functionally overexpressed, allowing for the first time to produce recombinant Lipl of high purity at a level of 150 U/mL culture medium. The physicochemical and catalytic properties of the recombinant lipase were compared with those of a commercial, nonrecombinant C. rugosa lipase preparation containing lipase isoforms.

Amino Acid Sequence↗

Regeneration of sugarcane elite breeding lines and engineering of stem borer resistance.

Five elite sugarcane breeding lines were tested for efficiency in embryogenesis and plant regeneration. All of them produced regenerative embryogenic calli but with varied efficiencies. To engineer strongly insect-resistant sugarcanes, the GC content of a truncated cry1Ac gene, which encodes the active region of Cry1Ac insecticidal delta-endotoxin, was increased from the original 37.4 to 47.5% following the sugarcane codon usage pattern. The synthetic cry1Ac gene (s-cry1Ac) was placed under the control of maize ubiquitin promoter and introduced by microprojectile bombardment into the embryogenic calli of sugarcane lines YT79-177 and ROC16. Southern blotting analysis showed that multicopies of s-cry1Ac were integrated into the genomes of transgenic sugarcane lines. Immunoblotting analysis identified 18 transgenic lines expressing detectable levels of s-Cry1Ac, which were estimated in the range of 1.8-10.0 ng mg(-1) total soluble proteins. Four transgenic and two parental lines were assayed for sugarcane stem borer resistance in leaf tissue feeding trials and greenhouse plant assays. The results showed that, while the untransformed control lines were severely damaged in both leaves and stems, the transgenic sugarcane lines expressing high levels of s-Cry1Ac proteins were highly resistant to sugarcane stem borer attack, resulting in complete mortality of the inoculated insects within 1 week after inoculation.

Animals↗

Novel alpha-conotoxins identified by gene sequencing from cone snails native to Hainan, and their sequence diversity.

Conotoxins (CTX) from the venom of marine cone snails (genus Conus) represent large families of proteins, which show a similar precursor organization with surprisingly conserved signal sequence of the precursor peptides, but highly diverse pharmacological activities. By using the conserved sequences found within the genes that encode the alpha-conotoxin precursors, a technique based on RT-PCR was used to identify, respectively, two novel peptides (LiC22, LeD2) from the two worm-hunting Conus species Conus lividus, and Conus litteratus, and one novel peptide (TeA21) from the snail-hunting Conus species Conus textile, all native to Hainan in China. The three peptides share an alpha4/7 subfamily alpha-conotoxins common cysteine pattern (CCX(4)CX(7)C, two disulfide bonds), which are competitive antagonists of nicotinic acetylcholine receptor (nAChRs). The cDNA of LiC22N encodes a precursor of 40 residues, including a propeptide of 19 residues and a mature peptide of 21 residues. The cDNA of LeD2N encodes a precursor of 41 residues, including a propeptide of 21 residues and a mature peptide of 16 residues with three additional Gly residues. The cDNA of TeA21N encodes a precursor of 38 residues, including a propeptide of 20 residues and a mature peptide of 17 residues with an additional residue Gly. The additional residue Gly of LeD2N and TeA21N is a prerequisite for the amidation of the preceding C-terminal Cys. All three sequences are processed at the common signal site -X-Arg- immediately before the mature peptide sequences. The properties of the alpha4/7 conotoxins known so far were discussed in detail. Phylogenetic analysis of the new conotoxins in the present study and the published homologue of alpha4/7 conotoxins from the other Conus species were performed systematically. Patterns of sequence divergence for the three regions of signal, proregion, and mature peptides, both nucleotide acids and residue substitutions in DNA and peptide levels, as well as Cys codon usage were analyzed, which suggest how these separate branches originated. Percent identities of the DNA and amino acid sequences of the signal region exhibited high conservation, whereas the sequences of the mature peptides ranged from almost identical to highly divergent between inter- and intra-species. Notably, the diversity of the proregion was also high, with an intermediate percentage of divergence between that observed in the signal and in the toxin regions. The data presented are new and are of importance, and should attract the interest of researchers in this field. The elucidated cDNAs of these toxins will facilitate a better understanding of the relationship of their structure and function, as well as the process of their evolutionary relationships.

Amino Acid Sequence↗

Cloning and sequencing of the yeast Saccharomyces cerevisiae SEC1 gene localized on chromosome IV.

The SEC1 gene of yeast Saccharomyces cerevisiae was cloned by complementing the temperature-sensitive mutation of sec1-1 at 37 degrees C, and its nucleotide sequence was determined. SEC1 is a single copy gene and encodes a protein of 724 amino acids and 83,490 daltons with a predicted pI value of 6.11. Hydrophobicity plotting showed no clearly hydrophobic regions suggesting a soluble nature for the protein. Amino acid sequence comparisons revealed no obvious homologies with the proteins in the SWISSPROT databank. Two consensus sequence for the cdc2 encoded protein kinase recognition site were revealed within Sec1p. The codon usage suggests a low expression level for SEC1. The 5' non-translated region contains two TATA-like sequences at -52 and -215 nucleotides from the translation start site. Two potential regulatory sequences for DNA binding proteins were found in the non-coding 5' region: a HAP2/HAP3 consensus recognition sequence at nucleotide-154 and a BAF1 consensus recognition sequence at nucleotide-136. The SEC1 specific probe detected a 2400 nucleotides long transcript, which was in reasonable agreement with the 2172 nucleotides long open reading frame.

Amino Acid Sequence↗

Sequence of a 7.8 kb segment on the left arm of yeast chromosome XI reveals four open reading frames, including the CAP1 gene, an intron-containing gene and a gene encoding a homolog to the mammalian UOG-1 gene.

We report here the DNA sequence of a segment of chromosome XI of Saccharomyces cerevisiae extending over 7.8 kb. The segment contains four long open reading frames, YKL150, YKL153, YKL155 and YKL156, YKL155 corresponds to the CAP1 gene. YKL153 contains an intron and shows an extremely biased codon usage suggestive of a highly expressed protein. YKL156 is a homolog to UOG-1, an open reading frame associated with the cDNA clone of the mammalian growth/differentiation factor 1. YKL150 reveals common motifs to both the RNA polymerase II elongation factor of Drosophila melanogaster and to the yeast PPR2 gene product.

Actin Capping Proteins↗

Three yeast genes, PIR1, PIR2 and PIR3, containing internal tandem repeats, are related to each other, and PIR1 and PIR2 are required for tolerance to heat shock.

We isolated three highly homologous genes, PIR1, PIR2 and PIR3, collectively called the PIR genes. The remarkable feature of their putative amino acid sequence is that they contain a sequence consisting of 18-19 amino acid residues repeated tandemly seven to ten times. Genes homologous to PIR were found in Kluyveromyces lactis and Zygosaccharomyces rouxii but not in Schizosaccharomyces pombe, suggesting that a set of PIR genes plays some role in budding yeast. Bias of codon usage seen in each of the PIR translation products suggests that they are expressed abundantly. The fact that disruption of each gene is viable indicates that none of them is essential. The double disruptants, pir1 pir2, were viable under various conditions, such as higher temperature (37 degrees C) or high salt concentration, but showed a slow-growing phenotype on an agar slab. Furthermore, they were sensitive to heat shock. Addition of a pir3 disruption to the pir1 pir2 double disruptant brought about no phenotypic difference from the original double mutant. PIR1 and PIR3 are closely linked to each other and are on chromosome XI.

Amino Acid Sequence↗

Two distinct yeast proteins are related to the mammalian ribosomal polypeptide L7.

The RLP7 gene of Saccharomyces cerevisiae was cloned, sequenced and localized to the right arm of chromosome XIV, close to the centromere. It encodes a predicted polypeptide (RLP7p) of 322 amino acids, with a calculated molecular mass of 36 kDa and an isoelectric point of 9.6. Putative open reading frames very similar to RLP7 are present in two other yeasts, Kluyveromyces lactis and Candida utilis. The RLP7p gene product has significant sequence similarity to the S. cerevisiae YL8 polypeptide of the large ribosomal subunit (Mizuta et al., 1992), itself homologous to the L7 subunit of mammalian ribosomes. However, RLP7p and YL8 do not functionally replace each other, since an rlp7-delta::HIS3 strain is completely inviable. Judging from its predicted mass, isoelectric point and amino acid sequence, RLP7p does not correspond to any ribosomal component biochemically identified so far in S. cerevisiae, and also differs from all known ribosomal proteins by the low codon usage bias of its gene.

Amino Acid Sequence↗

Nucleotide sequence analysis of an 8887 bp region of the left arm of yeast chromosome XIV, encompassing the centromere sequence.

The nucleotide sequencing of 8887 bp of the left arm of chromosome XIV is described. The sequence includes the centromeric region. Both strands were sequenced with an average redundancy of 5.09 per base pair. The overall G+C content is 37.3% (39.2% for putative coding regions versus 32.5% for non-coding regions). Six open reading frames (ORFs) greater than 100 amino acids were detected, all of which are completely confined to the 8.9 kbp region. Codon frequencies of the six ORFs agree with codon usage in Saccharomyces cerevisiae and all show the characteristics of low-level expressed genes. Comparison of the translated sequences with protein sequences in data bases suggests the presence of two ORFs (N2014 and N2007) encoding ribosomal proteins, the latter of which is the previously sequenced MRP7 gene. Another ORF (N2012) could encode a membrane-associated protein since it contains secretory signal sequence and two presumed transmembrane helices. This protein might be involved in mitochondrial energy transfer. ORF N2016 is immediately adjacent to the centromere, suggesting that it corresponds to the SPO1 gene, which is very tightly linked to the centromere at the left arm side of chromosome XIV (Mortimer et al., 1989).

Base Composition↗

Twelve open reading frames revealed in the 23.6 kb segment flanking the centromere on the Saccharomyces cerevisiae chromosome XIV right arm.

The nucleotide sequence of 23.6 kb of the right arm of chromosome XIV is described, starting from the centromeric region. Both strands were sequenced with an average redundancy of 4.87 per base pair. The overall G+C content is 38.8% (42.5% for putative coding regions versus 29.4% for non-coding regions). Twelve open reading frames (ORFs) greater than 100 amino acids were detected. Codon frequencies of the twelve ORFs agree with codon usage in Saccharomyces cerevisiae and all show the characteristics of low level expressed genes. Five ORFs (N2019, N2029, N2031, N2048 and N2050) are encoded by previously sequenced genes (the mitochondrial citrate synthase gene, FUN34, RPC34, PRP2 and URK1, respectively). ORF N2052 shows the characteristics of a transmembrane protein. Other elements in this region are a tRNA(Pro) gene, a tRNA(Asn) gene, a tau 34 and a truncated delta 34 element. Nucleotide sequence comparison results in relocation of the SIS1 gene to the left arm of the chromosome as confirmed by colinearity analysis.

Base Sequence↗

Two-dimensional protein map of Saccharomyces cerevisiae: construction of a gene-protein index.

This publication marks the beginning of the construction of a gene-protein index that relates proteins which are resolved on the two-dimensional protein map of Saccharomyces cerevisiae with their corresponding genes. We report the identification of 36 novel polypeptide spots on the yeast protein map. They correspond to the products of 26 genes. Together with the polypeptide spots previously identified, this raises to 41 the number of genes whose products have been identified on the protein map. The proteins identified here are concerned with four major areas of yeast cellular physiology: carbon metabolism, heat shock, amino acid biosynthesis and purine biosynthesis. Given the molecular weight and isoelectric point of the identified proteins, and the codon-usage bias of the corresponding genes, it can be estimated that 25 to 35% of all the soluble yeast proteins are detectable under the labelling and running gel conditions used in this study.

Amino Acid Sequence↗

Cassettes for PCR-mediated construction of green, yellow, and cyan fluorescent protein fusions in Candida albicans.

We have developed a set of plasmids containing fluorescent protein cassettes for use in PCR-mediated gene tagging in Candida albicans. We engineered YFP and CFP variants of the GFP sequence optimized for C. albicans codon usage. The fluorescent protein sequences, linked to C. albicans auxotrophic marker sequences, were amplified by PCR and transformed directly into yeast. Gene-specific sequence was incorporated into the PCR primers, such that the tag-cassette integrates by homologous recombination at the 3'-end of the gene of interest. This technique was used to tag Cdc3 and Tub1 with GFP, YFP and CFP, which were readily visualized by fluorescence microscopy and localized as expected. In addition, Tub1-YFP and Cdc3-CFP were visualized in the same cells. Thus, this technique directs one-step construction of multiple fluorescent protein fusions, facilitating the study of protein co-expression and co-localization in C. albicans cells in vivo.

Bacterial Proteins↗

Heterologous expression and characterization of a "Pseudomature" form of taxadiene synthase involved in paclitaxel (Taxol) biosynthesis and evaluation of a potential intermediate and inhibitors of the multistep diterpene cyclization reaction.

The diterpene cyclase taxadiene synthase from yew (Taxus) species transforms geranylgeranyl diphosphate to taxa-4(5),11(12)-diene as the first committed step in the biosynthesis of the anti-cancer drug Taxol. Taxadiene synthase is translated as a preprotein bearing an N-terminal targeting sequence for localization to and processing in the plastids. Overexpression of the full-length preprotein in Escherichia coli and purification are compromised by host codon usage, inclusion body formation, and association with host chaperones, and the preprotein is catalytically impaired. Since the transit peptide-mature enzyme cleavage site could not be determined directly, a series of N-terminally truncated enzymes was created by expression of the corresponding cDNAs from a suitable vector, and each was purified and kinetically evaluated. Deletion of up to 79 residues yielded functional protein; however, deletion of 93 or more amino acids resulted in complete elimination of activity, implying a structural or catalytic role for the amino terminus. The pseudomature form of taxadiene synthase having 60 amino acids deleted from the preprotein was found to be superior with respect to level of expression, ease of purification, solubility, stability, and catalytic activity with kinetics comparable to the native enzyme. In addition to the major product, taxa-4(5),11(12)-diene (94%), this enzyme produces a small amount of the isomeric taxa-4(20), 11(12)-diene ( approximately 5%), and a product tentatively identified as verticillene ( approximately 1%). Isotopically sensitive branching experiments utilizing (4R)-[4-(2)H(1)]geranylgeranyl diphosphate confirmed that the two taxadiene isomers, and a third (taxa-3(4),11(12)-diene), are derived from the same intermediate taxenyl C4-carbocation. These results, along with the failure of the enzyme to utilize 2, 7-cyclogeranylgeranyl diphosphate as an alternate substrate, indicate that the reaction proceeds by initial ionization of the diphosphate ester and macrocyclization to the verticillyl intermediate, followed by a secondary cyclization to the taxenyl cation and deprotonation (i.e., formation of the A-ring prior to B/C-ring closure). Two potential mechanism-based inhibitors were tested with recombinant taxadiene synthase but neither provided time-dependent inactivation nor afforded more than modest competitive inhibition.

Antineoplastic Agents↗

Renaturation of recombinant human pro-urokinase expressed in Escherichia coli.

A synthetic gene encoding human pro-urokinase (pro-UK) with E. coli-favored codon usage was cloned into plasmid pET-3d and expressed in E. coli BL21(DE3) LysS strain. The expressed products, which accumulated as inactive inclusion bodies, were denatured and renatured in vitro. A broad range of parameters such as pH, protein concentration, denaturant concentration, the use of cosolvent polyethylene glycol and presence of basic or acidic amino acid was examined. At optimal renaturation condition, pro-UK activity of more than 1000I.U was obtained from 1 milliliter cell culture.

Cloning, Molecular↗

Trichomonas vaginalis: characterization, expression, and phylogenetic analysis of a carbamate kinase gene sequence.

The gene encoding carbamate kinase (CBK, ATP:carbamate phosphotransferase, EC 2.7.2.2) from Trichomonas vaginalis has been sequenced and its expression in this protozoon has been verified using reverse-transcription polymerase chain reaction. The codon usage and percentage nucleotide composition in the coding and noncoding regions are consistent with other genes isolated from this parasite. Phylogenetic analysis of this gene has suggested possible speciation events that are congruent with other studies, with suggestions of lateral gene transfer. The gene was expressed in Escherichia coli using the pQE-30 expression system, and the recombinant protein was purified using affinity chromatography. The expression of the recombinant protein was identified via Western blotting and matrix-assisted laser desorption ionization mass spectrometry of tryptic peptides. Preliminary kinetic assays have revealed that the recombinant enzyme has a K(m) similar to that of the native enzyme and size-exclusion chromatography has shown that the enzyme is active as the homodimer.

Amino Acid Sequence↗

Structure and organization of the human neuronatin gene.

Neuronatin is a brain-specific human gene that we recently isolated and observed to be selectively expressed during brain development. In this report, the genomic structure and organization of human neuronatin is described. The human gene spans 3973 bases and contains three exons and two introns. Based on primer extension analysis, a single cap site is located 124 bases upstream from the methionine (ATG) initiation codon, in good context, GAACCATGG. The promoter contains a modified TATA box, CATAAA (-27), and a modified CAAT box, GGCGAAT (-59). The 5'-flanking region contains putative transcription factor binding sites for SP-1, AP-2 (two sites), delta-subunit, SRE-2, NF-A1, and ETS. In addition, a 21-base sequence highly homologous to the neural restrictive silence element that governs neuron-specific gene expression is observed at -421. Furthermore, SP-1 and AP-3 binding sites are present in intron 1. All splice donor and acceptor sites conformed to the GT/AG rule. Exon 1 encodes 24 amino acids, exon 2 encodes 27 amino acids, and exon 3 encodes 30 amino acids. At the 3'-end of the gene, the poly(A) signal, AATAAA, poly(A) site, and GT cluster are observed. The neuronatin gene is expressed as two mRNA species, alpha and beta, generated by alternative splicing. The alpha-form contains all three exons, whereas in the beta-form, the middle exon has been spliced out. The third nucleotide of all frequently used codons, except threonine, of neuronatin is either G or C, consistent with codon usage expected for Homo sapiens. This information about the structure of the human neuronatin gene will help in understanding the significance of this gene in brain development and human disease.

Amino Acid Sequence↗

The organization of intercistronic regions of the aerobactin operon of pColV-K30 may account for the differential expression of the iucABCD iutA genes.

The complete nucleotide sequence of the 8.3 kilobase operon of the enterobacterial virulence plasmid pColV-K30, which encodes a high-affinity iron transport system mediated by the hydroxamate siderophore aerobactin, has been determined. The region includes five open reading frames which correspond to the genes iucA, iucB, iucC and iucD, encoding the enzymes of the biosynthetic pathway for aerobactin, and iutA for the outer membrane receptor of ferri-aerobactin complexes. The sequences of the iucABCD genes are tightly coupled without any intervening non-coding sequence. The predicted secondary mRNA structures at the gene junctions within the iucABCD cluster, along with their codon usage, may account for the differential expression of each of the protein products, as observed in vivo with minicells. The genes iucA and iucC, which determine the two subunits of the aerobactin synthetase complex, showed a considerable homology within three stretches of their amino acid sequence. A potential operator sequence (iron box) for the binding of the iron(II)-responsive Fur repressor protein was found within the iucA coding region, suggesting that the operon is subjected to an additional level of transcriptional repression by iron (II).

Amino Acid Sequence↗

Uneven distribution of GATC motifs in the Escherichia coli chromosome, its plasmids and its phages.

This work reconsiders the GATC motif distribution in a 1.6 Mb segment of the Escherichia coli genome, compared to its distribution in phages and plasmids. At first sight the distribution of GATC words looks random. But when a realistic model of the chromosome (made of average genes having the same codon usage as in the real chromasome), is used as a theoretical reference, strong biasesare observed. GATC pairs such as GATCNNGATC are under-represented while there is a strong positive selection for motifs separated by 10, 19, 70 and 1100 bp. The last class is the only one present in E. coli parasites. It can be ascribed to the triggering sequences of the long-patch mismatch repair system. The 6 bp class overlaps with the consensus of CAP (catabolite activator protein) and FNR (fumarate/nitrate regulator) binding sites, thus accounting for counter-selection. The other classes, which could be targets for a nucleic acid-binding protein, are almost always present inside protein coding sequences, and are members of clusters of GATC motifs. Analysis of the genes containing these motifs suggests that they correspond to a regulatory process monitoring the shift from anaerobic to aerobic growth conditions. In particular this regulation, closing down transcription of a large number of genes involved in intermediary metabolism would be well suited for the cold and oxygen shift from the mammal's gut to the standard environmental conditions. In this process the methylation status of GATC clusters would be very important for tuning transcription, and a DNA binding protein, probably a member of the cold-shock proteins family would be needed for alleviating the effects mediated by slackening of the pace of methylation during the shift.

Bacteriophages↗

Molecular evolution of the histone 3 multigene family in the Drosophila melanogaster species subgroup.

Molecular evolution of the histone multigene family was studied by cloning and sequencing regions of the histone 3 gene in the Drosophila melanogaster species subgroup. Analysis of the nucleotide substitution pattern showed that in the coding region synonymous changes occurred more frequently to A or T in contrast to the GC-rich base composition, while in the 3' region the nucleotide substitutions were most likely in equilibrium. These results suggested that the base composition at the third codon position of the H3 gene, i.e., codon usage, has been changing to A or T in the Drosophila melanogaster species subgroup.

Animals↗