Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

A synthetic Rev-independent bovine immunodeficiency virus-based packaging construct.

Replication competent lentivirus (RCL) has been the major safety concern associated with applications of lentivirus-based gene transfer systems for human gene therapy. Minimization and elimination of overlaps between the packaging and the transfer vector constructs are expected to reduce the potential to generate RCL. We previously developed second- and third-generation bovine immunodeficiency virus (BIV)-based gene transfer systems. However, some sequence homologies between the vector and gag/pol packaging constructs remained. In order to minimize the sequence homologies, we recoded gag/pol with codon usage optimized for expression in human cells in this report. Expression of the recoded gag/pol was Rev/RRE independent. Thus, RRE was eliminated from the packaging construct, thereby removing a 312 bp block of homology. In addition, recoding gag/pol minimized overall homologies between the packaging and transfer vector constructs. Vectors generated by the recoded packaging construct with a four plasmid system had titers greater than 1 x 10(6) transducing units per milliliter, equivalent to those of the earlier generation systems. The vectors were functional in vitro and efficiently transduced rat pigment epithelial cells in vivo. Generation of the synthetic packaging construct provides further advances to the safety of lentiviral vectors for clinical applications.

Animals↗

A joint prediction of the folding types of 1490 human proteins from their genetic codons.

The codon usages for 1490 human proteins have been published by Wada et al. (1990). Based on these data, the frequencies of occurrence of 20 amino acids for each of the 1490 proteins have been calculated according to the genetic codes. Proteins are generally classified into five folding types, i.e. the alpha, beta, alpha + beta, alpha/beta and zeta (irregular) types. The folding type of a protein is correlated to its amino acid composition. By means of three methods established by different investigators, the folding type for each of the 1490 human proteins has been predicted. It has been demonstrated that the accuracy of prediction for the 1490 human proteins is at least 80% by examining the predicted results of some structurally known proteins with these methods. There are only six proteins for which there is uncertainty about their folding types as completely inconsistent results were obtained when predicted with the three different methods. For the remaining 1484 human proteins the numbers of alpha, beta, alpha + beta, alpha/beta, and zeta folding type proteins were found to be 128, 235, 169, 933 and 19, respectively, suggesting that the alpha/beta type proteins would predominate in this set of human proteins. The occurrence frequencies of bases in the first, second and third codon position for each folding type of protein have been calculated. It is shown that the folding type of a protein is strongly dependent on the ratio of frequency of base G in the first codon position with that in the second codon position. The biological implication of the results has been discussed.

Amino Acid Sequence↗

The complete maternal and paternal mitochondrial genomes of the Mediterranean mussel Mytilus galloprovincialis: implications for the doubly uniparental inheritance mode of mtDNA.

The maternal (F) and paternal (M) mitochondrial genomes of the mussel Mytilus galloprovincialis have diverged by about 20% in nucleotide sequence but retained identical gene content and gene arrangement and similar nucleotide composition and codon usage bias. Both lack the ATPase8 subunit gene, have two tRNAs for methionine and a longer open-reading frame for cox3 than seen in other mollusks. Between the F and M genomes, tRNAs are most conserved followed by rRNAs and protein-coding genes, even though the degree of divergence varies considerably among the latter. Divergence at nad3 is exceptionally low most likely because this gene includes the origin of transcription of the lagging strand (O(L)). Noncoding regions are the least conserved with the notable exception of the central domain of the main control region and a segment of another noncoding region immediately following nad3. The amino acid divergence (14%) of the two genomes is smaller than in two other pairs of conspecific genomes that are available in GenBank, that of the clam Venerupis philippinarum (34%) and of the fresh water mussel Inversidens japanensis (50%), suggesting that doubly uniparental inheritance of mtDNA emerged at different times in the three species or that there has been a relatively recent replacement of the male genome by the female in the Mytilus line. The latter hypothesis is supported from phylogenetic and population studies of Mytilidae. That the M genome contains a full complement of genes with no premature termination codons argues against it being a selfish element that rides with the sperm. It is shorter than the F by 118 bp, which apparently cannot account for the postulated replicative advantage of this genome over the F in male gonads. The high similarity of the two genomes explains why the F genome may assume the role of the M genome, but it does not exclude the possibility that for this to happen some M-specific sequences must be transferred on to the F genome by means of recombination. If such sequences exist they would most likely be located in noncoding regions.

Animals↗

Characterization of the Micromonospora rosaria pMR2 plasmid and development of a high G+C codon optimized integrase for site-specific integration.

pMR2, an 11.1 kb plasmid was isolated from Micromonospora rosaria SCC2095, NRRL3718, and its complete nucleotide sequence determined. Analysis revealed 13 ORFs including homologs of a KorSA regulatory protein and TraB plasmid transfer protein found on other actinomycete plasmids. pMR2 contains att/int functions consisting of an integrase, an excisionase, and a putative plasmid attachment site (attP). The integrase gene contained a high frequency of codons rarely used in high G+C actinomycete coding regions. The gene was codon optimized for actinomycete codon usage to create the synthetic gene int-OPT. pSPRX740, containing an rpsL promoter and the att/int-OPT region, was introduced into Micromonospora halophytica var. nigra ATCC33088. Analysis of DNA flanking the pSPRX740 integration site confirmed site-specific integration into a tRNA(Phe) gene in the M. halopytica var. nigra chromosome. The pMR2 attP element and chromosomal attachment (attB) site contain a 63 bp region of sequence identity overlapping the 3' end of the tRNA(Phe) gene. Plasmids comprising the site-specific att/int-OPT functions of pMR2 can be used to integrate genes into the chromosome of actinomycetes with an appropriate tRNA gene. The development of an integrative system for Micromonospora will expand our ability to study antibiotic biosynthesis in this important actinomycete genus.

Attachment Sites, Microbiological↗

Design of protective and therapeutic DNA vaccines for the treatment of allergic diseases.

The DNA vaccine revolution has opened a vast scope of novel approaches for protective and therapeutic treatments of type I allergy. This review gives an overview on the current status of allergy DNA vaccines and presents advances in the design of vaccine constructs. An immense number of concurring studies have proven the stimulation of Th1 cells and the induction of a balanced Th1/Th2 cytokine milieu as the fundamental mechanisms underlying the anti-allergic effects of DNA vaccines. Basic vaccine formulations thus can be optimized by improving the cellular immunogenicity via co-administration of cytokines, co-expression or co-application of immunostimulatory DNA sequences or adapting the codon usage. The latter is a frequent and major reason for impaired vaccine expression (e.g. translation of plant allergen genes in mammal cells). Because of unwanted side effects during conventional specific immunotherapy with allergen extracts, safety is increasingly demanded for both, protein and DNA vaccines for allergy treatment. We discuss the creation of hypoallergenic DNA vaccines based on deliberate allergen gene fragmentation, the use of mutations and the routine production of hypoallergenic DNA vaccines by forced ubiquitination. Furthermore, allergen-expressing DNA replicon vaccines are introduced, which enable a drastic reduction of the vaccine dose without loss of anti-allergic efficacy. Finally, the development of DNA multi vaccines and fusion vaccines for protective and therapeutic applications against certain groups of allergens is addressed.

Allergens↗

Unusual ciliate-specific codons in Tetrahymena mRNAs are translated correctly in a rabbit reticulocyte lysate supplemented with a subcellular fraction from Tetrahymena.

The codon usage of Tetrahymena thermophila and other ciliates deviates from the 'universal genetic code' in that UAA and probably UAG are not translational termination signals but code for glutamine. Therefore, translation in vitro of mRNA from Tetrahymena in a reticulocyte lysate is prematurely terminated if a UAA or UAG triplet is present in the reading frame of the mRNA. We show that the addition of a subcellular fraction from Tetrahymena thermophila enables a rabbit reticulocyte lysate to translate Tetrahymena mRNAs into full-sized proteins. The activity of the subcellular fraction is shown to depend on the combined function of a protein component(s) and a tRNA(s). The subcellular fraction is easily prepared and its usefulness for the identification of isolated mRNAs from Tetrahymena by their translation products in vitro is demonstrated.

Animals↗

Expression and characterization of soluble human erythropoietin receptor made in Streptomyces lividans 66.

A gene encoding the extracellular domain of the human erythropoietin receptor (EPO-R) was constructed using oligonucleotides, with a view to maintaining preferred codon usage for the Streptomycetes. The gene was subcloned into a multicopy Streptomyces-Escherichia coli shuttle vector, pCAN46 (derived from pIJ680), containing a strong constitutive promoter from the S. fradiae aph gene, a signal peptide coding region derived from the protease B gene of S. griseus, and a transcription terminator sequence also derived from the S. fradiae aph gene. Extracellular expression of authentic EPO-R by S. lividans was demonstrated using SDS-PAGE and Western blot analysis, followed by direct amino terminal sequencing of the purified product. Specific binding of S. lividans-expressed EPO-R to recombinant human glycosylated EPO was demonstrated using BIAcore (surface plasmon resonance) analysis and native gel shift assays.

Base Sequence↗

Evidence for horizontal gene transfer in Escherichia coli speciation.

After extracting more than 780 identified Escherichia coli genes from available data libraries, we investigated the codon usage of the corresponding coding sequences and extended the study of gene classes, thus obtained, to the nature and intensity of short nucleotide sequence selection, related to constraints operating at the nucleotide level. Using Factorial Correspondence Analysis we found that three classes ought to be included in order to match all data now available. The first two classes, as known, encompass genes expressed either continuously at a high level, or at a low level and/or rarely; the third class consists of genes corresponding to surface elements of the cell, genes coming from mobile elements as well as genes resulting in a high fidelity of DNA replication. This suggests that bacterial strains cultivated in the laboratory have been fixed by specific use of antimutator genes that are horizontally exchanged.

Amino Acids↗

Codon optimization, genetic insulation, and an rtTA reporter improve performance of the tetracycline switch.

The objective of this work was to further develop a tetracycline repressor (TetR) protein system that allows control of transgene expression. First, to circumvent the need for a binary approach, a single plasmid design was constructed and tested in tissue culture. To indirectly assay integrations that express the synthetic transcription factor (rtTA), a bicistronic gene was built which included an internal ribosome entry site (IRES) and a green fluorescent protein coding region (GFP) on the same expression cassette as the coding region of rtTA (pTetGREEN). This construct did not produce fluorescent colonies when stably integrated and provided minimal expression of GFP in the face of adequate expression of rtTA. The coding region for TetR was then altered by introducing 156 silent point mutations to simulate mammalian genes. Replacement of wild-type TetR gene (tetR) in pTetGREEN with 'mammalianized' tetR provided GFP expression. Adjustment of codon usage in the tetR region of rtTA nearly doubled the expression level of functional rtTA. To increase the number of rtTA expressing lines, the chicken egg-white lysozyme matrix attachment region (MAR) was introduced into the single plasmid design just upstream of the tetracycline operators (tetO). Inclusion of the MAR doubled the number of colonies that expressed rtTA (44% vs 88%). With the modifications described here, the number of lines that express rtTA and provide induction from a single plasmid design can be increased by the inclusion of a MAR and the level of rtTA expression can be further increased by adjusting the base composition of the TetR coding region. The MAR also insulates the inducible gene from the promoter driving rtTA.

Amino Acid Sequence↗

Molecular cloning, primary structure and disruption of the structural gene of aldolase from Saccharomyces cerevisiae.

A yeast cDNA genetic library in a bacteriophage expression vector was screened using an antiserum reacting with fructose 1,6-bisphosphate aldolase from Saccharomyces cerevisiae. Radio-labelled probes of selected immunopositive clones were used for screening of a yeast genomic library. From the genomic clones a yeast/Escherichia coli shuttle plasmid was constructed containing on a 1990-base-pair fragment the entire structural gene FBA1 coding for yeast aldolase. The primary structure of the FBA1 gene was determined. An open reading frame comprises 1077 base pairs coding for a protein of 359 amino acids with a predicted molecular mass of 39,608 Da. As observed for other strongly expressed yeast genes, codon usage is extremely biased. The 810 base pairs at the 5' end and the 90 base pairs at the 3' end of the coding region of the cloned FBA1 gene are sufficient for normal expression and show characteristic elements present in the noncoding sequences of other yeast genes. Aldolase is the major protein in yeast cells transformed with a high-copy-number plasmid containing the FBA1 gene. The aldolase gene was disrupted by insertion of the yeast URA3 gene into the coding region of one FBA1 allele in a homozygous diploid ura3 strain. The haploid offsprings with the defective aldolase allele fba1::URA3 lack aldolase enzymatic activity and fail to grow in media containing as a carbon source metabolites of only one side of the aldolase reaction.

Amino Acid Sequence↗

Cloning and DNA sequence of the Mycobacterium fortuitum var fortuitum plasmid pAL5000.

The complete nucleotide sequence of the Mycobacterium fortuitum var fortuitum plasmid pAL5000 has been determined. Computer analysis of this 4821-bp plasmid for protein coding regions, based on mycobacterial codon usage preferences, reveals the presence of two putative protein coding regions immediately downstream from typical mycobacterial promoter and ribosome binding sites. Both open reading frames, ORF1 and ORF2, produced proteins with the predicted respective sizes previously shown in minicell expression experiments [A. Labidi et al. (1985) FEMS Microbiol. Lett. 30, 221-225]. ORF1 encodes a putative 20-kDa basic protein with characteristics of a DNA binding protein involved in plasmid DNA replication. ORF2 encodes a 67-kDa protein with an amino-terminal sequence suggestive of a transported protein and a possible transmembrane anchor near its carboxyl-terminal. The current sequence and its analysis are more consistent with the minicell expression experiments than the previously published sequence of the pAL5000 plasmid [J. Rauzier et al. (1988) Gene 71, 315-321].

Amino Acid Sequence↗

The biosynthesis of the ubiquinol-cytochrome c reductase complex in yeast. DNA sequence analysis of the nuclear gene coding for the 14-kDa subunit.

The nuclear gene coding for the imported 14-kDa subunit of the ubiquinol-cytochrome c reductase of yeast mitochondria has been sequenced in an attempt to define regulatory and protein topogenic elements. The gene has a length of 381 base pairs and is potentially capable of encoding a polypeptide of 14561 Da. It is transcribed into a single low-abundance RNA of 680 nucleotides whose 5' and 3' termini map, respectively, 30-35 nucleotides upstream and 180-190 nucleotides downstream of the initiator and termination codons. Consistent with the estimated low level of the mRNA, codon usage in the gene is not strongly biased and other features, characteristic of highly expressed genes in yeast, are absent. The 14-kDa protein is predicted to be a predominantly hydrophilic protein, with only a single, short hydrophobic stretch located between positions 19-38. Comparison with other imported mitochondrial proteins so far sequenced has failed to reveal unifying features that might serve as targeting elements. Steady-state levels of the 14-kDa and 11-kDa subunits are reduced in mit- mutants which synthesize truncated forms of apocytochrome b and in these, newly synthesized subunits exhibit a specifically increased turnover rate. We suggest that association of these two subunits with the complex may be mediated or enhanced by interaction with other subunits, in particular cytochrome b.

Amino Acid Sequence↗

Horizontal gene transfer in bacterial and archaeal complete genomes.

There is growing evidence that horizontal gene transfer is a potent evolutionary force in prokaryotes, although exactly how potent is not known. We have developed a statistical procedure for predicting whether genes of a complete genome have been acquired by horizontal gene transfer. It is based on the analysis of G+C contents, codon usage, amino acid usage, and gene position. When we applied this procedure to 17 bacterial complete genomes and seven archaeal ones, we found that the percentage of horizontally transferred genes varied from 1.5% to 14.5%. Archaea and nonpathogenic bacteria had the highest percentages and pathogenic bacteria, except for Mycoplasma genitalium, had the lowest. As reported in the literature, we found that informational genes were less likely to be transferred than operational genes. Most of the horizontally transferred genes were only present in one or two lineages. Some of these transferred genes include genes that form part of prophages, pathogenecity islands, transposases, integrases, recombinases, genes present only in one of the two Helicobacter pylori strains, and regions of genes functionally related. All of these findings support the important role of horizontal gene transfer in the molecular evolution of microorganisms and speciation.

Amino Acids↗

Nucleotide sequences from the colicin E8 operon: homology with plasmid ColE2-P9.

The primary structures of the immunity (Imm) and lysis (Lys) proteins, and the C-terminal 205 amino acid residues of colicin E8 were deduced from nucleotide sequencing of the 1,265 bp ClaI-PvuI DNA fragment of plasmid ColE8-J. The gene order is col-imm-lys confirming previous genetic data. A comparison of the colicin E8 peptide sequence with the available colicin E2-P9 sequence shows an identical receptor-binding domain but 20 amino acid replacements and a clustering of synonymous codon usage in the nuclease-active region. Sequence homology of the two colicins indicates that they are descended from a common ancestral gene and that colicin E8, like colicin E2, may also function as a DNA endonuclease. The native ColE8 imm (resident copy) is 258 bp long and is predicted to encode an acidic protein of 9,604 mol. wt. The six amino acid replacements between the resident imm and the previously reported non-resident copy of the ColE8 imm ([E8 imm]) found in the ribonuclease-producing ColE3-CA38 plasmid offer an explanation for the incomplete protection conferred by [E8 Imm] to exogenously added colicin E8. Except for one nucleotide and amino acid change in the putative signal peptide sequence, the ColE8 lys structure is identical to that present in ColE2-P9 and ColE3-CA38.

Amino Acid Sequence↗

Lateral gene transfer and ancient paralogy of operons containing redundant copies of tryptophan-pathway genes in Xylella species and in heterocystous cyanobacteria.

BACKGROUND: Tryptophan-pathway genes that exist within an apparent operon-like organization were evaluated as examples of multi-genic genomic regions that contain phylogenetically incongruous genes and coexist with genes outside the operon that are congruous. A seven-gene cluster in Xylella fastidiosa includes genes encoding the two subunits of anthranilate synthase, an aryl-CoA synthetase, and trpR. A second gene block, present in the Anabaena/Nostoc lineage, but not in other cyanobacteria, contains a near-complete tryptophan operon nested within an apparent supraoperon containing other aromatic-pathway genes. RESULTS: The gene block in X. fastidiosa exhibits a sharply delineated low-GC content. This, as well as bias of codon usage and 3:1 dinucleotide analysis, strongly implicates lateral gene transfer (LGT). In contrast, parametric studies and protein tree phylogenies did not support the origination of the Anabaena/Nostoc gene block by LGT. CONCLUSIONS: Judging from the apparent minimal amelioration, the low-GC gene block in X. fastidiosa probably originated by LGT at a relatively recent time. The surprising inability to pinpoint a donor lineage still leaves room for alternative, albeit less likely, explanations other than LGT. On the other hand, the large Anabaena/Nostoc gene block does not seem to have arisen by LGT. We suggest that the contemporary Anabaena/Nostoc array of divergent paralogs represents an ancient ancestral state of paralog divergence, with extensive streamlining by gene loss occurring in the lineage of descent representing other (unicellular) cyanobacteria.

Amino Acid Sequence↗

Multiple divergent mRNAs code for a single human calmodulin.

The isolation of a novel complementary DNA (cDNA) clone coding for human calmodulin (CaM) is reported. Although it encodes a protein indistinguishable from the only known higher vertebrate calmodulin, its nucleotide sequence varies extensively from that of two previously reported human CaM cDNAs (Wawrzynczak and Perham, 1984; SenGupta et al., 1987). Only 82 and 81% identity, respectively, is found between the newly isolated and the two known human mRNAs in their coding regions. No striking homology is present in their noncoding regions. Codon usage in the three CaM mRNAs is also surprisingly divergent. A 2.3-kilobase mRNA corresponding to the newly isolated clone is expressed to varying extents in several human tissues, together with an approximately 0.8-kilobase mRNA species presumably arising from alternative polyadenylation of the same primary transcript. The results indicate that the human genome contains at least three divergent CaM genes that are under selective pressure to encode an identical protein while maintaining maximally divergent nucleotide sequences. Partial characterization of a genomic clone specifying the 3' portion of the newly identified CaM mRNA shows that this gene contains introns at identical positions as the previously characterized bona fide vertebrate CaM genes. Evolutionary implications of the presence of a CaM multigene family are discussed.

Amino Acid Sequence↗

SUF12 suppressor protein of yeast. A fusion protein related to the EF-1 family of elongation factors.

Mutations at the suf12 locus were isolated in Saccharomyces cerevisiae as extragenic suppressors of +1 frameshift mutations in glycine (GGX) and proline (CCX) codons, as well as UGA and UAG nonsense mutations. To identify the SUF12 function in translation and to understand the relationship between suf12-mediated misreading and translational frameshifting, we have isolated an SUF12+ clone from a centromeric plasmid library by complementation. SUF12+ is an essential, single-copy gene that is identical with the omnipotent suppressor gene SUP35+. The 2.3 x 10(3) base SUF12+ transcript contains an open reading frame sufficient to encode a 88 x 10(3) Mr protein. The pattern of codon usage and transcript abundance suggests that SUF12+ is not a highly expressed gene. The linear SUF12 amino acid sequence suggests that SUF12 has evolved as a fusion protein of unique N-terminal domains fused to domains that exhibit essentially co-linear homology to the EF-1 family of elongation factors. Beginning internally at amino acid 254, homology is more extensive between the SUF12 protein and EF-1 alpha of yeast (36% identity; 65% with conservative substitutions) than between EF-1 alpha of yeast and EF-Tu of Escherichia coli. The most extensive regions of SUF12/EF-1 alpha homology are those regions that have been conserved in the EF-1 family, including domains involved in GTP and tRNA binding. It is clear that SUF12 and EF-1 alpha are not functionally equivalent, since both are essential in vivo. The N-terminal domains of SUF12 are unique and may reflect, in part, the functional distinction between these proteins. These domains exhibit unusual amino acid composition and extensive repeated structure. The behavior of suf12-null/SUF12+ heterozygotes indicates that suf12 is co-dominantly expressed and suggests that suf12 allele-specific suppression may result from functionally distinct mutant proteins rather than variation in residual wild-type SUF12+ activity. We propose a model of suf12-mediated frameshift and nonsense suppression that is based on a primary defect in the normal process of codon recognition.

Amino Acid Sequence↗

Genetic characterization and partial sequence determination of a Treponema pallidum operon expressing two immunogenic membrane proteins in Escherichia coli.

A detailed physical and genetic map of a previously cloned 5.5-kilobase segment of Treponema pallidum DNA is described. This segment expressed two proteins that are cell membrane associated in Escherichia coli. The structural genes of these treponemal membrane proteins, tmpA and tmpB, are coordinately expressed, and transcription in E. coli can start from at least two different treponemal promoters. The tmpA and tmpB proteins are the products of in vivo proteolytic cleavage from precursor proteins which are 2 and 4 kilodaltons larger, respectively, than the mature proteins. Because the sizes of the corresponding proteins produced in T. pallidum were identical to those of the mature membrane proteins in E. coli, we concluded that a similar proteolytic processing takes place in both E. coli and T. pallidum. Although tmpA and tmpB were controlled by the same transcription signals, tmpB was expressed to a higher extent than tmpA, and only the tmpB product could be overproduced by placing the left lambda promoter in front of the structural genes. The nucleotide sequence of the T. pallidum tmpA gene was established. This is the first T. pallidum gene sequenced. Codon usage and the nature of transcriptional and translational signals are discussed. The deduced amino acid sequence indicated the presence of a sequence that was characteristic for a signal peptide. This sequence information allowed the construction of hybrid genes coding for proteins having beta-galactosidase enzyme activity as well as TmpA epitopes. The enzyme-linked antigen was expressed at a high level in E. coli when transcriptional and translational signals from coliphage lambda were used. In this case the protein produced was a sandwich protein consisting of 21 amino acids of the lambda cro protein, 204 amino acids of the T. pallidum TmpA protein, and 1,020 amino acids of the E. coli lambda-galactosidase. The potential use of this enzyme-linked antigen for the serodiagnosis of syphilis is discussed.

Amino Acid Sequence↗