Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Structure and comparative analysis of the genes encoding component C of methyl coenzyme M reductase in the extremely thermophilic archaebacterium Methanothermus fervidus.

A 6-kilobase-pair (kbp) region of the genome of the extremely thermophilic arachaebacterium Methanothermus fervidus which encodes the alpha, beta, and gamma subunit polypeptides of component C of methyl coenzyme M reductase was cloned and sequenced. Genes encoding the beta (mcrB) and gamma (mcrG) subunits were separated by two open reading frames (designated mcrC and mcrD) which encode unknown gene products. The M. fervidus genes were preceded by ribosome-binding sites, separated by short A + T-rich intergenic regions, contained unexpectedly few NNC codons, and exhibited inflexible codon usage at some locations. Sites of transcription initiation and termination flanking the mcrBDCGA cluster of genes in M. fervidus were identified. The sequences of the genes, the encoded polypeptides, and transcription regulatory signals in M. fervidus were compared with the functionally equivalent sequences from two mesophilic methanogens (Methanococcus vannielii and Methanosarcina barkeri) and from a moderate thermophile (Methanobacterium thermoautotrophicum Marburg). The amino acid sequences of the polypeptides encoded by the mcrBCGA genes in the two thermophiles were approximately 80% identical, whereas all other pairs of these gene products contained between 50 and 60% identical amino acid residues. The mcrD gene products have diverged more than the products of the other mcr genes. Identification of highly conserved regions within mcrA and mcrB suggested oligonucleotide sequences which might be developed as hybridization probes which could be used for identifying and quantifying all methanogens.

Amino Acid Sequence↗

Sequences and homology analysis of two genes encoding beta-glucosidases from Bacillus polymyxa.

The nucleotide sequences of the bglA and bglB genes encoding beta-glucosidases from Bacillus polymyxa have been determined. Both genes contain coding regions of 1344 bp, corresponding to polypeptides with Mrs of 51,643 and 51,547, respectively. Patterns of codon usage indicate that both genes are expressed at a low frequency. Previous data suggested that the proteins encoded by bglA and bglB were intra- and extracellular enzymes, respectively; however, neither of the two deduced amino acid sequences has N termini with the typical features of a leader peptide. The proteins encoded by bglA and bglB show remarkable homology to each other and to other beta-glucosidases (Bgl) and beta-galactosidases (beta Gal). On the basis of the observed homologies, we can define two groups of microbial Bgl: one of them, type I, including most bacterial Bgl, and type II, including enzymes from different yeast species and one from Clostridium thermocellum. Likewise, at least two groups of beta Gal can be distinguished: type I, including enzymes homologous to type-I Bgl, and type II, showing no homology to any of the previous groups.

Amino Acid Sequence↗

Codon-improved Cre recombinase (iCre) expression in the mouse.

By applying the mammalian codon usage to Cre recombinase, we improved Cre expression, as determined by immunoblot and functional analysis, in three different mammalian cell lines. The improved Cre (iCre) gene was also designed to reduce the high CpG content of the prokaryotic coding sequence, thereby reducing the chances of epigenetic silencing in mammals. Transgenic iCre expressing mice were obtained with good frequency, and in these mice loxP-mediated DNA recombination was observed in all cells expressing iCre. Moreover, iCre fused to two estrogen receptor hormone binding domains for temporal control of Cre activity could also be expressed in transgenic mice. However, Cre induction after administration of tamoxifen yielded only low Cre activity. Thus, whereas efficient activation of Cre fusion proteins in the brain needs further improvements, our studies indicate that iCre should facilitate genetic experiments in the mouse.

Animals↗

Design and expression of a synthetic phyC gene encoding the neutral phytase in Pichia pastoris.

The 1074-bp phyCs gene (optimized phyC gene) encoding neutral phytase was designed and synthesized according to the methylotrophic yeast Pichia pastoris codon usage bias without altering the protein sequence. The expression vector, pP9K-phyCs, was linearized and transformed in P. pastoris. The yield of total extracellular phytase activity was 17.6 U/ml induced in Buffered Methanol-complex Medium (BMMY) and 18.5 U/ml in Wheat Bran Extract Induction (WBEI) medium at the flask scale, respectively, improving over 90 folds compared with the wild-type isolate. Purified enzyme showed temperature optimum of 70 degrees and pH optimum of 7.5. The enzyme activity retained 97% of the relative activity after incubation at 80 degrees for 5 min. Because of the heavy glycosylation the expressed phytase had a molecular size of approximately 51 kDa. After deglycosylation by endoglycosylase H (EndoH(f)), the enzyme had an apparent molecular size of 42 kDa. Its property and thermostability was affected by the glycosylation.

6-Phytase↗

The causes of synonymous rate variation in the rodent genome. Can substitution rates be used to estimate the sex bias in mutation rate?

Miyata et al. have suggested that the male-to-female mutation rate ratio (alpha) can be estimated by comparing the neutral substitution rates of X-linked (X), Y-linked (Y), and autosomal (A) genes. Rodent silent site X/A comparisons provide very different estimates from X/Y comparisons. We examine three explanations for this discrepancy: (1) statistical biases and artifacts, (2) nonneutral evolution, and (3) differences in mutation rate per germline replication. By estimating errors and using a variety of methodologies, we tentatively reject explanation 1. Our analyses of patterns of codon usage, synonymous rates, and nonsynonymous rates suggest that silent sites in rodents are evolving neutrally, and we can therefore reject explanation 2. We find both base composition and methylation differences between the different sets of chromosomes, a result consistent with explanation 3, but these differences do not appear to explain the observed discrepancies in estimates of alpha. Our finding of significantly low synonymous substitution rates in genomically imprinted genes suggests a link between hemizygous expression and an adaptive reduction in the mutation rate, which is consistent with explanation 3. Therefore our results provide circumstantial evidence in favor of the hypothesis that the discrepancies in estimates of alpha are due to differences in the mutation rate per germline replication between different parts of the genome. This explanation violates a critical assumption of the method of Miyata et al., and hence we suggest that estimates of alpha, obtained using this method, need to be treated with caution.

Animals↗

Understanding the adaptation of Halobacterium species NRC-1 to its extreme environment through computational analysis of its genome sequence.

The genome of the halophilic archaeon Halobacterium sp. NRC-1 and predicted proteome have been analyzed by computational methods and reveal characteristics relevant to life in an extreme environment distinguished by hypersalinity and high solar radiation: (1) The proteome is highly acidic, with a median pI of 4.9 and mostly lacking basic proteins. This characteristic correlates with high surface negative charge, determined through homology modeling, as the major adaptive mechanism of halophilic proteins to function in nearly saturating salinity. (2) Codon usage displays the expected GC bias in the wobble position and is consistent with a highly acidic proteome. (3) Distinct genomic domains of NRC-1 with bacterial character are apparent by whole proteome BLAST analysis, including two gene clusters coding for a bacterial-type aerobic respiratory chain. This result indicates that the capacity of halophiles for aerobic respiration may have been acquired through lateral gene transfer. (4) Two regions of the large chromosome were found with relatively lower GC composition and overrepresentation of IS elements, similar to the minichromosomes. These IS-element-rich regions of the genome may serve to exchange DNA between the three replicons and promote genome evolution. (5) GC-skew analysis showed evidence for the existence of two replication origins in the large chromosome. This finding and the occurrence of multiple chromosomes indicate a dynamic genome organization with eukaryotic character.

Adaptation, Biological↗

PF-IND: probability algorithm and software for separation of plant and fungal sequences.

The separation of plant and fungal sequences in EST pools by bioinformatic methods is difficult because of sequence similarities between plants and fungi, lack of enough sequence information, and the short length of the isolated fragments. An algorithm and software that utilize the differences in codon usage bias to discriminate between plant and fungal sequences are described. The software (PF-IND) includes five pairs of fungi and their host plants that can be used to analyze a large number of related species. Analysis of a sequence provides an arbitrary value that defines the likelihood that a sequence will be a fungal or a plant gene. The software can distinguish between homologous fungal and plant genes and it helps identify the correct reading frame of unknown expressed sequence tags (ESTs) for which BLAST analyses do not provide clear information. Short sequences of 100-150 bp can be analyzed with high confidence. PF-IND analysis of 100 sequences derived from fungal infected plants identified the origin of 94 sequences. Only 66 sequences were identified by a BLASTX analysis of the same 100 ESTs. Overall, PF-IND is a novel bioinformatic tool aimed at assisting the research of fungus-plant interactions.

Algorithms↗

Comparison between Pyrococcus horikoshii and Pyrococcus abyssi genome sequences reveals linkage of restriction-modification genes with large genome polymorphisms.

Recent work suggests that restriction-modification gene complexes are mobile genetic elements that insert themselves into the genome and cause various genome rearrangements. In the present work, the complete genome sequences of Pyrococcus horikoshii and Pyrococcus abyssi, two species in a genus of hyperthermophilic archaeon (archaebacterium), were compared to detect large genome polymorphisms linked with restriction-modification gene homologs. Sequence alignments, GC content analysis, and codon usage analysis demonstrated the diversity of these homologs and revealed a possible case of relatively recent acquisition (horizontal transfer). In two cases out of the six large polymorphisms identified, there was insertion of a DNA segment with a modification gene homolog, accompanied by target deletion (simple substitution). In two other cases, homologous DNA segments carrying a modification gene homolog were present at different locations in the two genomes (transposition). In both cases, substitution (insertion/deletion) in one of the two loci was accompanied by inversion of adjacent chromosomal segment. In the fifth case, substitution by a DNA segment carrying type I restriction, modification, and specificity gene homologs was likewise accompanied by adjacent inversion. In the last case, two homologous DNA segments, were found at different loci in the two genomes (transposition), but only one of them had insertion of a modification homolog and an unknown ORF. The possible relationship of these polymorphisms to attack by restriction enzymes on the chromosome will be discussed.

Chromosome Inversion↗

Structure of genes and an insertion element in the methane producing archaebacterium Methanobrevibacter smithii.

DNA fragments cloned from the methanogenic archaebacterium Methanobrevibacter smithii which complement mutations in the purE and proC genes of E. coli have been sequenced. Sequence analyses, transposon mutagenesis and expression in E. coli minicells indicate that purE and proC complementations result from the synthesis of M. smithii polypeptides with molecular weights of 36,697 and 27,836 respectively. The encoding genes appear to be located in operons. The M. smithii genome contains 69% A/T basepairs (bp) which is reflected in unusual codon usages and intergenic regions containing approximately 85% A/T bp. An insertion element, designated ISM1, was found within the cloned M. smithii DNA located adjacent to the proC complementing region. ISM1 is 1381 bp in length, has 29 bp terminal inverted repeat sequences and contains one major ORF encoded in 87% of the ISM1 sequence. ISM1 is mobile, present in approximately 10 copies per genome and integration duplicates 8 bp at the site of insertion. The duplicated sequences show homology with sequences within the 29 bp terminal repeat sequence of ISM1. Comparison of our data with sequences from halophilic archaebacteria suggests that 5'GAANTTTCA and 5'TTTTAATATAAA may be consensus promoter sequences for archaebacteria. These sequences closely resemble the consensus sequences which precede Drosophila heat-shock genes (Pelham 1982; Davidson et al. 1983). Methanogens appear to employ the eubacterial system of mRNA: 16SrRNA hybridization to ensure initiation of translation; the consensus ribosome binding sequence is 5'AGGTGA.

Amino Acid Sequence↗

Nucleocapsid protein of cell culture-adapted Seoul virus strain 80-39: analysis of its encoding sequence, expression in yeast and immuno-reactivity.

Seoul virus (SEOV) is a hantavirus causing a mild to moderate form of hemorrhagic fever with renal syndrome that is distributed mainly in Asia. The nucleocapsid (N) protein-encoding sequence of SEOV (strain 80-39) was RT-PCR-amplified and cloned into a yeast expression vector containing a galactose-inducible promoter. A survey of the pattern of synonymous codon preferences for a total of 22 N protein-encoding hantavirus genes including 13 of SEOV strains revealed that there is minor variation in codon usage by the same gene in different viral genomes. Introduction of the expression plasmid into yeast Saccharomyces cerevisiae resulted in the high-level expression of a hexahistidine-tagged N protein derivative. The nickel-chelation chromatography purified, yeast-expressed SEOV N protein reacted in the immunoblot with a SEOV-specific monoclonal antibody and certain HTNV- and PUUV-cross-reactive monoclonal antibodies. The immunization of a rabbit with the recombinant N protein resulted in the induction of a high-titered antibody response. In ELISA studies, the N protein was able to detect antibodies in sera of experimentally infected laboratory rats and in human anti-hantavirus-positive sera or serum pools of patients from different geographical origin. The yeast-expressed SEOV N protein represents a promising antigen for development of diagnostic tools in serology, sero prevalence studies and vaccine development.

Animals↗

The human COL11A2 gene structure indicates that the gene has not evolved with the genes for the major fibrillar collagens.

The human COL11A2 gene was analyzed from two overlapping cosmid clones that were previously isolated in the course of searching the human major histocompatibility region (Janatipour, M., Naumov, Y., Ando, A., Sugimura, K., Okamoto, N., Tsuji, K., Abe, K., and Inoko, H. (1992) Immunogenetics 35, 272-278). Nucleotide sequencing defined over 28,000 base pairs of the gene. It was shown to contain 66 exons. As with most genes for fibrillar collagens, the first intron was among the largest, and the introns at the 5'-end of the gene were in general larger than the introns at the 3'-end. Analysis of the exons coding for the major triple helical domain indicated that the gene structure had not evolved with the genes for the major fibrillar collagens in that there were marked differences in the number of exons, the exon sizes, and codon usage. The gene was located close to the gene for the retinoic X receptor beta in a head-to-tail arrangement similar to that previously seen with the two mouse genes (P. Vandenberg and D. J. Prockop, submitted for publication). Also, there was marked interspecies homology in the intergenic sequences. The amino acid sequences and the pattern of charged amino acids in the major triple helix of the alpha 2(XI) chain suggested that the chain can be incorporated into the same molecule as alpha 1(XI) and alpha 1(V) chains but not into the same molecule as the alpha 3(XI)/alpha 1(II) chain. The structure of the carboxyl-terminal propeptide was similar to the carboxyl-terminal propeptides of the pro alpha 1(XI) chain and pro alpha chains of other fibrillar collagens, but it was shorter because of internal deletions of about 30 amino acids.

Amino Acid Sequence↗

Biochemical indication for myristoylation-dependent conformational changes in HIV-1 Nef.

The accessory HIV-1 Nef protein is essential for viral replication, high virus load, and progression to AIDS. These functions are mediated by the alteration of signaling and trafficking pathways and require the membrane association of Nef by its N-terminal myristoylation. However, a large portion of Nef is also found in the cytosol, in line with the observation that myristoylation is only a weak lipidation anchor for membrane attachment. We performed biochemical studies to analyze the implications of myristoylation on the conformation of Nef in aqueous solution. To establish an in vivo myristoylation assay, we first optimized the codon usage of Nef for Escherichia coli expression, which resulted in a 15-fold higher protein yield. Myristoylation was achieved by coexpression with the N-myristoyltransferase and confirmed by mass spectrometry. The myristoylated protein was soluble, and proton NMR spectra confirmed proper folding. Size exclusion chromatography revealed that myristoylated Nef appeared of smaller size than the unmodified form but not as small as an N-terminally truncated from of Nef that omits the anchor domain. Western blot stainings and limited proteolysis of both forms showed different recognition profiles and degradation pattern. Analytical ultracentrifugation revealed that myristoylated Nef prevails in a monomeric state while the unmodified form exists in an oligomeric equilibrium of monomer, dimer, and trimer associations. Finally, fluorescence correlation spectroscopy using multiphoton excitation revealed a shorter diffusion time for the lipidated protein compared to the unmodified form. Taken together, our data indicated myristoylation-dependent conformational changes in Nef, suggesting a rather compact and monomeric form for the lipidated protein in solution.

Base Sequence↗

Predicted highly expressed and putative alien genes of Deinococcus radiodurans and implications for resistance to ionizing radiation damage.

Predicted highly expressed (PHX) and putative alien genes determined by codon usages are characterized in the genome of Deinococcus radiodurans (strain R1). Deinococcus radiodurans (DEIRA) can survive very high doses of ionizing radiation that are lethal to virtually all other organisms. It has been argued that DEIRA is endowed with enhanced repair systems that provide protection and stability. However, predicted expression levels of DNA repair proteins with the exception of RecA tend to be low and do not distinguish DEIRA from other prokaryotes. In this paper, the capability of DEIRA to resist extreme doses of ionizing and UV radiation is attributed to an unusually high number of PHX chaperone/degradation, protease, and detoxification genes. Explicitly, compared with all current complete prokaryotic genomes, DEIRA contains the greatest number of PHX detoxification and protease proteins. Other sources of environmental protection against severe conditions of UV radiation, desiccation, and thermal effects for DEIRA are the several S-layer (surface structure) PHX proteins. The top PHX gene of DEIRA is the multifunctional tricarboxylic acid (TCA) gene aconitase, which, apart from its role in respiration, also alerts the cell to oxidative damage.

Chromosomes, Bacterial↗

Near-critical behavior of aminoacyl-tRNA pools in E. coli at rate-limiting supply of amino acids.

The rates of consumption of different amino acids in protein synthesis are in general stoichiometrically coupled with coefficients determined by codon usage frequencies on translating ribosomes. We show that when the rates of synthesis of two or more amino acids are limiting for protein synthesis and exactly matching their coupled rates of consumption on translating ribosomes, the pools of aminoacyl-tRNAs in ternary complex with elongation factor Tu and GTP are hypersensitive to a variation in the rate of amino acid supply. This high sensitivity makes a macroscopic analysis inconclusive, because it is accompanied by almost free and anticorrelated diffusion in copy numbers of ternary complexes. This near-critical behavior is relevant for balanced growth of Escherichia coli cells in media that lack amino acids and for adaptation of E. coli cells after downshifts from amino-acid-containing to amino-acid-lacking growth media. The theoretical results are used to discuss transcriptional control of amino acid synthesis during multiple amino acid limitation, the recovery of E. coli cells after nutritional downshifts and to propose a robust mechanism for the regulation of RelA-dependent synthesis of the global effector molecule ppGpp.

Algorithms↗

Synthesis of a lacI gene analogue with reduced CpG content.

A lacI gene analogue with reduced CpG content has been synthesized. Codon usage in the lacI gene was manipulated to remove most CpG sites (82/95; 86%) while maintaining wild-type amino acid sequence. The double-stranded gene sequence was synthesized using standard beta-cyanoethyl phosphoramidite chemistry and subsequently cloned into pBR322. Bacterial promoter sequences with different levels of activity were attached upstream of the modified coding region to study its expression in E. coli. Production of lacI protein was confirmed in a lacI- E. coli strain by Western blot analysis and by measuring repression of the lacZ gene with the chromogenic lacZ indicator, 5-bromo-4-chloro-3-indolyl-beta-D-galactopyranoside (X-gal). The modified lacI gene construct can be used as a genetic target in cultured mammalian cells or in transgenic animals to avoid high levels of background mutation associated with methylated CpG sequences. The construction scheme described here provides a general approach to remove CpG sequences from gene constructs when methylation is undesirable.

Amino Acid Sequence↗

Isolation of alpha- and beta-tubulin genes of Plasmodium falciparum using a single oligonucleotide probe.

An oligonucleotide probe (315) specific for the alpha- and beta-tubulin genes of Plasmodium falciparum was synthesized utilizing codon usage of P. falciparum determined from published gene sequences. By screening genomic and cDNA libraries with the oligonucleotide probe, alpha- and beta-tubulin clones were isolated. Positive clones were identified by partial sequencing and comparing the deduced amino acid sequence with the chicken brain alpha- and beta-tubulin amino acid sequences. The beta-tubulin gene was completely sequenced at the genomic level and partially at cDNA level. The deduced polypeptide is 445 amino acids long, shares 88% homology with chicken brain beta-tubulin, and contains two introns of 362 and 163 bp long, respectively. alpha- and beta-tubulin genes of P. falciparum are unlinked and dispersed; more than one copy of each gene may be present. Northern blot analysis of total RNA of the blood-stage parasite indicates the presence of three transcripts of alpha-tubulin (3.3 kb, 2.6 kb, 1.9 kb) and three transcripts of beta-tubulin gene (3.6 kb, 2.9 kb, 2.0 kb). The significance of these transcripts is presently unknown.

Amino Acid Sequence↗

Expression and purification of the synthetic preS1 gene of Hepatitis B Virus with preferred Escherichia coli codon preference.

To produce high levels of hepatitis B virus (HBV) preS1 protein at low cost, a DNA fragment encoding the preS1 region, residues 1-119, of HBV adr subtype was synthesized by overlapping-PCR according to Escherichia coli (E. coli) B preferred codon usage. The synthetic preS1 gene (spreS1) was cloned into the bacterial expression vector pET-30a and transferred into the expression strain E. coli BL21(DE3). Recombinant preS1 protein with an N-terminal His6 tag was expressed at high levels in soluble form, yielding about 44% of the total cellular protein. This technique overcomes problems that existed in previously reported expression systems of preS1 or its epitope, i.e., low-level expression or expression in inclusion bodies. Using this His-tagged preS1 expression system, recombinant protein was purified by single-step affinity chromatography on a Ni-NTA column resulting in a yield was about 28 mg recombinant protein per liter culture. Furthermore, Western blotting and indirect ELISA analysis demonstrate that the reactivity of preS1-specific antibody is comparable between the recombinant and commercialized preS1 protein. Thus, our improved expression system could be used for practical, low-cost large-scale production of recombinant preS1 without refolding steps.

Base Sequence↗

Synthesis of cholera toxin B subunit gene: cloning and expression of a functional 6XHis-tagged protein in Escherichia coli.

Cholera toxin B subunit (CTB) has been extensively studied as immunogen, adjuvant, and oral tolerance inductor depending on the antigen conjugated or coadministered. It has been already expressed in several bacterial and yeast systems. In this study, we synthesized a versatile gene coding a 6XHis-tagged CTB (359bp). The sequence was designed according to codon usage of Escherichia coli, Lactobacillus casei, and Salmonella typhimurium. The gene assembly was based on a polymerase chain reaction, in which the polymerase extends DNA fragments from a pool of overlapping oligonucleotides. The synthetic gene was amplified, cloned, and expressed in E. coli in an insoluble form, reaching levels about 13 mg of purified active pentameric rCTB per liter of induced culture. Western blot and ELISA analyses showed that recombinant CTB is strongly and specifically recognized by polyclonal antibodies against the cholera toxin. The ability to form the functional pentamers was observed in cell culture by the inhibition of cholera toxin activity on Y1 adrenal cells in the presence of recombinant CTB. The 6XHis-tagged CTB provides a simple way to obtain functional CTB through Ni(2+)-charged resin after refolding and also free of possible CTA contaminants as in the case of CTB obtained from Vibrio cholerae cultures.

Adrenal Glands↗