Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Complete genome structure of the nitrogen-fixing symbiotic bacterium Mesorhizobium loti.

The complete nucleotide sequence of the genome of a symbiotic bacterium Mesorhizobium loti strain MAFF303099 was determined. The genome of M. loti consisted of a single chromosome (7,036,071 bp) and two plasmids, designated as pMLa (351,911 bp) and pMLb (208, 315 bp). The chromosome comprises 6752 potential protein-coding genes, two sets of rRNA genes and 50 tRNA genes representing 47 tRNA species. Fifty-four percent of the potential protein genes showed sequence similarity to genes of known function, 21% to hypothetical genes, and the remaining 25% had no apparent similarity to reported genes. A 611-kb DNA segment, a highly probable candidate of a symbiotic island, was identified, and 30 genes for nitrogen fixation and 24 genes for nodulation were assigned in this region. Codon usage analysis suggested that the symbiotic island as well as the plasmids originated and were transmitted from other genetic systems. The genomes of two plasmids, pMLa and pMLb, contained 320 and 209 potential protein-coding genes, respectively, for a variety of biological functions. These include genes for the ABC-transporter system, phosphate assimilation, two-component system, DNA replication and conjugation, but only one gene for nodulation was identified.

ATP-Binding Cassette Transporters↗

Insights into a dinoflagellate genome through expressed sequence tag analysis.

BACKGROUND: Dinoflagellates are important marine primary producers and grazers and cause toxic "red tides". These taxa are characterized by many unique features such as immense genomes, the absence of nucleosomes, and photosynthetic organelles (plastids) that have been gained and lost multiple times. We generated EST sequences from non-normalized and normalized cDNA libraries from a culture of the toxic species Alexandrium tamarense to elucidate dinoflagellate evolution. Previous analyses of these data have clarified plastid origin and here we study the gene content, annotate the ESTs, and analyze the genes that are putatively involved in DNA packaging. RESULTS: Approximately 20% of the 6,723 unique (11,171 total 3'-reads) ESTs data could be annotated using Blast searches against GenBank. Several putative dinoflagellate-specific mRNAs were identified, including one novel plastid protein. Dinoflagellate genes, similar to other eukaryotes, have a high GC-content that is reflected in the amino acid codon usage. Highly represented transcripts include histone-like (HLP) and luciferin binding proteins and several genes occur in families that encode nearly identical proteins. We also identified rare transcripts encoding a predicted protein highly similar to histone H2A.X. We speculate this histone may be retained for its role in DNA double-strand break repair. CONCLUSION: This is the most extensive collection to date of ESTs from a toxic dinoflagellate. These data will be instrumental to future research to understand the unique and complex cell biology of these organisms and for potentially identifying the genes involved in toxin production.

Amino Acid Sequence↗

Nucleotide sequence of the pnp gene of Escherichia coli encoding polynucleotide phosphorylase. Homology of the primary structure of the protein with the RNA-binding domain of ribosomal protein S1.

The pnp gene is located at 69 min on the Escherichia coli chromosome adjacent to the rpsO gene which encodes the ribosomal protein S15. In this paper, we present the sequence of a 3030-nucleotide DNA fragment containing the open reading frames coding for ribosomal protein S15 and polynucleotide phosphorylase. Translation of pnp is initiated by 5'-UUG-3' codon separated by 7 nucleotides from a good ribosome binding site. Codon usage in this gene is typical of highly expressed proteins of E. coli. Some of the transcripts of the pnp gene terminate just after the stem of the terminator t2 visible in the nucleotide sequence. However, a very strong read-through occurs at this site, thus permitting many of the pnp transcripts to extend beyond this transcription terminator. We also describe the primary structure homologies between a 69-amino-acid stretch of polynucleotide phosphorylase and the four homologous stretches of ribosomal protein S1 which form its RNA binding site. The possibility that this 69-amino-acid stretch constitutes the polynucleotide binding domain of polynucleotide phosphorylase is discussed.

Amino Acid Sequence↗

Genomic structures of viral agents in relation to the biosynthesis of selenoproteins.

The genomes of both bacteria and eukaryotic organisms are known to encode selenoproteins, using the UGA codon for seleno-cysteine (SeC), and a complex cotranslational mechanism for SeC incorporation into polypeptide chains, involving RNA stem-loop structures. These common features and similar codon usage strongly suggest that this is an ancient evolutionary development. However, the possibility that some viruses might also encode selenoproteins remained unexplored until recently. Based on an analysis of the genomic structure of the human immunodeficiency virus HIV-1, we demonstrated that several regions overlapping known HIV genes have the potential to encode selenoproteins (Taylor et al. [31], J. Med. Chem. 37, 2637-2654 [1994]). This is provocative in the light of overwhelming evidence of a role for oxidative stress in AIDS pathogenesis, and the fact that a number of viral diseases have been linked to selenium (Se) deficiency, either in humans or by in vitro and animal studies. These include HIV-AIDS, hepatitis B linked to liver disease and cancer, Coxsackie virus B3, Keshan disease, and the mouse mammary tumor virus (MMTV), against which Se is a potent chemoprotective agent. There are also established biochemical mechanisms whereby extreme Se deficiency can induce a proclotting or hemorrhagic effect, suggesting that hemorrhagic fever viruses should also be examined for potential virally encoded selenoproteins. In addition to the RNA stem-loop structures required for SeC insertion at UGA codons, genomic structural features that may be required for selenoprotein synthesis can also include ribosomal frameshift sites and RNA pseudoknots if the potential selenoprotein module overlaps with another gene, which may prove to be the rule rather than the exception in viruses. One such pseudoknot that we predicted in HIV-1 has now been verified experimentally; a similar structure can be demonstrated in precisely the same location in the reverse transcriptase coding region of hepatitis B virus. Significant new findings reported here include the existence of highly distinctive glutathione peroxidase (GSH-Px)-related sequences in Coxsackie B viruses, new theoretical data related to a previously proposed potential selenoprotein gene overlapping the HIV protease coding region, and further evidence in support of a novel frameshift site in the HIV nef gene associated with a well-conserved UGA codon in the 1-reading frame.

Amino Acid Sequence↗

UpGene: Application of a web-based DNA codon optimization algorithm.

Although DNA codon optimization is a standard molecular biology strategy to overcome poor gene expression, to date no public software exists to facilitate this process. Among the uses of codon optimization, human immunodeficiency virus (HIV) vaccine development represents one of the most difficult challenges. A key obstacle to an effective DNA-based vaccine is the low-level expression of HIV genes in mammalian cells, which is due primarily to the instability of HIV mRNAs resulting from AU-rich elements and rare codon usage. In this report we describe the development of a DNA optimization algorithm integrated with a PCR primer design program to redesign specific coding sequences for maximal gene expression. Using this algorithm combination, together with PCR-based gene assembly, we have successfully optimized gene sequences for simian immunodeficiency virus (SIV) strain mac239 structural antigenic proteins gag and env, resulting in high-level gene expression in eukaryotic cells. Our findings demonstrate that our user-friendly algorithm is a valuable tool for DNA-based HIV vaccine development. Moreover, it can be used to optimize any other genes of interest and is freely available online at http://www.vectorcore.pitt.edu/upgene.html.

Algorithms↗

Cloning and expression of malarial pyrimidine enzymes.

We have cloned genes encoding three enzymes of the de novo pyrimidine pathway using genomic DNA from Plasmodium falciparum and sequence information from the Malarial Genome Project. Genes encoding dihydroorotase (reaction 3), orotate phosphoribosyltransferase (reaction 5), and OMP decarboxylase (reaction 6) have been cloned into the plasmid pET 3a or 3d with a thrombin cleavable 9xHis tag at the C-terminus and the enzymes were expressed in Escherichia coli. To overcome the toxicity of malarial OMP decarboxylase when expressed in E. coli, and the unusual codon usage of the malarial gene, a hybrid plasmid, pMICO, was constructed which expresses low levels of T7 lysozyme to inhibit T7 RNA polymerase used for recombinant expression, and extra copies of rare tRNAs. Catalytically-active OMP decarboxylase has been purified in tens of milligrams by chromatography on Ni-NTA. The gene encoding orotate phosphoribosyltransferase includes an extension of 66 amino acids from the N-terminus when compared with sequences for this enzyme from other organisms. We have found that other pyrimidine enzymes also contain unusual protein inserts. Milligram quantities of pure recombinant malarial enzymes from the pyrimidine pathway will provide targets for development of novel antimalarial drugs.

Animals↗

Neocallimastix frontalis enolase gene, enol: first report of an intron in an anaerobic fungus.

A DNA clone containing a putative enolase gene was isolated from a genomic DNA library of the anaerobic fungus Neocallimastix frontalis. It was deduced from sequence comparisons that the enolase gene was interrupted by a large 331 bp intron. The enolase gene, termed enol, has an ORF of 1308 bp and encodes a predicted 436 amino acid protein. The deduced amino acid sequence shows high identity (71.5-71%) to those of enolases from the yeasts Saccharomyces cerevisiae and Candida albicans. The G+C content of the enolase coding sequence (43.8 mol%) is considerably higher than the G+C content of the intervening sequence (14.2 mol%) or the 5' and 3' non-translated flanking sequences (15.2 and 4.7 mol%, respectively). The codon usage of the N. frontalis enolase gene was very biased as has been found for the highly expressed genes of yeast and filamentous fungi. The gene has all the canonical features (polyadenylation signal, intron splicing boundaries) of genes isolated from aerobic filamentous fungi. Only one enolase gene could be detected in N. frontalis genomic DNA by Southern analysis with a homologous probe. RNA analysis detected a single enolase transcript of about 1.6 kb. When mycelium was grown on glucose, levels of enolase mRNA were markedly increased by comparison with enolase mRNA levels in mycelium grown on cellulose, suggesting that expression of the N. frontalis enolase gene was transcriptionally regulated by the carbon source.

Base Sequence↗

Mitochondrial genomes of Vanhornia eucnemidarum (Apocrita: Vanhorniidae) and Primeuchroeus spp. (Aculeata: Chrysididae): Evidence of rearranged mitochondrial genomes within the Apocrita (Insecta: Hymenoptera).

We sequenced most of the mitochondrial (mt) genomes of 2 apocritan taxa: Vanhornia eucnemidarum and Primeuchroeus spp. These mt genomes have similar nucleotide composition and codon usage to those of mt genomes reported for other Hymenoptera, with a total A + T content of 80.1% and 78.2%, respectively. Gene content corresponds to that of other metazoan mt genomes, but gene organization is not conserved. There are a total of 6 tRNA genes rearranged in V. eucnemidarum and 9 in Primeuchroeus spp. Additionally, several noncoding regions were found in the mt genome of V. eucnemidarum, as well as evidence of a sustained gene duplication involving 3 tRNA genes. We also report an inversion of the large and small ribosomal RNA genes in Primeuchroeus spp. mt genome. However, none of the rearrangements reported are phylogenetically informative with respect to the current taxon sample.

Animals↗

OspA, a lipoprotein antigen of the obligate intracellular bacterial pathogen Piscirickettsia salmonis.

No effective recombinant vaccines are currently available for any rickettsial diseases. In this regard the first non-ribosomal DNA sequences from the obligate intracellular pathogen Piscirickettsia salmonis are presented. Genomic DNA isolated from Percoll density gradient purified P. salmonis, was used to construct an expression library in lambda ZAP II. In the absence of preexisting DNA sequence, rabbit polyclonal antiserum raised against P. salmonis, with a bias toward P. salmonis surface antigens, was used to identify immunoreactive clones. Catabolite repression of the lac promoter was required to obtain a stable clone of a 4,983 bp insert in Escherichia coli due to insert toxicity exerted by the accompanying radA open reading frame (ORF). DNA sequence analysis of the insert revealed 1 partial and 4 intact predicted ORF's. A 486 bp ORF, ospA, encoded a 17 kDa antigenic outer surface protein (OspA) with 62% amino acid sequence homology to the genus common 17 kDa outer membrane lipoprotein of Rickettsia prowazekii, previously thought confined to members of the genus Rickettsia. Palmitate incorporation demonstrated that OspA is posttranslationally lipidated in E. coli, albeit poorly expressed as a lipoprotein even after replacement of the signal sequence with the signal sequence from lpp (Braun lipoprotein) or the rickettsial 17 kDa homologue. To enhance expression, ospA was optimized for codon usage in E. coli by PCR synthesis. Expression of ospA was ultimately improved (approximately 13% of total protein) with a truncated variant lacking a signal sequence. High level expression (approximately 42% tot. prot.) was attained as an N-terminal fusion protein with the fusion product recovered as inclusion bodies in E. coli BL21. Expression of OspA in P. salmonis was confirmed by immunoblot analysis using polyclonal antibodies generated against a synthetic peptide of OspA (110-129) and a strong antibody response against OspA was detected in convalescent sera from coho salmon (Oncorhynchus kisutch).

Amino Acid Sequence↗

Quantitative exploration of the occurrence of lateral gene transfer by using nitrogen fixation genes as a case study.

Lateral gene transfer (LGT) is now accepted as an important factor in the evolution of prokaryotes. Establishment of the occurrence of LGT is typically attempted by a variety of methods that includes the comparison of reconstructed phylogenetic trees, the search for unusual GC composition or codon usage within a genome, and identification of similarities between distant species as determined by best blast hits. We explore quantitative assessments of these strategies to study the prokaryotic trait of nitrogen fixation, the enzyme-catalyzed reduction of N(2) to ammonia. Phylogenies constructed on nitrogen fixation genes are not in agreement with the tree-of-life based on 16S rRNA but do not conclusively distinguish between gene loss and LGT hypotheses. Using a series of analyses on a set of complete genomes, our results distinguish two structurally distinct classes of MoFe nitrogenases whose distribution cuts across lines of vertical inheritance and makes us believe that a conclusive case for LGT has been made.

Codon↗

Transcript mapping reveals different expression strategies for the bicistronic RNAs of the geminivirus wheat dwarf virus.

We have characterised the major transcripts of the Czech isolate of wheat dwarf virus (WDV-CJI) which show that WDV uses two different mechanisms for expressing overlapping open reading frames (ORFs). Mapping of the virion sense RNAs identified a single polyadenylated transcript of 1.1kb spanning the overlapping ORFs V1 and V2 which encode cell-cell spread functions and the coat protein respectively. This finding distinguishes WDV from other monocot-infecting geminiviruses studied so far which were shown to encode two 3' co-terminal transcripts capable of expressing either the V1 or V2 ORF. A survey of codon usage at the junction between the V1 and V2 ORF has led us to propose that translational frame shifting analogous to that in the yeast Ty element may occur. Analysis of polymerase chain reaction (PCR) amplified complementary sense cDNA clones has revealed the presence of mature spliced and unspliced RNAs which could encode products of an intron mediated C1:C2 ORF fusion or the C1 ORF product alone. Mapping of the 5' and 3' extremities of the major WDV encoded transcripts has allowed us to identify putative transcription regulatory sequences and the presence of multiple overlapping transcripts may suggest temporal regulation of transcription.

Amino Acid Sequence↗

Primary human T lymphocytes engineered with a codon-optimized IL-15 gene resist cytokine withdrawal-induced apoptosis and persist long-term in the absence of exogenous cytokine.

IL-15 is a common gamma-chain cytokine that has been shown to be more active than IL-2 in several murine cancer immunotherapy models. Although T lymphocytes do not produce IL-15, murine lymphocytes carrying an IL-15 transgene demonstrated superior antitumor activity in the immunotherapy of B16 melanoma. Thus, we sought to investigate the biological impact of constitutive IL-15 expression by human lymphocytes. In this report we describe the generation of a retroviral vector encoding a codon-optimized IL-15 gene. Alternate codon usage significantly enhanced the translational efficiency of this tightly regulated gene in retroviral vector-transduced cells. Activated human CD4+ and CD8+ human lymphocytes expressed IL-15Ralpha and produced high levels of cytokine upon retroviral transduction with the IL-15 vector. IL-15-transduced lymphocytes remained viable for up to 180 days in the absence of exogenous cytokine. IL-15 vector-transduced T cells showed continued proliferation after cytokine withdrawal and resistance to apoptosis while retaining specific Ag recognition. In the setting of adoptive cell transfer, IL-15-transduced lymphocytes may prolong lymphocyte survival in vivo and could potentially enhance antitumor activity.

3T3 Cells↗

Codon-modifications and an endoplasmic reticulum-targeting sequence additively enhance expression of an Aspergillus phytase gene in transgenic canola.

Transgenic plants offer advantages for biomolecule production because plants can be grown on a large scale and the recombinant macromolecules can be easily harvested and extracted. We introduced an Aspergillus phytase gene into canola (Brassica napus) (line 9412 with low erucic acid and low glucosinolates) by Agrobacterium-mediated transformation. Phytase expression in transgenic plant was enhanced with a synthetic phytase gene according to the Brassica codon usage and an endoplasmic reticulum (ER) retention signal KDEL that confers an ER accumulation of the recombinant phytase. Secretion of the phytase to the extracellular fluid was also established by the use of the tobacco PR-S signal peptide. Phytase accumulation in mature seed accounted for 2.6% of the total soluble proteins. The enzyme can be glycosylated in the seeds of transgenic plants and retain a high stability during storage. These results suggest a commercial feasibility of producing a stable recombinant phytase in canola at a high level for animal feed supplement and for reducing phosphorus eutrophication problems.

6-Phytase↗

Cloning, sequencing, and expression of the mig gene of Mycobacterium avium, which codes for a secreted macrophage-induced protein.

Mycobacterium avium is an intracellular pathogen that has evolved to be a frequent cause of disseminated infection in immunocompromised patients. Although these bacilli are readily phagocytized, they are able to survive and even multiply within human macrophages. The process whereby mycobacteria circumvent the lytic functions of the macrophages is currently not well understood, but this is a key aspect in the pathogenicity of all pathogenic mycobacteria. Previously, we identified a gene in M. avium, designated mig (for macrophage-induced gene), the expression of which is induced when the bacilli grow in human macrophages (G. Plum and J. E. Clark-Curtiss, Infect. Immun. 62:476-483, 1994). In the present study we show that (i) the nucleotide sequence of the mig gene has an open reading frame of 295 amino acids with a strong bias for mycobacterial codon usage, (ii) the mig gene also codes for a putative signal peptide of 19 amino acid residues, (iii) mig is induced by acidity to be expressed as an early-secreted 30-kDa protein, and (iv) the Mig protein exhibits an AMP-binding domain signature. However, beyond this motif which is common to enzymes that activate a large variety of substrates, no homologies to known sequences are found. We also show that (v) Mycobacterium smegmatis strains expressing the Mig protein have a limited advantage for survival in macrophages. These findings may be concordant with a role of the mig gene in the virulence of M. avium.

Amino Acid Sequence↗

Overproduction and localization of components of the polyketide synthase of Streptomyces glaucescens involved in the production of the antibiotic tetracenomycin C.

Three proteins, including the beta-keto acyl synthase and the acyl carrier protein, involved in the synthesis of the polyketide antibiotic tetracenomycin C by Streptomyces glaucescens GLA.0 were produced in Escherichia coli by using the T7 RNA polymerase-dependent pT7-7 expression vector. Changing the N-terminal codon usage of two of the genes greatly increased the level of protein produced without affecting mRNA levels, suggesting improvements in translational efficiency. Western immunoblot analysis of cytoplasmic and membrane fractions of S. glaucescens with antibodies raised to synthetic oligopeptides corresponding to the two presumed components of the beta-keto acyl synthase indicated that both proteins were membrane bound; one appears to be proteolytically cleaved before or during association with the membrane. The beta-keto acyl synthase could be detected in stationary-phase cultures but not in rapidly growing cultures, correlating with the time of appearance of tetracenomycin C in the medium.

3-Oxoacyl-(Acyl-Carrier-Protein) Synthase↗

The putative acetyl-CoA synthetase gene of Cryptosporidium parvum and a new conserved protein motif in acetyl-CoA synthetases.

We determined the nucleotide (nt) sequence of the putative gene encoding acetyl-coenzyme A synthetase (ACS) from the parasitic protozoan Cryptosporidium parvum. The gene is single copy, located on a chromosome of approximately 1.08 mb, and has no introns. The gene is characterized by low codon usage bias and encodes a 694-amino acid (aa) protein with a predicted molecular size of 78 kDa, similar to other ACSs from different prokaryotic and eukaryotic species. Comparison of multiple protein alignments of ACSs revealed a new conserved sequence motif PKT(R/V/L)SGK(I/V/T)(T/M/V/K)R(R/N) near the C-terminus, which may be a signature for ACSs. This motif shares significant homology with sequences from other members of the AMP-binding family, has secondary structure similar to the purine-binding motif of ATP- and GTP-ases, and may play a role in the enzymatic activity of proteins from the AMP-binding family.

Acetate-CoA Ligase↗

The phosphofructokinase genes of yeast evolved from two duplication events.

Yeast phosphofructokinase (PFK) is an octameric enzyme composed of four alpha-subunits and four beta-subunits, encoded by the genes PFK1 and PFK2, respectively. PFK1 was mapped 23 cM distal to ADE3 on chromosome VII, and PFK2 30 cM proximal to RNA1 on chromosome XIII. The entire nucleotide sequences for the two genes were obtained by sequencing both DNA strands. Only one major open reading frame was found for each gene. They encode 987 aa for PFK1 (Mr 107,984) and 959 aa for PFK2 (Mr 104,589). Both genes show a biased codon usage. The deduced amino acid sequences showed: (i) 20% homology between the N- and the C-terminal halves of each subunit, (ii) 55% homology between the two subunits, and (iii) significant homologies to the PFK sequences from human and rabbit muscle (42%), Escherichia coli (34%), and Bacillus (36%). These data support the view that two gene duplication events occurred in the evolution of the yeast PFK genes. The first duplication event took place soon after the separation of prokaryotic and eukaryotic lineage and the second in Saccharomyces later in the phylogeny. Functional domains in the yeast subunits were deduced by comparison to the rabbit muscle enzyme.

Amino Acid Sequence↗

Nucleotide sequence of an Escherichia coli chromosomal hemolysin.

We determined the DNA sequence of an 8,211-base-pair region encompassing the chromosomal hemolysin, molecularly cloned from an O4 serotype strain of Escherichia coli. All four hemolysin cistrons (transcriptional order, C, A, B, and D) were encoded on the same DNA strand, and their predicted molecular masses were, respectively, 19.7, 109.8, 79.9, and 54.6 kilodaltons. The identification of pSF4000-encoded polypeptides in E. coli minicells corroborated the assignment of the predicted polypeptides for hlyC, hlyA, and hlyD. However, based on the minicell results, two polypeptides appeared to be encoded on the hlyB region, one similar in size to the predicted molecular mass of 79.9 kilodaltons, and the other a smaller 46-kilodalton polypeptide. The four hemolysin gene displayed similar codon usage, which is atypical for E. coli. This reflects the low guanine-plus-cytosine content (40.2%) of the hemolysin DNA sequence and suggests the non-E. coli origin of the hemolysin determinant. In vitro-derived deletions of the hemolysin recombinant plasmid pSF4000 indicated that a region between 433 and 301 base pairs upstream of the putative start of hlyC is necessary for hemolysin synthesis. Based on the DNA sequence, a stem-loop transcription terminator-like structure (a 16-base-pair stem followed by seven uridylates) in the mRNA was predicted distal to the C-terminal end of hlyA. A model for the general transcriptional organization of the E. coli hemolysin determinant is presented.

Amino Acid Sequence↗