Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

The GC-rich transposon Bytmar1 from the deep-sea hydrothermal crab, Bythograea thermydron, may encode three transposase isoforms from a single ORF.

Mariner-like elements (MLEs) are classII transposons with highly conserved sequence properties and are widespread in the genome of animal species living in continental environments. We describe here the first full-length MLE found in the genome of a marine crustacean species, the deep-sea hydrothermal crab Bythograea thermydron (Crustacea), named Bytmar1. A comparison of its sequence features with those of the MLEs contained in the genomes of continental species reveals several distinctive characteristics. First, Bytmar1 elements contains an ORF that may encode three transposase isoforms 349, 379, and 398 amino acids (aa) in long. The two biggest proteins are due to the presence of a 30- and 49-aa flag, respectively, at the N-terminal end of the 349-aa cardinal MLE transposase. Their GC contents are also significantly higher than those found in continental MLEs. This feature is mainly due to codon usage in the transposase ORF and directly interferes with the curvature propensities of the Bytmar1 nucleic acid sequence. Such an elevated GC content may interfere with the ability of Bytmar 1 to form an excision complex and, in consequence, with its efficiency to transpose. Finally, the origin of these characteristics and their possible consequences on transposition efficiency are discussed.

Amino Acid Sequence↗

The DNA sequence and genetic organization of a Neurospora mitochondrial plasmid suggest a relationship to introns and mobile elements.

We have determined the complete 3581 bp sequence of the mitochondrial plasmid from Neurospora crassa strain Mauriceville-1c. The plasmid contains a long open reading frame that is expressed in its major transcript and could encode a hydrophilic protein of 710 amino acids. Two characteristics of the plasmid--codon usage and the presence of conserved sequence elements--suggest that it is related to Group I mtDNA introns. The major transcripts of the plasmid are approximately full-length, colinear RNAs that have heterogenous 5' ends and a single major 3' end. The major 5' and 3' ends are adjacent and slightly overlapping. The Mauriceville plasmid may belong to a class of genetic elements that were or are the progenitors of mtDNA introns.

Base Sequence↗

Comparative analysis of the cDNA sequences derived from the larval and the adult alpha 1-globin mRNAs of Xenopus laevis.

The complete nucleotide sequences of cloned cDNA segments derived from the larval and the adult alpha 1-globin mRNA of Xenopus laevis have been determined. These sequences comprise part of the 5' noncoding region, the entire coding region and the 3' noncoding region, including the polyadenylation site. The larval sequence differs from the adult one by a much longer 3' noncoding region. The sequences diverge by 47%, but codon usage is similar. Comparison of the amino acid sequences of vertebrates shows that the sites of heme contact are highly conserved, whereas the alpha 1/beta 1- and the alpha 1/beta 11-interfaces diverge to different degrees. In these regions the larval alpha 1-globin diverges less from embryonic than from adult alpha-like globins of vertebrates. This suggests that these sites are mainly responsible for the functional peculiarities of the larval amphibian hemoglobins.

Animals↗

The mitochondrial genome of the smaller tea tortrix Adoxophyes honmai (Lepidoptera: Tortricidae).

Almost complete mitochondrial genome from a representative of an insect order Lepidoptera, the smaller tea tortrix Adoxophyes honmai was determined. The 15,680 bp long A. honmai genome encodes 13 putative proteins, two ribosomal RNAs and 22 tRNAs. The nucleotide sequences of A. honmai mitochondrial DNA have been compared with those of five species from the Lepidoptera and insects in the other orders that are available in the databases. The orientation and gene order of A. honmai is almost same to that of other insects with a few minor exceptions in the order of tRNAs and distribution of non-coding regions. Nucleotide composition, amino acid composition and codon usage are in the range of values estimated from other insect mitogenomes. In AT rich region of A. honmai, tandem reiterations are observed with repeats of TAA. In a preliminary phylogenetic analysis based on the concatenated 7 protein coding genes, A. honmai, an apoditrysian tortricid moth joined basally within the monophyly of Lepidoptera, supporting its relationship with other more derived species including obtectomeran Ostrinia species.

Animals↗

Synonymous substitutions in the Xdh gene of Drosophila: heterogeneous distribution along the coding region.

The Xdh (rosy) region of Drosophila subobscura has been sequenced and compared to the homologous region of D. pseudoobscura and D. melanogaster. Estimates of the numbers of synonymous substitutions per site (Ks) confirm that Xdh has a high synonymous substitution rate. The distributions of both nonsynonymous and synonymous substitutions along the coding region were found to be heterogeneous. Also, no relationship has been detected between Ks estimates and codon usage bias along the gene, in contrast with the generally observed relationship among genes. This heterogeneous distribution of synonymous substitutions along the Xdh gene, which is expression-level independent, could be explained by a differential selection pressure on synonymous sites along the coding region acting on mRNA secondary structure. The synonymous rate in the Xdh coding region is lower in the D. subobscura than in the D. pseudoobscura lineage, whereas the reverse is true for the Adh gene.

Animals↗

Sequence comparison of ACE-1, the gene encoding acetylcholinesterase of class A, in the two nematodes Caenorhabditis elegans and Caenorhabditis briggsae.

The ace-1 gene, which encodes acetylcholinesterase of class A, has been cloned and sequenced in C. briggsae and compared to its homologue in C. elegans. Both genes present an open reading frame of 1860 nucleotides. The percentages of identity are 80% and 95% at the nucleotide and aminoacid levels respectively. All residues characteristic of an acetylcholinesterase are found in conserved positions in C. briggsae ACE-1. The deduced C-terminus is hydrophilic, thus resembling the catalytic peptide T of vertebrate cholinesterases. Codon usage in both ace-1 genes appears to be lowly biased. This may indicate that these genes are lowly expressed. The splicing sites of the eight introns of ace-1 in C. elegans are conserved in C. briggsae, but introns are shorter in C. briggsae. No homology was found between intronic sequences in both species, except for the consensus border sequences.

Acetylcholinesterase↗

The human and mouse GATA-6 genes utilize two promoters and two initiation codons.

GATA-6 has been implicated in the regulation of myocardial differentiation during cardiogenesis. To determine how its expression is controlled, we have characterized the human and mouse genes. We have mapped their transcriptional start sites and demonstrate that two alternative promoters and 5' noncoding exons are utilized. Both transcript isoforms are expressed in the same tissue-specific and developmental stage-specific pattern, and their ratio appears similar wherever examined. The more upstream noncoding exon showed a substantial degree of homology between the two mammalian species, suggesting a conserved regulatory function. Moreover, in transfection assays we show that elements within this exon act to promote its transcription. Positive regulatory elements that effect transcription from the more downstream exon were not apparent in this assay, revealing a regulatory distinction between the two promoters. We also demonstrate alternative initiator codon usage in both the human and mouse GATA-6 genes. Both isoforms of the protein are synthesized in vitro regardless of which 5' noncoding exon is present in the RNA, although the larger protein has greater transcriptional activation potential in transfection assays. Thus, GATA-6 function in the cell is controlled by a complex interplay of transcriptional and translational regulation.

Alternative Splicing↗

Optimizing doped libraries by using genetic algorithms.

The insertion of random sequences into protein-encoding genes in combination with biological selection techniques has become a valuable tool in the design of molecules that have useful and possibly novel properties. By employing highly effective screening protocols, a functional and unique structure that had not been anticipated can be distinguished among a huge collection of inactive molecules that together represent all possible amino acid combinations. This technique is severely limited by its restriction to a library of manageable size. One approach for limiting the size of a mutant library relies on 'doping schemes', where subsets of amino acids are generated that reveal only certain combinations of amino acids in a protein sequence. Three mononucleotide mixtures for each codon concerned must be designed, such that the resulting codons that are assembled during chemical gene synthesis represent the desired amino acid mixture on the level of the translated protein. In this paper we present a doping algorithm that "reverse translates' a desired mixture of certain amino acids into three mixtures of mononucleotides. The algorithm is designed to optimally bias these mixtures towards the codons of choice. This approach combines a genetic algorithm with local optimization strategies based on the downhill simplex method. Disparate relative representations of all amino acids (and stop codons) within a target set can be generated. Optional weighing factors are employed to emphasize the frequencies of certain amino acids and their codon usage, and to compensate for reaction rates of different mononucleotide building blocks (synthons) during chemical DNA synthesis. The effect of statistical errors that accompany an experimental realization of calculated nucleotide mixtures on the generated mixtures of amino acids is simulated. These simulations show that the robustness of different optima with respect to small deviations from calculated values depends on their concomitant fitness. Furthermore, the calculations probe the fitness landscape locally and allow a preliminary assessment of its structure.

Algorithms↗

Gene electro-transfer of an improved erythropoietin plasmid in mice and non-human primates.

BACKGROUND: Anemia due to impaired erythropoietin (EPO) production is associated with kidney failure. Recombinant proteins are commonly administered to alleviate the symptoms of this dysfunction, whereas gene therapy approaches envisaging the delivery of EPO genes have been tried in animal models in order to achieve stable and long-lasting EPO protein production. Naked DNA intramuscular injection is a safe approach for gene delivery; however, transduction levels show high inter-individual variability in rodents and very poor efficiency in non-human primates. Transduction can be improved in several animal models by application of electric pulses after DNA injection. METHODS: We have designed a modified EPO gene version by changing the EPO leader sequence and optimizing the gene codon usage. This modified gene was electro-injected into mice, rabbits and cynomolgus monkeys to test for protein production and biological effect. CONCLUSIONS: The modified EPO gene yields higher levels of circulating transgene product and a more significant biological effect than the wild-type gene in all the species tested, thus showing great potential in clinically developable gene therapy approaches for EPO delivery.

Animals↗

Evolution and degeneration of eukaryotic DNA replication system.

Several molecular forms of DNA polymerases have been identified in eukaryotic cells. Although three DNA polymerases alpha, delta, and epsilon, have been well studied and indicated to be involved in nuclear DNA replication process, it remains unclear how this hetero-polymerase system might have arisen. Here I wish to consider its past and future, viewed in the context of molecular evolution. Comparative analysis has revealed some nucleotides and/or amino acids to be conserved in DNA polymerase delta, in polymerase domains III and IV, which have disappeared in DNA polymerase alpha. Furthermore, the codon usage for serine residues in conserved domains of DNA polymerase alpha varies and is not as conservative as for DNA polymerase delta. Recently and in the present study, I have reported that DNA polymerase delta could substitute for the function of DNA polymerase alpha in vitro, and proposed the hypothesis that eukaryotic DNA polymerase alpha arose due to symbiotic contacts. This 'exogenous' polymerase would be expected to be excluded from the eukaryotic DNA replication system, and my analysis in the present study suggests it is about to degenerate.

Codon↗

G+C3 structuring along the genome: a common feature in prokaryotes.

The heterogeneity of gene nucleotide content in prokaryotic genomes is commonly interpreted as the result of three main phenomena: (1) genes undergo different selection pressures both during and after translation (affecting codon and amino acid choice); (2) genes undergo different mutational pressure whether they are on the leading or lagging strand; and (3) genes may have different phylogenetic origins as a result of lateral transfers. However, this view neglects the necessity of organizing genetic information on a chromosome that needs to be replicated and folded, which may add constraints to single gene evolution. As a consequence, genes are potentially subjected to different mutation and selection pressures, depending on their position in the genome. In this paper, we analyze the structuring of different codon usage measures along completely sequenced bacterial genomes. We show that most of them are highly structured, suggesting that genes have different base content, depending on their location on the chromosome. A peculiar pattern of genome structure, with a tendency toward an A+T-enrichment near the replication terminus, is found in most bacterial phyla and may reflect common chromosome constraints. Several species may have lost this pattern, probably because of genome rearrangements or integration of foreign DNA. We show that in several species, this enrichment is associated with an increase of evolutionary rate and we discuss the evolutionary implications of these results. We argue that structural constraints acting on the circular chromosome are not negligible and that this natural structuring of bacterial genomes may be a cause of overestimation in lateral gene transfer predictions using codon composition indices.

Base Composition↗

Proteome analysis of the plant pathogen Xylella fastidiosa reveals major cellular and extracellular proteins and a peculiar codon bias distribution.

The bacteria Xylella fastidiosa is the causative agent of a number of economically important crop diseases, including citrus variegated chlorosis. Although its complete genome is already sequenced, X. fastidiosa is very poorly characterized by biochemical approaches at the protein level. In an initial effort to characterize protein expression in X. fastidiosa we used one- and two-dimensional gel electrophoresis and mass spectrometry to identify the products of 142 genes present in a whole cell extract and in an extracellular fraction of the citrus isolated strain 9a5c. Of particular interest for the study of pathogenesis are adhesion and secreted proteins. Homologs to proteins from three different adhesion systems (type IV fimbriae, mrk pili and hsf surface fibrils) were found to be coexpressed, the last two being detected only as multimeric complexes in the high molecular weight region of one-dimensional electrophoresis gels. Using a procedure to extract secreted proteins as well as proteins weakly attached to the cell surface we identified 30 different proteins including toxins, adhesion related proteins, antioxidant enzymes, different types of proteases and 16 hypothetical proteins. These data suggest that the intercellular space of X. fastidiosa colonies is a multifunctional microenvironment containing proteins related to in vivo bacterial survival and pathogenesis. A codon usage analysis of the most expressed proteins from the whole cell extract revealed a low biased distribution, which we propose is related to the slow growing nature of X. fastidiosa. A database of the X. fastidiosa proteome was developed and can be accessed via the internet (URL: www.proteome.ibi.unicamp.br).

Antioxidants↗

'Stop-codon-specific' restriction endonucleases: their use in mapping and gene manipulation.

Certain restriction endonucleases recognise target sequences that contain the stop triplet TAG and are commonly either 4 or 6 bp in length. Interestingly, these restriction targets do not occur at the frequency expected on the basis of base composition and size. For example, the tetranucleotide MaeI recognition sequence (CTAG) occurs considerably less commonly (5-8-fold) in the genome of Escherichia coli (and many other eubacteria) than expected from mononucleotide frequencies. This surprising rarity is particularly evident in protein-encoding genes and is largely dictated by codon usage. Thus, amber (TAG) nonsense mutations frequently give rise to novel MaeI (CTAG) sites which are unique within a translated region. Such amber/MaeI sites, whether arising spontaneously or created in vitro by site-directed mutagenesis, act as a useful physical marker for the presence of the nonsense mutation and are a convenient startpoint for a range of diverse procedures. These features provide a useful supplement to protein engineering methods which use nonsense suppression to mediate amino acid replacements.

Bacteria↗

Reprogramming neuroblastoma by diet-enhanced polyamine depletion.

Neuroblastoma is a highly lethal childhood tumour derived from differentiation-arrested neural crest cells1,2. Like all cancers, its growth is fuelled by metabolites obtained from either circulation or local biosynthesis3,4. Neuroblastomas depend on local polyamine biosynthesis, and the inhibitor difluoromethylornithine has shown clinical activity5. Here we show that such inhibition can be augmented by dietary restriction of upstream amino acid substrates, leading to disruption of oncogenic protein translation, tumour differentiation and profound survival gains in the Th-MYCN mouse model. Specifically, an arginine- and proline-free diet decreases the amount of the polyamine precursor ornithine and enhances tumour polyamine depletion by difluoromethylornithine. This polyamine depletion causes ribosome stalling, unexpectedly specifically at codons with adenosine in the third position. Such codons are selectively enriched in cell cycle genes and low in neuronal differentiation genes. Thus, impaired translation of these codons, induced by combined dietary and pharmacological intervention, favours a pro-differentiation proteome. These results suggest that the genes of specific cellular programmes have evolved hallmark codon usage preferences that enable coherent translational rewiring in response to metabolic stresses, and that this process can be targeted to activate differentiation of paediatric cancers.

Animals↗

A novel JC virus variant found in the Highlands of Papua New Guinea has a 21-base pair deletion in the agnoprotein gene.

OBJECTIVES: This paper describes a unique JC virus (JCV) variant recovered from the Highlands of Papua New Guinea that contains an inframe 21-bp deletion in the agnoprotein gene. We characterize the mutation and suggest possible roles for the deletion in JCV evolution. STUDY DESIGN/METHODS: JCV DNA was extracted from urine and polymerase chain reaction (PCR) amplified using whole genome primers. PCR products were cloned, and multiple clones were sequenced. The JCV agnogene was PCR amplified to verify the presence of the agnogene deletion. RESULTS: This mutation creates a 21-bp deletion near the 3' end, which alters the predicted secondary structure of the messenger RNA and changes local codon usage at the 3' end of the agnogene. Protein secondary structure predictions suggest the deleted portion of the agnoprotein may be a flexible surface feature. CONCLUSIONS: We describe the first stable coding region deletion in JCV that presumably signifies a single evolutionary event that led to the split from other Highlands viral groups and occurred well after the human expansions that led to the peopling of the Southwest Pacific.

Base Sequence↗

Scrambled duplications in the feline leukemia virus gag gene: a putative pattern for molecular evolution.

The present study is a detailed computer-assisted analysis of the feline leukemia virus gag gene nucleotide sequence together with its flanking sequences (ST-FeLV GAG) that is compared with the aligned sectors of the Moloney strain of murine leukemia virus (Mo-MuLV GAG) and of three strains of feline sarcoma virus. It shows that perfectly matched repeated oligomers up to 13 nucleotides long are overrepresented and scattered throughout both ST-FeLV GAG and Mo-MuLV GAG, in noncoding and coding sectors, with no stringent correlation to codon usage in ST-FeLV gPr80gag. Many repeated oligomers share a core consensus that is intriguingly part of the inverted repeat at the termini of the long terminal repeat. Local scrambled repetitions of nucleotide subsequences have been found; they suggest a model of molecular evolution by slippage-like mechanisms. Thus, viral genomes could be subject to the same evolutionary mechanisms that are now known to be operating extensively in eukaryotic genomes. The data are discussed in light of putative patterns of molecular evolution.

Antigens, Viral↗

Gene organization deduced from the complete sequence of liverwort Marchantia polymorpha mitochondrial DNA. A primitive form of plant mitochondrial genome.

Analysis of the mitochondrial DNA of a liverwort Marchantia polymorpha by electron microscopy and restriction endonuclease mapping indicated that the liverwort mitochondrial genome was a single circular molecule of about 184,400 base-pairs. We have determined the complete sequence of the liverwort mitochondrial DNA and detected 94 possible genes in the sequence of 186,608 base-pairs. These included genes for three species of ribosomal RNA, 29 genes for 27 species of transfer RNA and 30 open reading frames (ORFs) for functionally known proteins (16 ribosomal proteins, 3 subunits of H(+)-ATPase, 3 subunits of cytochrome c oxidase, apocytochrome b protein and 7 subunits of NADH ubiquinone oxidoreductase). Three ORFs showed similarity to ORFs of unknown function in the mitochondrial genomes of other organisms. Furthermore, 29 ORFs were predicted as possible genes by using the index of G + C content in first, second and third letters of codons (42.0 +/- 10.9%, 37.0 +/- 13.2% and 26.4 +/- 9.4%, respectively) obtained from the codon usages of identified liverwort genes. To date, 32 introns belonging to either group I or group II intron have been found in the coding regions of 17 genes including ribosomal RNA genes (rrn18 and rrn26), a transfer RNA gene (trnS) and a pseudogene (psi nad7). RNA editing was apparently lacking in liverwort mitochondria since the nucleotide sequences of the liverwort mitochondrial DNA were well-conserved at the DNA level.

Base Sequence↗

High-level expression of an altered cDNA encoding human isovaleryl-CoA dehydrogenase in Escherichia coli.

Isovaleryl-CoA dehydrogenase (IVD) catalyzes the conversion of isovaleryl-CoA to 3-methylcrotonyl-CoA in the leucine catabolism pathway. The cDNA encoding the mature human IVD polypeptide was cloned in a prokaryotic expression vector, but the level of expression in Escherichia coli was extremely low and attempts to purify the enzyme to homogeneity were unsuccessful. To enhance expression, the nucleotide sequence of 22 codons within the 111-bp region at the 5'-end of the cDNA was altered to accommodate E. coli codon usage without altering the amino-acid coding sequence. The altered IVD cDNA was synthesized by PCR, using a primer containing the desired modifications. Following overnight induction of the E. coli transformed with this cDNA, the enzyme was purified to homogeneity using diethylaminoethyl agarose and high-pressure ceramic hydroxyapatite resins. IVD activity was increased 165-fold in the crude extract of cells containing the modified cDNA, as compared to that containing the wild-type cDNA.

Amino Acid Sequence↗