Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

Composition strand asymmetries in prokaryotic genomes: mutational bias and biased gene orientation.

Most prokaryotic genomes display strand compositional asymmetries, but the reasons for these biases remain unclear. When the distribution of gene orientation is biased, as it often is, this may induce a bias in composition, as codon frequencies are not identical. We show here that this effect can be estimated and removed, and that the residual base skews are the highest at third base codon positions and lower at first and second positions. This strongly suggests that compositional asymmetries result from 1) a replication-related mutational bias that is filtered through selective pressure and/or from 2) an uneven distribution of gene orientation. In most cases, the mutational bias alters the codon usage and amino acid frequencies of the leading and the lagging strand. However, these features are not ubiquitous amongst prokaryotes, and the biological reasons for them remain to be found.

Bacillus subtilis↗

Expression systems for use in actinomycetes and related organisms.

There have been significant advances in genetic and molecular approaches to understanding the physiology of organisms belonging to the genera Mycobacterium, Corynebacterium, Nocardia and Streptomyces. This review discusses recent advances in heterologous protein expression in members of the actinomycete group, including codon usage, post-translational modification and inducible gene expression.

Actinomycetales↗

Evolution of Shc functions from nematode to human.

The Shc protein family is characterized by the (CH2)-PTB-CH1-SH2 modularity. Its complexity increased during evolution from one locus in Drosophila (dShc), to at least three loci in mammals (shc, rai and sli). The three mammalian loci encode, because of alternative initiation codon usage and splicing pattern, at least six Shc-like proteins. Genetic and biological evidence indicates that the mammalian Shc isoforms regulate functions as diverse as growth (p52/p46Shc), apoptosis (p66Shc) and life-span (p66Shc). Available structure-function data and analysis of sequence similarities of Shc-like genes and proteins suggest complex diversification of Shc functions during evolution. Notably, Ras activation, the best-characterized Shc activity, appears to be a recent evolutionary acquisition.

Adaptor Proteins, Signal Transducing↗

Engineered GFP as a vital reporter in plants.

BACKGROUND: The green-fluorescent protein (GFP) of the jellyfish Aequorea victoria has recently been used as a universal reporter in a broad range of heterologous living cells and organisms. Although successful in some plant transient expression assays based on strong promoters or high copy number viral vectors, further improvement of expression efficiency and fluorescent intensity are required for GFP to be useful as a marker in intact plants. Here, we report that an extensively modified GFP is a versatile and sensitive reporter in a variety of living plant cells and in transgenic plants. RESULTS: We show that a re-engineered GFP gene sequence, with the favored codons of highly expressed human proteins, gives 20-fold higher GFP expression in maize leaf cells than the original jellyfish GFP sequence. When combined with a mutation in the chromophore, the replacement of the serine at position 65 with a threonine, the new GFP sequence gives more than 100-fold brighter fluorescent signals upon excitation with 490 nm (blue) light, and swifter chromophore formation. We also show that this modified GFP has a broad use in various transient expression systems, and allows the easy detection of weak promoter activity, visualization of protein targeting into the nucleus and various plastids, and analysis of signal transduction pathways in living single cells and in transgenic plants. CONCLUSIONS: The modified GFP is a simple and economical new tool for the direct visualization of promoter activities with a broad range of strength and cell specificity. It can be used to measure dynamic responses of signal transduction pathways, transfection efficiency, and subcellular localization of chimeric proteins, and should be suitable for many other applications in genetically modified living cells and tissues of higher plants. The data also suggest that the codon usage effect might be universal, allowing the design of recombinant proteins with high expression efficiency in evolutionarily distant species such as humans and maize.

Allium↗

Reporters for the analysis of gene regulation in fungi pathogenic to man.

In the past few years, highly sensitive gene reporters have been developed for the infectious fungi including gene reporters with altered codon usage. The tools are, therefore, now at hand for functionally characterizing the promoters of genes regulated by the bud-hypha transition, high frequency switching and cues from the cellular environment.

Candida albicans↗

Neutral effect of recombination on base composition in Drosophila.

Recombination is thought to have various evolutionary effects on genome evolution. In this study, we investigated the relationship between the base composition and recombination rate in the Drosophila melanogaster genome. Because of a current debate about the accuracy of the estimates of recombination rate in Drosophila, we used eight different measures of recombination rate from recent work. We confirmed that the G + C content of large introns and flanking regions is positively correlated with recombination rate, suggesting that recombination has a neutral effect on base composition in Drosophila. We also confirmed that this neutral effect of recombination is the main determinant of the correlation between synonymous codon usage bias and recombination rate in Drosophila.

Animals↗

Nucleotide composition in protein-coding and non-coding DNA in the zygomycete Phycomyces blakesleeanus.

The zygomycete Phycomyces blakesleeanus has a 30 Mb genome with a 35% content of guanine and cytosine (GC). We determined the GC content in Phycomyces genes and fragments of genes available in public databases, the frequency of nucleotides in each codon position, and the codon usage. We observed a difference of 18% between the GC content of protein-coding and non-coding DNA. This large difference allowed the visualization of protein-coding DNA by plotting the GC content along a segment of Phycomyces DNA. We have identified a high GC DNA segment linked to the pyrG genes of the zygomycete genera Phycomyces, Mucor, and Blakeslea that corresponds to the 3' end of the gene responsible for the protein kinase C.

Amino Acid Sequence↗

Porphobilinogen synthase from pea: expression from an artificial gene, kinetic characterization, and novel implications for subunit interactions.

Porphobilinogen synthase (PBGS) is present in all organisms that synthesize tetrapyrroles such as heme, chlorophyll, and vitamin B(12). The homooctameric metalloenzyme catalyzes the condensation of two 5-aminolevulinic acid molecules to form the tetrapyrrole precursor porphobilinogen. An artificial gene encoding PBGS of pea (Pisum sativum L.) was designed to overcome previous problems during bacterial expression caused by suboptimal codon usage and was constructed by recursive polymerase chain reaction from synthetic oligonucleotides. The recombinant 330 residue enzyme without a putative chloroplast transit peptide was expressed in Escherichia coli and purified in 100-mg quantities. The specific activity is protein concentration dependent, which indicates that a maximally active octamer can dissociate into less active smaller units. The enzyme is most active at slightly alkaline pH; it shows two pK(a) values of 7.4 and 9.7. Atomic absorption spectroscopy shows maximal binding of three Mg(II) per subunit; kinetic data support two functionally distinct types of Mg(II) and the third appears to be nonphysiologic and inhibitory. Analysis of the protein concentration dependence of the specific activity suggests that the minimal functional unit is a tetramer. A model of octameric pea PBGS was built to predict the location of intermolecular disulfide linkages that were revealed by nonreducing sodium dodecyl sulfate-polyacrylamide gel electrophoresis. As verified by site-specific mutagenesis, disulfide linkages can form between four cysteines per octamer, each located five amino acids from the C-terminus. These data are consistent with the protein undergoing conformational changes and the idea that whole-body motion can occur between subunits.

Amino Acid Sequence↗

High-level expression and mutagenesis of recombinant human phosphatidylcholine transfer protein using a synthetic gene: evidence for a C-terminal membrane binding domain.

Phosphatidylcholine transfer protein (PC-TP) is a 214-amino acid cytosolic protein that promotes intermembrane transfer of phosphatidylcholines, but no other phospholipid class. To probe mechanisms for membrane interactions and phosphatidylcholine binding, we expressed recombinant human PC-TP in Escherichia coli using a synthetic gene. Optimization of codon usage for bacterial protein translation increased expression of PC-TP from trace levels to >10% of the E. coli cytosolic protein mass. On the basis of secondary structure predictions of an amphipathic alpha-helix (residues 198-212) in proximity to a hydrophobic alpha-helix (residues 184-193), we explored whether the C-terminus might interact with membranes and promote binding of phosphatidylcholines. Consistent with this possibility, truncation of five residues from the C-terminus shortened the predicted amphipathic alpha-helix and decreased PC-TP activity by 50%, whereas removal of 10 residues eliminated the alpha-helix, abolished activity, and markedly decreased the level of membrane binding. Circular dichroic spectra of synthetic peptides containing one ((196-214)PC-TP) or both ((183-214)PC-TP) predicted C-terminal alpha-helices in aqueous buffer were most consistent with random coil structures. However, both peptides adopted alpha-helical configurations in the presence of trifluoroethanol or phosphatidylcholine/phosphatidylserine small unilamellar vesicles. The helical content of (196-214)PC-TP increased in proportion to vesicle phosphatidylserine content, consistent with stabilization of the alpha-helix at the membrane surface. In contrast, the helical content of (183-214)PC-TP was not influenced by vesicle composition, implying that the more hydrophobic of the alpha-helices penetrated into the membrane bilayer. These studies suggest that tandem alpha-helices located near the C-terminus of PC-TP facilitate membrane binding and extraction of phosphatidylcholines.

Amino Acid Sequence↗

Sequence determination and analysis of the 3' region of chicken pro-alpha 1(I) and pro-alpha 2(I) collagen messenger ribonucleic acids including the carboxy-terminal propeptide sequences.

Three pro-alpha 1 collagen cDNA clones, pCg1, pCg26, and pCg54, and two pro-alpha 2 collagen cDNA clones, pCg 13 and pCg45, were subjected to extensive DNA sequence determination. The combined sequences specified the amino acid sequences for chicken pro-alpha 1 and pro-alpha 2 type I collagens starting at residue 814 in the collagen triple-helical region and continuing to the procollagen C-termini as determined by the first in-phase termination codon. Thus, the sequences of 272 pro-alpha 1 C-terminal, 260 pro-alpha 2 C-terminal, 201 pro-alpha 1 helical, and 201 pro-alpha 2 helical amino acids were established. In addition, the sequences of several hundred nucleotides corresponding to noncoding regions of both procollagen mRNAs were determined. In total, 1589 pro-alpha 1 base pairs and 1691 pro-alpha 2 base pairs were sequenced, corresponding to approximately one-third of the total length of each mRNA. Both procollagen mRNA sequences have a high G+C content. The pro-alpha 1 mRNA is 75% G+C in the helical coding region sequenced and 61% G&C in the C-terminal coding region while the pro-alpha 2 mRNA is 60% and 48% G+C, respectively, in these regions. The dinucleotide sequence pCG occurs at a higher frequence in both sequences than is normally found in vertebrate DNAs and is approximately 5 times more frequent in the pro-alpha 1 sequence than in the pro-alpha 2 sequence. Nucleotide homology in the helical coding regions is very limited given that these sequences code for the repeating Gly-X-Y tripeptide in a region where X and Y residues are 50% conserved. These differences are clearly reflected in the preferred codon usages of the two mRNAs.

Amino Acid Sequence↗

Modifying the substrate specificity of staphylococcal lipases.

The lipase from Staphylococcus hyicus (SHL) displays a high phospholipase activity whereas the homologous S. aureus lipase (SAL) is not active or hardly active on phospholipid substrates. Previously, it has been shown that elements within the region comprising residues 254-358 are essential for the recognition of phospholipids by SHL. To specifically identify the important residues, nine small clusters of SHL were individually replaced by the corresponding SAL sequence within region 254-358. For cloning convenience, a synthetic gene fragment of SHL was assembled, thereby introducing restriction sites into the SHL gene and optimizing the codon usage. All nine chimeras were well-expressed as active enzymes. Eight chimeras showed lipase and phospholipase activities within a factor of 2 comparable to WT-SHL in standard activity assays. Exchange of the polar SHL region 293-300 by the more hydrophobic SAL region resulted in a 32-fold increased k(cat)/K(m) value for lipase activity and a concomitant 68-fold decrease in k(cat)/K(m) for phospholipase activity. Both changes are due to effects on catalytic turnover as well as on substrate affinity. Subsequently, six point mutants were generated; G293N, E295F, T297P, K298F, I299V, and L300I. Residue E295 appeared to play a minor role whereas K298 was the major determinant for phospholipase activity. The mutation K298F caused a 60-fold decrease in k(cat)/K(m) on the phospholipid substrate due to changes in both k(cat) and K(m). Substitution of F298 by a lysine in SAL resulted in a 4-fold increase in phospholipase activity. Two additional hydrophobic to polar substitutions further increased the phospholipase activity 23-fold compared to WT-SAL.

Amino Acid Sequence↗

Analysis of major ampullate silk cDNAs from two non-orb-weaving spiders.

Compared to other arthropods, spiders are unique in their use of silk throughout their life span and the extraordinary mechanical properties of the silk threads they produce. Studies on orb-weaving spider silk proteins have shown that silk proteins are composed of highly repetitive regions, characterized by alanine and glycine-rich units. We have isolated and sequenced four partial cDNA clones representing major ampullate spider silk gene transcripts from two non-orb weavers: three for Kukulcania hibernalis and one for Agelenopsis aperta. These cDNA sequences were compared to each other, as well as to the previously published orb-weaver silk gene sequences. The results indicate that the repeats encoding conserved amino acid motifs such as polyA and polyGA that are characteristic of some orb-weaving spider silks are also found in some of the cDNAs reported in this study. However, we also found other motifs such as polyGS and polyGV in the cDNA sequences from the two non-orb-weaving spiders. The amino acid composition of the silk gland extracts shows that alanine and glycine are the major components of the silk of these two non-orb weavers as is the case in orb-weaver silks. Sequence alignment shows that A. aperta's cDNA displays a C-terminal encoding region that is about 44% similar to the one present in N. clavipes's MaSp1 cDNA. In addition, as previously observed for spider silk sequences, the analysis of the codon usage for these four cDNAs demonstrates a bias for A or T in the wobble base position.

Amino Acid Sequence↗

Specific sequence modifications of a cry3B endotoxin gene result in high levels of expression and insect resistance.

Solanum melongena (eggplant) cv. Picentia and the wild species Solanum integrifolium were transformed with both a wild type (wt) and four mutagenized versions of Bacillus thuringiensis (Bt) gene Bt43 belonging to the cry3 class. The Bt gene was partly modified in its nucleotide sequence by replacing four target regions (W: +1 to +170; X: +592 to +1057; Y: +1203 to +1376; Z: +1376 to +1984) with synthetic fragments obtained by polymerase chain reaction amplification of crude oligonucleotides. The synthetic Bt genes were designed to avoid, in their modified regions, sequences such as ATTTA sequence, polyadenylation sequences and splicing sites, which might destabilize the messenger RNA. Furthermore, the codon usage was improved for a better expression in the plant system. The amino acid composition was not altered. Four versions of the modified Bt gene were obtained, BtE, BtF, BtH and BtI, with a nucleotide subtitution percentage of 8.2, 8.6, 14, and 16%, respectively, in comparison to the wt gene Bt43. Modified versions contained different subsets of substituted regions: BtE-W + Z, BtF - Y + Z, BtH-X + Y + Z, BtI - W + X + Y + Z. In the final modified version (BtI), overall guanine+cytosine was increased from the 34.1% of the wt gene to 45.5%, and most of the destabilizing sequences were eliminated. Transgenic plants obtained with the more modified versions, BtH and BtI, were fully resistant to Leptinotarsa decemlineata Say first- and third-instar larvae, while Bt43 wt, BtE and BtF genotypes did not cause mortality and did not affect larval development.

Animals↗

Complete sequence of the mitochondrial DNA of Chlamydomonas eugametos.

The complete nucleotide sequence of the Chlamydomonas eugametos (Chlamydomonadales, Chlorophyceae, sensu Mattox and Stewart) mitochondrial genome has been determined (22,897 bp, 34.6% G + C). The genes identified in this circular-mapping genome include those for apocytochrome b, subunit 1 of the cytochrome oxidase complex, subunits 1, 2, 4, 5, and 6 of the NADH dehydrogenase complex, discontinuous large and small subunit ribosomal rRNAs and three tRNAs whose anticodons CAU, CCA and UUG are specific for methionine, tryptophan and glutamine, respectively. The C. eugametos mitochondrial DNA (mtDNA), therefore, shares almost the same reduced set of coding functions and similar unusual features of rRNA gene organization with the linear 15.8 kb mtDNA of Chlamydomonas reinhardtii, the only other completely sequenced chlamydomonadalean mtDNA. However, sequence analysis of the C. eugametos mtDNA has revealed the following distinguishing features relative to those of C. reinhardtii: (1) the absence of a reverse transcriptase-like gene homologue, (2) the presence of an additional gene for tRNA(met) that may be a pseudogene, (3) a completely different gene order, (4) transcription of all genes from the same mtDNA strand, (5) a lower G + C content, (6) less pronounced bias in codon usage, and (7) nine group I introns, several of which contain open reading frames coding for potential maturases/endonucleases and two have a nucleotide at the 5' or 3' splice site of the deduced precursor RNAs that deviates from highly conserved nucleotides reported in other group I introns. The features of mitochondrial genome organization and gene content shared by C. eugametos and C. reinhardtii contrast with those of other green algal mtDNAs that have been characterized in detail. The deep evolutionary divergence between these two Chlamydomonas taxa within the Chlamydomonadales suggests that their shared features of mitochondrial genome organization evolved prior to the origin of this group.

Animals↗

Origin, adaptation and evolutionary pathways of fungal viruses.

Fungal viruses or mycoviruses are widespread in fungi and are believed to be of ancient origin. They have evolved in concert with their hosts and are usually associated with symptomless infections. Mycoviruses are transmitted intracellularly during cell division, sporogenesis and cell fusion, and they lack an extracellular phase to their life cycles. Their natural host ranges are limited to individuals within the same or closely related vegetative compatibility groups. Typically, fungal viruses are isometric particles 25-50 nm in diameter, and possess dsRNA genomes. The best characterized of these belong to the family Totiviridae whose members have simple undivided dsRNA genomes comprised of a coat protein (CP) gene and an RNA dependent RNA polymerase (RDRP) gene. A recently characterized totivirus infecting a filamentous fungus was found to be more closely related to protozoan totiviruses than to yeast totiviruses suggesting these viruses existed prior to the divergence of fungi and protozoa. Although the dsRNA viruses at large are polyphyletic, based on RDRP sequence comparisons, the totiviruses are monophyletic. The theory of a cellular self-replicating mRNA as the origin of totiviruses is attractive because of their apparent ancient origin, the close relationships among their RDRPs, genome simplicity and the ability to use host proteins efficiently. Mycoviruses with bipartite genomes (partitiviruses), like the totiviruses, have simple genomes, but the CP and RDRP genes are on separate dsRNA segments. Because of RDRP sequence similarity, the partitiviruses are probably derived from a totivirus ancestor. The mycoviruses with unencapsidated dsRNA-like genomes (hypoviruses) and those with bacilliform (+) strand RNA genomes (barnaviruses) have more complex genomes and appear to have common ancestry with plant (+) strand RNA viruses in supergroup 1 with potyvirus and sobemovirus lineages, respectively. The La France isometric virus (LIV), an unclassified virus with multipartite dsRNA genome, is associated with a severe die-back disease of the cultivated mushroom. LIV appears to be of recent origin since it differs from its host in codon usage.

Adaptation, Biological↗

Methods for the detection of non-random base substitution in virus genes: models of synonymous nucleotide substitution in picornavirus genes.

A substantial fraction of phylogenetic divergence between closely related RNA virus genes is generally accounted for by synonymous (non-amino acid changing) point mutation. Viral evolution may be a complicated phenomena, governed by many different processes. However in this study we ask whether there are any properties in the patterns of synonymous nucleotide substitutions in three different Picornavirus genes that permit the process of accumulation of synonymous point mutation in these genes to be distinguished from some of the simplest most basic evolutionary models. We conclude that while the observed patterns in the occurrence of synonymous point substitution are consistent with those predicted by a model in which base mutation is equi-probable along a gene, and the probability of synonymous substitution determined only by local codon usage, some patterns in the actual nucleotides exchanged remain to be explained.

Amino Acid Substitution↗

Molecular anatomy of Tupaia (tree shrew) adenovirus genome; evolution of viral genes and viral phylogeny.

Adenoviruses are globally spread and infect species in all five taxons of vertebrates. Outstanding attention is focused on adenoviruses because of their transformation potential, their possible usability as vectors in gene therapy and their applicability in studies dealing with, e.g. cell cycle control, DNA replication, transcription, splicing, virus-host interactions, apoptosis, and viral evolution. The accumulation of genetic data provides the basis for the increase of our knowledge about adenoviruses. The Tupaia adenovirus (TAV) infects members of the genus Tupaiidae that are frequently used as laboratory animals in behavior research dealing with questions about biological and molecular processes of stress in mammals, in neurobiological and physiological studies, and as model organisms for human hepatitis B and C virus infections. In the present study the TAV genome underwent an extensive analysis including determination of codon usage, CG depletion, gene content, gene arrangement, potential splice sites, and phylogeny. The TAV genome has a length of 33,501 bp with a G+C content of 49.96%. The genome termini show a strong CG depletion that could be due to methylation of these genome regions during the viral replication cycle. The analysis of the coding capacity of the complete TAV genome resulted in the identification of 109 open reading frames (ORFs), of which 38 were predicted to be real viral genes. TAV was classified within the genus Mastadenovirus characterized by typical gene content, arrangement, and homology values of 29 conserved ORFs. Phylogenetic trees show that TAV is part of a separate evolutionary lineage and no mastadenovirus species can be considered as the most related. In contrast to other mastadenoviruses a direct ancestor of TAV captured a DUT gene from its mammalian host, presumably controlling local dUTP levels during replication and enhance viral replication in non-dividing host tissues. Furthermore, TAV possesses a second DNA-binding protein gene, that is likely to play a role in the determination of the host range. In view of these data it is conceivable that TAV underwent evolutionary adaptations to its biological environment resulting in the formation of special genomic components that provided TAV with the ability to expand its host range during viral evolution.

Adenoviridae Infections↗

A ROS repressor-mediated binary regulation system for control of gene expression in transgenic plants.

We describe a novel binary system to control transgene expression in plants. The system is based on the prokaryotic repressor, ROS, from Agrobacterium tumefaciens, optimized for plant codon usage and for nuclear targeting (synROS). The ROS protein bound in vitro to double stranded DNA comprising the ROS operator sequence, as well as to single stranded ROS operator DNA sequences, in an orientation-independent manner. A synROS-GUS fusion protein was localized to the nucleus, whereas wtROS-GUS fusion remained in the cytoplasm. The ability of synROS to repress transgene expression was validated in transgenic Arabidopsis thaliana and Brassica napus. When expressed constitutively under the actin2 promoter, synROS repressed the expression of the reporter gene gusA linked to a modified CaMV35S promoter containing ROS operator sequences in the vicinity of the TATA box and downstream of the transcription initiation signal. Repression ranged from 32 to 87% in A. thaliana, and from 23 to 76% in B. napus. These results are discussed in relation to the potential application of synROS in controlling the expression of transgenes and endogenous genes in plants and other organisms.

Agrobacterium tumefaciens↗