Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Population, evolutionary and genomic consequences of interference selection.

Weakly selected mutations are most likely to be physically clustered across genomes and, when sufficiently linked, they alter each others' fixation probability, a process we call interference selection (IS). Here we study population genetics and evolutionary consequences of IS on the selected mutations themselves and on adjacent selectively neutral variation. We show that IS reduces levels of polymorphism and increases low-frequency variants and linkage disequilibrium, in both selected and adjacent neutral mutations. IS can account for several well-documented patterns of variation and composition in genomic regions with low rates of crossing over in Drosophila. IS cannot be described simply as a reduction in the efficacy of selection and effective population size in standard models of selection and drift. Rather, IS can be better understood with models that incorporate a constant "traffic" of competing alleles. Our simulations also allow us to make genome-wide predictions that are specific to IS. We show that IS will be more severe at sites in the center of a region containing weakly selected mutations than at sites located close to the edge of the region. Drosophila melanogaster genomic data strongly support this prediction, with genes without introns showing significantly reduced codon bias in the center of coding regions. As expected, if introns relieve IS, genes with centrally located introns do not show reduced codon bias in the center of the coding region. We also show that reasonably small differences in the length of intermediate "neutral" sequences embedded in a region under selection increase the effectiveness of selection on the adjacent selected sequences. Hence, the presence and length of sequences such as introns or intergenic regions can be a trait subject to selection in recombining genomes. In support of this prediction, intron presence is positively correlated with a gene's codon bias in D. melanogaster. Finally, the study of temporal dynamics of IS after a change of recombination rate shows that nonequilibrium codon usage may be the norm rather than the exception.

Animals↗

Biological nitrogen fixation: primary structure of the Klebsiella pneumoniae nifH and nifD genes.

A DNA fragment carrying the Klebsiella pneumoniae nifK, D, and H genes was isolated from the nif- strain UNF841 (Tn5::nifK) by molecular cloning into the Escherichia coli plasmid pBR325. The nucleotide sequences of both the nifH gene, which encodes the Fe protein of the nitrogenase enzyme complex, and 622 nucleotides of the nifD gene, which encodes the alpha-subunit of the Mo-Fe protein, were determined by direct DNA sequencing by both the chemical and chain termination methods. A comparison of the primary structure of the Klebsiella nifH gene and its product with that recently determined for the blue-green alga Anabaena demonstrates that the gene sequences are more divergent than the protein sequence data would suggest. This implies that despite the strong, presumably functional, constraints that act at the protein structure level, the nucleotide sequence of the gene and its mRNA are only restrained by the coding requirements, allowing substantial drift in codon usage.

Amino Acid Sequence↗

Internal correspondence analysis of codon and amino-acid usage in thermophilic bacteria.

Starting from two datasets of codon usage in coding sequences from mesophilic and thermophilic bacteria, we used internal correspondence analysis to study the variability of codon usage within and between species, and within and between amino acids. The first dataset included 18,958,458 codons from 58,482 coding sequences from completely sequenced genomes of 25 species, along with 6,793,581 dinucleotides from 21,876 intergenic spaces. The second dataset, with partially sequenced genomes, included 97,095,873 codons from 293 bacterial species. Results were consistent between the two datasets. The trend for the amino-acid composition of thermophilic proteins was found to be under the control of a pressure at the nucleic acid level, not a selection at the protein level. This effect was not present in intergenic spaces, ruling out a pressure at the DNA level. The pattern at the mRNA level was more complex than a simple purine enrichment of the sense strand of coding sequences. Outliers in the partial genome dataset introduced a note of caution about the interpretation of temperature as the direct determinant of the trend observed in thermophiles. The surprising lack of selection on the amino-acid content of thermophilic proteins suggests that the amino-acid repertoire was set up in a hot environment.

Bacteria↗

The complete nucleotide sequence and gene organization of the mitochondrial genome of the oriental mole cricket, Gryllotalpa orientalis (Orthoptera: Gryllotalpidae).

The complete nucleotide sequences of the mitochondrial genome (mitogenome) of the oriental mole cricket, Gryllotalpa orientalis (Orthoptera: Gryllotalpidae), were determined. The 15,521-bp-long G. orientalis mitogenome contains typical gene complement, base composition, and codon usage found in metazoan mitogenomes. The G. orientalis mitogenome contains the third lowest A+T content (70.5%) among the complete insects mt genome sequences. The initiation codon for the G. orientalis COI gene appears to be ATG, instead of the tetranucleotides, which have been postulated to act as initiation codon for Locusta migratoria and some lepidopteran COI genes. The initiation codon for ND2 appears to be GTG, which is rare, but has been designated as an initiator of Tricholepidion gertschi ND2. All anticodons of G. orientalis tRNAs were identical to Drosophila yakuba and L. migratoria. The tRNA(Ser)(AGN) could not form a stable stem loop structure in the DHU arm as shown in many other insect tRNA(Ser)(AGN). Phylogenetic analysis of nucleotide sequence information from all mt genes supported a monophyletic Diptera, a monophyletic Lepidoptera, a monophyletic Coleoptera, a monophyletic Mecopterida (Diptera+Lepidoptera), and a monophyletic Endopterygota (Diptera+Lepidoptera+Coleoptera), suggesting that the complete insect mitogenome sequence has a resolving power to the diversification events within Endopterygota. However, the relationships of ancient insect orders were unstable, indicating the limited use of mitogenome information at deeper phylogenetic depth.

Animals↗

Polynucleotide viral vaccines: codon optimisation and ubiquitin conjugation enhances prophylactic and therapeutic efficacy.

Papillomavirus infection is a major antecedent of anogenital malignancy. We have previously established that the L1 and L2 capsid genes of papillomavirus have suboptimal codon usage for expression in mammalian cells. We now show that the lack of immunogenicity of polynucleotide vaccines based on the L1 gene can be overcome with codon modified L1, which induces strong immune responses, including conformational virus neutralising antibody and delayed type hypersensitivity. Conjugation of a ubiquitin gene to a hybrid gene incorporating L1 and the E7 non-structural papillomavirus protein improved E7 specific CTL responses, and induced protection against an E7 expressing tumour, but induced little neutralising antibody. However, a mixture of ubiquitin conjugated and non-ubiquitin conjugated polynucleotides induced virus neutralising antibody and E7 specific CD8 T cells. An optimal combined prophylactic/therapeutic viral vaccine might therefore comprise ubiquitin conjugated and non-ubiquitinated genes, to induce prophylactic neutralising antibody and therapeutic cell mediated immune responses.

Animals↗

Development of a GFP reporter gene for Chlamydomonas reinhardtii chloroplast.

Reporter genes have been successfully used in chloroplasts of higher plants, and high levels of recombinant protein expression have been reported. Reporter genes have also been used in the chloroplast of Chlamydomonas reinhardtii, but in most cases the amounts of protein produced appeared to be very low. We hypothesized that the inability to achieve high levels of recombinant protein expression in the C. reinhardtii chloroplast was due to the codon bias seen in the C. reinhardtii chloroplast genome. To test this hypothesis, we synthesized a gene encoding green fluorescent protein (GFP) de novo, optimizing its codon usage to reflect that of major C. reinhardtii chloroplast-encoded proteins. We monitored the accumulation of GFP in C. reinhardtii chloroplasts transformed with the codon-optimized GFP cassette (GFPct), under the control of the C. reinhardtii rbcL 5'- and 3'-UTRs. We compared this expression with the accumulation of GFP in C. reinhardtii transformed with a non-optimized GFP cassette (GFPncb), also under the control of the rbcL 5'- and 3'-UTRs. We demonstrate that C. reinhardtii chloroplasts transformed with the GFPct cassette accumulate approximately 80-fold more GFP than GFPncb-transformed strains. We further demonstrate that expression from the GFPct cassette, under control of the rbcL 5'- and 3'-UTRs, is sufficiently robust to report differences in protein synthesis based on subtle changes in environmental conditions, showing the utility of the GFPct gene as a reporter of C. reinhardtii chloroplast gene expression.

Amino Acid Sequence↗

Distribution and evolution of sequence characteristics in the E. coli genome.

The mean (G + C) composition (51.0%) and standard deviation (+/- 3.8%) of published DNA sequences accounting for 10% of the E. coli genome is in excellent agreement with the principal overall distribution determined by high resolution melting. While differences in base and neighbor characteristics are small and uniform throughout all regions of the genome, it is found that the (G + C) content of sequences varies in segmented fashion within boundaries corresponding to coding (53% G + C) and noncoding (46% G + C) regions; with variances in the latter being six-fold greater than in coding regions. The variance in different regions shows a strong negative dependence on (G + C) content of the region, reflecting the condition that A-T and G-C base pairs are preferred neighbors of A-T and C-G pairs, respectively; with the bias increasing with decreasing (G + C) content. Neighbor analysis indicates the most extreme positive biases occur in AA, TT, GC and CG throughout all regions, but particularly in noncoding regions. Extraordinary numbers of oligomeric strings of (A)n, etc., are the further consequence of this bias. These and other characteristics point to the existence of inherent biases in neighbor frequencies levied during replication or repair, and which reflect, in turn, neighbor influences during mutation. The bias in codon usage noted by Grantham and others is seen here as due, in part, to the adaptation of coding sequences to this microenvironment through selection among synonymous codons so as to preserve inherent neighbor biases.

Base Composition↗

Cloning of bovine prolactin cDNA and evolutionary implications of its sequence.

Prolactin, growth hormone, and chorionic somatomammotropin (placental lactogen) constitute a set of related polypeptides believed to derive from a common evolutionary ancestor protein. We have cloned and sequenced DNA complementary to the mRNA coding for bovine prolactin. This cDNA contains 702 bases corresponding to 10 amino acids in the leader peptide, all 199 amino acids of the hormone, and 75 nucleotides in the 3' untranslated region of the mRNA. Nucleotide sequence analysis of this cDNA permitted the identification of 10 amino acids in the signal peptide, plus the correction or elucidation of amino acid assignments at 16 sites where aspartic and glutamic acids had not been distinguished from their amides by amino acid sequencing. Codon usage in bovine prolactin mRNA is nonrandom, but, similarly to rat and human prolactins, it does not exhibit the strong preference for G or C in codon third positions seen in bovine, rat, and human growth hormone mRNAs. The translational termination signal in bovine prolactin in UAA, also the same as in rat and human prolactins and differing from the UAG "stop" codon used in bovine, rat, and human growth hormones and human chorionic somatomammotropin. The amino acid and mRNA nucleotide sequences of bovine, rat, and human prolactins and growth hormones were compared by several techniques based on various theories of molecular evolution. The comparison of prolactin to growth hormone is consistent in all three species, suggesting that the genes for these two hormones diverged about 350 million years ago. However, comparisons among the three prolactins or among the three growth hormones to determine the times of evolutionary divergence of the three species generated values that were inconsistent with each other and with the fossil record. Analysis of these discrepancies suggests that the genes for prolactin and growth hormone may now be evolving by different mechanisms.

Amino Acid Sequence↗

Synonymous mutations in the human dopamine receptor D2 (DRD2) affect mRNA stability and synthesis of the receptor.

Although changes in nucleotide sequence affecting the composition and the structure of proteins are well known, functional changes resulting from nucleotide substitutions cannot always be inferred from simple analysis of DNA sequence. Because a strong synonymous codon usage bias in the human DRD2 gene, suggesting selection on synonymous positions, was revealed by the relative independence of the G+C content of the third codon positions from the isochoric G+C frequencies, we chose to investigate functional effects of the six known naturally occurring synonymous changes (C132T, G423A, T765C, C939T, C957T, and G1101A) in the human DRD2. We report here that some synonymous mutations in the human DRD2 have functional effects and suggest a novel genetic mechanism. 957T, rather than being 'silent', altered the predicted mRNA folding, led to a decrease in mRNA stability and translation, and dramatically changed dopamine-induced up-regulation of DRD2 expression. 1101A did not show an effect by itself but annulled the above effects of 957T in the compound clone 957T/1101A, demonstrating that combinations of synonymous mutations can have functional consequences drastically different from those of each isolated mutation. C957T was found to be in linkage disequilibrium in a European-American population with the -141C Ins/Del and TaqI 'A' variants, which have been reported to be associated with schizophrenia and alcoholism, respectively. These results call into question some assumptions made about synonymous variation in molecular population genetics and gene-mapping studies of diseases with complex inheritance, and indicate that synonymous variation can have effects of potential pathophysiological and pharmacogenetic importance.

Dopamine↗

Optimized FaeG expression and a thermolabile enterotoxin DNA adjuvant enhance priming of an intestinal immune response by an FaeG DNA vaccine in pigs.

One of the problems hindering the development of DNA vaccines is the relatively low immunogenicity often seen in humans and large animals compared to that in mice. In the present study, we tried to enhance the immunogenicity of a pcDNA1/faeG19 DNA vaccine in pigs by optimizing the FaeG expression plasmid and by coadministration of the plasmid vectors encoding the A and B subunits of the Escherichia coli thermolabile enterotoxin (LT). The insertion of a Kozak sequence and optimization of vector (cellular localization and expression) and both vector and codon usage were all shown to enhance in vitro FaeG expression compared to that of pcDNA1/faeG19. Subsequently, pcDNA1/faeG19 and the vector-optimized and the vector-codon-optimized construct were tested for their immunogenicity in pigs. In line with the in vitro results, antibody responses were better induced with increasing expression. The LT vectors additionally enhanced the antibody response, although not significantly, and were necessary to induce an F4-specific cellular response. These vectors were also added because LT has been described to direct the systemic response towards a mucosal immunoglobulin A (IgA) response in mice. Here, however, the intradermal FaeG DNA prime-oral F4 boost immunization resulted in a mainly systemic IgG response, with only a marginal but significant reduction in F4+ E. coli fecal excretion when the piglets were primed with pWRGFaeGopt and pWRGFaeGopt with the LT vectors.

Adhesins, Escherichia coli↗

Correlations between Shine-Dalgarno sequences and gene features such as predicted expression levels and operon structures.

This work assesses relationships for 30 complete prokaryotic genomes between the presence of the Shine-Dalgarno (SD) sequence and other gene features, including expression levels, type of start codon, and distance between successive genes. A significant positive correlation of the presence of an SD sequence and the predicted expression level of a gene based on codon usage biases was ascertained, such that predicted highly expressed genes are more likely to possess a strong SD sequence than average genes. Genes with AUG start codons are more likely than genes with other start codons, GUG or UUG, to possess an SD sequence. Genes in close proximity to upstream genes on the same coding strand in most genomes are significantly higher in SD presence. In light of these results, we discuss the role of the SD sequence in translation initiation and its relationship with predicted gene expression levels and with operon structure in both bacterial and archaeal genomes.

Archaeal Proteins↗

Construction and validation of the Rhodobacter sphaeroides 2.4.1 DNA microarray: transcriptome flexibility at diverse growth modes.

A high-density oligonucleotide DNA microarray, a genechip, representing the 4.6-Mb genome of the facultative phototrophic proteobacterium, Rhodobacter sphaeroides 2.4.1, was custom-designed and manufactured by Affymetrix, Santa Clara, Calif. The genechip contains probe sets for 4,292 open reading frames (ORFs), 47 rRNA and tRNA genes, and 394 intergenic regions. The probe set sequences were derived from the genome annotation generated by Oak Ridge National Laboratory after extensive revision, which was based primarily upon codon usage characteristic of this GC-rich bacterium. As a result of the revision, numerous missing ORFs were uncovered, nonexistent ORFs were deleted, and misidentified start codons were corrected. To evaluate R. sphaeroides transcriptome flexibility, expression profiles for three diverse growth modes--aerobic respiration, anaerobic respiration in the dark, and anaerobic photosynthesis--were generated. Expression levels of one-fifth to one-third of the R. sphaeroides ORFs were significantly different in cells under any two growth modes. Pathways involved in energy generation and redox balance maintenance under three growth modes were reconstructed. Expression patterns of genes involved in these pathways mirrored known functional changes, suggesting that massive changes in gene expression are the major means used by R. sphaeroides in adaptation to diverse conditions. Differential expression was observed for genes encoding putative new participants in these pathways (additional photosystem genes, duplicate NADH dehydrogenase, ATP synthases), whose functionality has yet to be investigated. The DNA microarray data correlated well with data derived from quantitative reverse transcription-PCR, as well as with data from the literature, thus validating the R. sphaeroides genechip as a powerful and reliable tool for studying unprecedented metabolic versatility of this bacterium.

Adaptation, Physiological↗

Detecting overlapping coding sequences in virus genomes.

BACKGROUND: Detecting new coding sequences (CDSs) in viral genomes can be difficult for several reasons. The typically compact genomes often contain a number of overlapping coding and non-coding functional elements, which can result in unusual patterns of codon usage; conservation between related sequences can be difficult to interpret--especially within overlapping genes; and viruses often employ non-canonical translational mechanisms--e.g. frameshifting, stop codon read-through, leaky-scanning and internal ribosome entry sites--which can conceal potentially coding open reading frames (ORFs). RESULTS: In a previous paper we introduced a new statistic--MLOGD (Maximum Likelihood Overlapping Gene Detector)--for detecting and analysing overlapping CDSs. Here we present (a) an improved MLOGD statistic, (b) a greatly extended suite of software using MLOGD, (c) a database of results for 640 virus sequence alignments, and (d) a web-interface to the software and database. Tests show that, from an alignment with just 20 mutations, MLOGD can discriminate non-overlapping CDSs from non-coding ORFs with a typical accuracy of up to 98%, and can detect CDSs overlapping known CDSs with a typical accuracy of 90%. In addition, the software produces a variety of statistics and graphics, useful for analysing an input multiple sequence alignment. CONCLUSION: MLOGD is an easy-to-use tool for virus genome annotation, detecting new CDSs--in particular overlapping or short CDSs--and for analysing overlapping CDSs following frameshift sites. The software, web-server, database and supplementary material are available at http://guinevere.otago.ac.nz/mlogd.html.

Algorithms↗

The role of selection in the evolution of human mitochondrial genomes.

High mutation rate in mammalian mitochondrial DNA generates a highly divergent pool of alleles even within species that have dispersed and expanded in size recently. Phylogenetic analysis of 277 human mitochondrial genomes revealed a significant (P < 0.01) excess of rRNA and nonsynonymous base substitutions among hotspots of recurrent mutation. Most hotspots involved transitions from guanine to adenine that, with thymine-to-cytosine transitions, illustrate the asymmetric bias in codon usage at synonymous sites on the heavy-strand DNA. The mitochondrion-encoded tRNAThr varied significantly more than any other tRNA gene. Threonine and valine codons were involved in 259 of the 414 amino acid replacements observed. The ratio of nonsynonymous changes from and to threonine and valine differed significantly (P = 0.003) between populations with neutral (22/58) and populations with significantly negative Tajima's D values (70/76), independent of their geographic location. In contrast to a recent suggestion that the excess of nonsilent mutations is characteristic of Arctic populations, implying their role in cold adaptation, we demonstrate that the surplus of nonsynonymous mutations is a general feature of the young branches of the phylogenetic tree, affecting also those that are found only in Africa. We introduce a new calibration method of the mutation rate of synonymous transitions to estimate the coalescent times of mtDNA haplogroups.

Amino Acid Substitution↗

Top DNA polymerase from Thermus thermophilus HB27: gene cloning, sequence determination, and physicochemical properties.

A gene, top encoding Thermus thermophilus HB27 (Top) DNA polymerase, was cloned in E. coli and its nucleotide sequence was determined. Based on its deduced amino acid sequence, Top DNA polymerase is a 93.8 kDa protein comprising 834 amino acid residues. Top DNA polymerase showed high amino acid homology with those of other DNA polymerases from the Thermus sp., for example, 87.3% identity with Taq DNA polymerase. Codon usage in the top gene was similar to those of the proteins from other Thermus strains. The G + C content in the third position of the codons was as high as 93%. The top gene under the control of the tac promoter was expressed in E. coli [plasmid pTOP9]. DNA amplification using the recombinant Top DNA polymerase performed the same as other thermostable DNA polymerases from Thermus strains. The optimum temperature for its reaction was 76 degrees C. An interesting observation was that the recombinant Top DNA polymerase was slowly cleaved into two fragments of about 60 kDa and 35 kDa at 4 degrees C and -20 degrees C. The larger fragment possessed polymerase activity like the Klenow fragment of E. coli DNA polymerase I. To prevent the cleavage of the Top DNA polymerase, a variety of protecting agents were examined. Among those examined, (NH4)2SO4 (100 mM) solution demonstrated an outstanding ability to block its cleavage for a prolonged period.

Amino Acid Sequence↗

Correlation analysis of amino acid usage in protein classes.

We present a comparative study of residue usage correlations of various organism protein sets of diverse phylogenetic species and of open reading frames of several large human viral genomes. Our correlation analysis reveals three major tendencies: (i) charge compensation reflected by the high correlation of basic with acidic residues; (ii) the positive correlations of functionally and structurally similar amino acids including many pairs of hydrophobic amino acids, all pairs of aromatic amino acids, the anionic pair (glutamate and aspartate), but not the cationic pair (lysine and arginine), moderately the hydroxyl pair (serine and threonine), the small amino acids (glycine and alanine), and many (but not all) of those having high values in the Dayhoff substitutability matrix (characteristics such as amino acid polarity or codon usage agreement, except for the wobble position, do not necessarily imply significant positive correlations); (iii) a widespread negative correlation of the aggregate strong codon group amino acids (Ala, Gly, Pro) versus the weak codon group amino acids (Lys, Ile, Tyr, Asn, Phe). Discussion and speculations relate amino acid usage correlations to protein function/structure, cellular localization, proximity in amino acid biosynthetic pathways, amino acid relative abundances, tRNA and aminoacyl synthetase availabilities, and evolutionary processes.

Amino Acids↗

Structural and putative regulatory sequences of the gene encoding ribosomal protein L25 in Candida utilis.

Using a heterologous probe containing a fragment of the L25-gene from Saccharomyces carlsbergensis we have isolated a DNA-fragment of Candida utilis carrying the gene encoding ribosomal protein L25. This gene is present in a single copy on the C. utilis genome, though as two distinguishable alleles. Both alleles have been isolated and sequenced including their flanking regions. The nucleotide sequence of the amino acid coding region of the C. utilis gene turned out to be highly homologous (83%) to the L25-gene of S. carlsbergensis. At the protein level the degree of homology is about 87%. Codon usage in both organisms appears to be somewhat different. Just like the Saccharomyces gene, the L25 gene in C. utilis appears to be split in its 5th codon, though the identity of this codon has changed. Intron as well as 5'- and 3'-flanking sequences have almost completely diverged, with some notable exceptions. Of the intervening sequences the 5'- and 3'-splice sites as well as the putative lariat branch site are conserved. In the 5'-flanking region, at a distance of about 330 n from the initiation codon, a conserved nucleotide element is present that is very similar to the upstream transcription activation site previously found in front of the ribosomal protein genes in Saccharomyces.

Amino Acid Sequence↗

Sequence, evolution and differential expression of the two genes encoding variant small subunits of ribulose bisphosphate carboxylase/oxygenase in Chlamydomonas reinhardtii.

We have sequenced the two genes for the small subunit of ribulose bisphosphate carboxylase/oxygenase (Rubisco) in Chlamydomonas reinhardtii and analyzed their expression. The two genes encode variant small subunits that differ by four amino acid residues. Both genes are expressed and each is transcribed into an RNA of distinct size. The accumulation of the two RNAs changes depending on the growth conditions, so the small subunit composition of Rubisco may be expected to differ in response to the environment. The C. reinhardtii small subunit sequence is homologous to those of vascular plants or cyanobacteria, but is longer at the amino terminus and in internal positions. The number and location of the intervening sequences in the genes from C. reinhardtii and from other plants differ. In several cases, internal length differences in the polypeptide coincide with the positions of introns in the coding sequence. Thus, changes in the exon structure of the genes during evolution may have been accompanied by substantial changes in the encoded protein. The translation and splicing signals in C. reinhardtii are similar to those of other eukaryotes, but the transcription signals are less conserved and the highly biased codon usage is very unusual.

Amino Acid Sequence↗