Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Preferred amino acids and thermostability.

Most organisms grow at temperatures from 20 to 50 degrees C, but some prokaryotes, including Archaea and Bacteria, are capable of withstanding higher temperatures, from 60 to >100 degrees C. Their biomolecules, especially proteins, must be sufficiently stable to function under these extreme conditions; however, the basis for thermostability remains elusive. We investigated the preferential usage of certain groupings of amino acids and codons in thermally adapted organisms, by comparative proteome analysis, using 28 complete genomes from 18 mesophiles (M), 4 thermophiles (T), and 6 hyperthermophiles (HT). Whenever the percent of glutamate (E) and lysine (K) increased in the HT proteomes, the percent of glutamine (Q) and histidine (H) decreased, so that the E + K/Q + H ratio was >4.5; it was <2.5 in the M proteomes, and 3.2 to 4.6 in T. The E + K/Q + H ratios for chaperonins, potentially thermostable proteins, were higher than their proteome ratios, whereas for DNA ligases, which are not necessarily thermostable, they followed the proteome ratios. Analysis of codon usage revealed that HT had more AGR codons for Arg than they did CGN codons, which were more common in mesophiles. The E + K/Q + H ratio may provide a useful marker for distinguishing HT, T and M prokaryotes, and the high percentage of the amino acid couple E + K, consistently associated with a low percentage of the pair Q + H, could contribute to protein thermostability. The preponderance of AGR codons for Arg is a signature of all HT so far analyzed. The E + K/Q + H ratio and the codon bias for Arg are apparently not related to phylogeny. HT members of the Bacteria show the same values as the HT members of the Archaea; the values for T organisms are related to their lifestyle (intermediate temperature) and not to their domain (Archaea) and the values for M are similar in Eukarya, Bacteria and Archaea.

Adaptation, Biological↗

Automated Machine Learning Tools to Build Regression Models for Schizosaccharomyces pombe Omics Data.

Machine learning is a powerful tool for analyzing biological data and making useful predictions. The surge of biological data from high-throughput omics technologies has raised the need for modeling approaches capable of tackling such amounts of data, which is pivotal to understanding the nature of complex molecular systems. Here, we show how to construct a simple model using automated machine learning (AutoML) to predict protein abundance in Schizosaccharomyces pombe, using data obtained from codon usage bias and quantitative proteomics.

Machine Learning↗

Hurdles to horizontal gene transfer: species-specific effects of synonymous variation and plasmid copy number determine antibiotic resistance phenotype.

Could codon composition condition the immediate success and the orientation of horizontal gene transfer? Horizontal gene transfer represents a change in the genome of expression of the transferred gene, and experimental evidence has accumulated indicating that the codon composition of a sequence is an important determinant of its compatibility with the translation machinery of the genome in which it is expressed. This suggests that codon composition influences the phenotype and the fitness conferred by a transferred gene and thus the immediate success of the transfer. To directly test this hypothesis, we characterized the resistance conferred by synonymous variants of a gentamicin resistance gene in three bacterial species: Escherichia coli, Acinetobacter baylyi and Pseudomonas aeruginosa. The strongest determinant of the resistance level conferred was the species in which the resistance gene was transferred, very likely because of important differences in the copy number of the plasmid carrying the gene. Significant differences in resistance were also found between synonymous variants within each of the three species, but more importantly, there was a strong interaction between species and variant: variants conferring high resistance in one species confer low resistance in another. However, the similarity in codon usage between the synonymous variants and the host genome only explained part of the phenotypic differences between variants in one species, P. aeruginosa. Further investigation of alternative explanations did not reveal common universal mechanisms across our three bacterial species. We conclude that codon composition can be a determinant of post-horizontal gene transfer success. However, there are multiple paths leading from synonymous sequence to phenotype, and sensitivity to these different paths is species-specific.

Gene Transfer, Horizontal↗

Negative effect of sequential serine codons on expression of foreign genes in Escherichia coli.

Herpes simplex virus encodes a 1298-residue protein designated ICP4 that regulates transcription of viral genes. Structural and functional analyses of ICP4 have been facilitated by production of portions of ICP4 in Escherichia coli. We previously observed that expression of most truncated forms of ICP4 in E. coli was relatively efficient, with the exception of portions of the ICP4 gene approximately between codons 160 and 220. We have now localized the portion of ICP4 that inhibits expression to a serine-rich region from position 176 to 199. Our experimental results suggest that codons within the serine-rich domain do not induce termination of transcription, do not alter the intrinsic stability of mRNA, and do not create a proteolytically sensitive site in this portion of ICP4. Silent mutations that alter codon usage of many of the 19 serine codons in this region had no effect on expression. However, we observed that the level of protein expression was inversely proportional to the number of serine codons in this region. The results are consistent with a model in which the serine-rich domain induces premature termination of translation. This effect is not due to any specific secondary structure in the mRNA or lack of sufficient seryl-tRNA synthetase. It remains to be determined whether premature termination can result from insufficient seryl-charged tRNAs. Our results suggest that foreign genes with more than 20 consecutive serine codons may be poorly expressed in E. coli.

Amino Acid Sequence↗

Expected frequencies of codon use as a function of mutation rates and codon fitnesses.

A method is shown to determine the expected pattern of codon use for any given set of mutation rates between nucleotides and any set of fitnesses for the codons. If it is assumed that mutations to stop codons are lethal then those codons which can mutate in one step to a stop codon tend to be used less frequently. This tendency is however, a very small one and is not likely to be observable within a single gene. Nor is it necessarily a general tendency. For example, the leucine pretermination codons may be used preferentially when mutations to proline are deleterious. It is shown that different mutation rates (eg: transitions occurring more frequently than transversions) may have as large an effect on codon usage as would strong selection for particular codons. For the model presented, an increase in the rate of transitions strongly decreases the expected frequency of UGG and CRR codons. Other codes are moderately affected by such a change in the mutation rates. Many other models can be examined using this method.

Amino Acids↗

scsB, a cDNA encoding the hydrogenosomal beta subunit of succinyl-CoA synthetase from the anaerobic fungus Neocallimastix frontalis.

A clone containing a Neocallimastix frontalis cDNA assumed to encode the beta subunit of succinyl-CoA synthetase (SCSB) was identified by sequence homology with prokaryotic and eukaryotic counter-parts. An open reading frame of 1311 bp was found. The deduced 437 amino acid sequence showed a high degree of identity to the beta-succinyl-CoA synthetase of Escherichia coli (46%), the mitochondrial beta-succinyl-CoA synthetase from pig (48%) and the hydrogenosomal beta-succinyl-CoA synthetase from Trichomonas vaginalis (49%). The G + C content of the succinyl-CoA synthetase coding sequence (43.8%) was considerably higher than that of the 5' (14.8%) and 3' (13.3%) non-translated flanking sequences, as has been observed for other genes from N. frontalis. The codon usage pattern was biased, with only 34 codons used and a strong preference for a pyrimidine (T) in the third positions of the codons. The coding sequence of the beta-succinyl-CoA synthetase cDNA was cloned in an E. coli expression vector encoding a 6(His) tag. The recombinant protein was purified by affinity binding and used to produce polyclonal antibodies. The anti-succinyl-CoA synthetase serum recognized a 45 kDa protein from a N. frontalis fraction enriched for hydrogenosomes and similar polypeptides in two related anaerobic fungi, Piromyces rhizinflata (45 kDa) and Caecomyces communis (47 kDa). Immunocytochemical experiments suggest that succinyl-CoA synthetase is located in the hydrogenosomal matrix. Staining for SCS activity in native electrophoretic gels revealed a band with an apparent molecular weight of approximately 330 kDa. The C-terminus of the succinyl-CoA synthetase sequence was devoid of the typical targeting signals identified so far in microbody proteins, indicating that N. frontalis uses a different signal for sorting SCSB into hydrogenosomes. Based on comparisons with other proteins we propose a putative N-terminal targeting signal for succinyl-CoA synthetase of N. frontalis that shows some of the features of mitochondrial targeting sequences.

Amino Acid Sequence↗

A phylogenomic study of the OCTase genes in Pseudomonas syringae pathovars: the horizontal transfer of the argK-tox cluster and the evolutionary history of OCTase genes on their genomes.

Phytopathogenic Pseudomonas syringae is subdivided into about 50 pathovars due to their conspicuous differentiation with regard to pathogenicity. Based on the results of a phylogenetic analysis of four genes (gyrB, rpoD, hrpL, and hrpS), Sawada et al. (1999) showed that the ancestor of P. syringae had diverged into at least three monophyletic groups during its evolution. Physical maps of the genomes of representative strains of these three groups were constructed, which revealed that each strain had five rrn operons which existed on one circular genome. The fact that the structure and size of genomes vary greatly depending on the pathovar shows that P. syringae genomes are quite rich in plasticity and that they have undergone large-scale genomic rearrangements. Analyses of the codon usage and the GC content at the codon third position, in conjunction with phylogenomic analyses, showed that the gene cluster involved in phaseolotoxin synthesis (argK-tox cluster) expanded its distribution by conducting horizontal transfer onto the genomes of two P. syringae pathovars (pv. actinidiae and pv. phaseolicola) from bacterial species distantly related to P. syringae and that its acquisition was quite recent (i.e., after the ancestor of P. syringae diverged into the respective pathovars). Furthermore, the results of a detailed analysis of argK [an anabolic ornithine carbamoyltransferase (anabolic OCTase) gene], which is present within the argK-tox cluster, revealed the plausible process of generation of an unusual composition of the OCTase genes on the genomes of these two phaseolotoxin-producing pathovars: a catabolic OCTase gene (equivalent to the orthologue of arcB of P. aeruginosa) and an anabolic OCTase gene (argF), which must have been formed by gene duplication, have first been present on the genome of the ancestor of P. syringae; the catabolic OCTase gene has been deleted; the ancestor has diverged into the respective pathovars; the foreign-originated argK-tox cluster has horizontally transferred onto the genomes of pv. actinidiae and pv. phaseolicola; and hence two copies of only the anabolic OCTase genes (argK and argF) came to exist on the genomes of these two pathovars. Thus, the horizontal gene transfer and the genomic rearrangement were proven to have played an important role in the pathogenic differentiation and diversification of P. syringae.

Evolution, Molecular↗

Nucleotide sequence and deduced amino acid sequence of a Plasmodium falciparum actin gene.

The nucleotide sequence of a Plasmodium falciparum actin gene has been established. The gene codes for a protein of 376 amino acids and is not interrupted by introns. The nucleotide sequence reveals an extreme bias in codon usage. Not less than 85% of the codons possess an A or T at the third position. As has been found for the actins in other unicellular eukaryotes, P. falciparum actin is related both to vertebrate cytoplasmic and vertebrate muscle specific actins. However, the malarial actin is one of the most alpha-like actins hitherto found in lower eukaryotes.

Actins↗

The effective number of codons for individual amino acids: some codons are more optimal than others.

The aim of this study was to evaluate the codon bias using the effective number of codons for individual amino acids (N(c)(AA)) and to assess the codon bias in relation to the definition of optimal codons, using Escherichia coli as a model organism. We show that a general correlation exists between the effective number of codons (Ncirc(c)) or codon adaptation index (CAI) and N(c)(AA), but that this correlation is not equally strong for all amino acids within a degeneracy group. For example, leucine codons contribute more to Ncirc(c) and the codon adaptation index than serine codons. A possible explanation is that some optimal codons are more optimal than others, in terms of the selectional advantage they offer. This hypothesis is confirmed by further analysis on the correlations that exist between values for relative synonymous codon usage (RSCU), N(c)(AA), and the codon adaptation index.

Amino Acids↗

Nucleotide sequence of a gene from the Pseudomonas transposon Tn501 encoding mercuric reductase.

We have determined the nucleotide sequence of the merA gene from the mercury-resistance transposon Tn501 and have predicted the structure of the gene product, mercuric reductase. The DNA sequence predicts a polypeptide of Mr 58 660, the primary structure of which shows strong homologies to glutathione reductase and lipoamide dehydrogenase, but mercuric reductase contains as additional N-terminal region that may form a separate domain. The implications of these comparisons for the tertiary structure and mechanism of mercuric reductase are discussed. The DNA sequence presented here has an overall G+C content of 65.1 mol%, typical of the bulk DNA of Pseudomonas aeruginosa from which Tn501 was originally isolated. Analysis of the codon usage in the merA gene shows that codons with C or G at the third position are preferentially utilized.

Amino Acid Sequence↗

Occurrence of unmodified adenine and uracil at the first position of anticodon in threonine tRNAs in Mycoplasma capricolum.

Codon usage pattern in the threonine four-codon (ACN) box in Mycoplasma capricolum is strongly biased towards adenine and uracil for the third base of codons. Codons ending in uracil or adenine, especially ACU, predominate over ACC and ACG. This bacterium contains two isoacceptor threonine tRNAs having anticodon sequences AGU and UGU, both with unmodified first nucleotides. It would thus appear that ACN codons are translated in an unusual way; tRNA(Thr)(AGU) would translate the most abundantly used codon ACU exclusively, because adenine at the first anticodon position can, according to the wobble rule, pair only with uracil of the third codon position. The tRNA(Thr)(UGU) would mainly be responsible for translation of three other codons, ACA, ACG, and ACC. Anticodon UGU would also be used for reading codon ACU as a redundancy of tRNA(Thr)-(AGU), as deduced from the mitochondrial code where unmodified uracil at the first anticodon position can pair with adenine, cytosine, guanine, and uracil by four-way wobble. The tRNA(Thr)(AGU) has much higher sequence homology to tRNA(Thr)(UGU) from M. capricolum (88%), Bacillus subtilis (77%) and Escherichia coli (86%) than to tRNA(Thr)(GGU) from B. subtilis (66%) and E. coli (63%), suggesting that tRNA(Thr)-(AGU) has been derived from tRNA(Thr)(UGU), but not from tRNA(Thr)(GGU).

Adenine↗

Drosophila melanogaster mitochondrial DNA: gene organization and evolutionary considerations.

The sequence of a 8351-nucleotide mitochondrial DNA (mtDNA) fragment has been obtained extending the knowledge of the Drosophila melanogaster mitochondrial genome to 90% of its coding region. The sequence encodes seven polypeptides, 12 tRNAs and the 3' end of the 16S rRNA and CO III genes. The gene organization is strictly conserved with respect to the Drosophila yakuba mitochondrial genome, and different from that found in mammals and Xenopus. The high A + T content of D. melanogaster mitochondrial DNA is reflected in a reiterative codon usage, with more than 90% of the codons ending in T or A, G + C rich codons being practically absent. The average level of homology between the D. melanogaster and D. yakuba sequences is very high (roughly 94%), although insertion and deletions have been detected in protein, tRNA and large ribosomal genes. The analysis of nucleotide changes reveals a similar frequency for transitions and transversions, and reflects a strong bias against G + C on both strands. The predominant type of transition is strand specific.

Amino Acid Sequence↗

Ureaplasma urealyticum urease genes; use of a UGA tryptophan codon.

Nucleotide sequence analysis of a Ureaplasma urealyticum DNA fragment, homologous to cloned urease genes of other prokaryotes, revealed three consecutive open reading frames. The molecular weights of the three deduced polypeptides are 11.2 kD, 13.6 kD and 66.6 kD. These values are consistent with the size of the three subunits previously reported for purified native urease. A significant sequence homology was found between the three polypeptides of the ureaplasmal urease and the single polypeptide of jack bean (Canavalia ensiformis) urease. Codon usage indicates that UGA is a tryptophan codon in this mollicute. Use of polymerase chain reactions has disclosed the existence of genetic polymorphism among the urease genes of different serotypes of U. urealyticum.

Amino Acid Sequence↗

Identification, genetic analysis and DNA sequence of a 7.8-kb virulence region of the Salmonella typhimurium virulence plasmid.

The 90-kilobase (kb) virulence plasmid of Salmonella typhimurium is responsible for invasion from the intestines to mesenteric lymph nodes and spleens of orally inoculated mice. We used Tn5 and aminoglycoside phosphotransferase (aph) gene insertion mutagenesis and deletion mutagenesis of a previously identified 14-kb virulence region to reduce this virulence region to 7.8kb. The 7.8-kb virulence region subcloned into a low copy-number vector conferred a wild-type level of splenic infection to virulence plasmid-cured S. typhimurium and conferred essentially a wild-type oral LD50. Insertion mutagenesis identified five loci essential for virulence, and DNA sequence analysis of the virulence region identified six open reading frames. Expected protein products were identified from four of the six genes, with three of the proteins identified as doublet bands in Escherichia coli minicells. Three of the five mutated genes were able to be complemented by clones containing only the corresponding wild-type gene. Only one of the five deduced amino acid sequences, that of the positive regulatory element, SpvR, possessed significant homology to other proteins. The codon usage for the virulence genes showed no codon bias, which is consistent with the low levels of expression observed for the corresponding proteins. Consensus promoters for several different sigma factors were identified upstream of several of the genes, whereas only consensus Rho-dependent termination sequences were observed between certain of the genes. The operon structure of this virulence region therefore appears to be complex. The construction of the cloned 7.8-kb virulence region and the determination of the DNA sequence will aid in the further genetic analysis of the five plasmid-encoded virulence genes of S. typhimurium.

Amino Acid Sequence↗

Genome sequencing and annotation: an overview.

Many microbial genome sequences have been determined, and more new genome projects are ongoing. Shotgun sequencing of randomly cloned short pieces of genomic DNA can provide a simple way of determining whole genome sequences. This process requires sequencing of many fragments, compilation of the separate sequences into one contiguous sequence, and careful editing of the assembled sequence. The genes present on the microbial genome are then predicted using clues derived from typical gene features, such as codon usage, ribosomal binding sequences, and bacterial initiation codons. Function of genes is predicted by homology searches performed against either public or well-established protein databases. This chapter discusses each of these stages in a genome-sequencing project.

Amino Acid Sequence↗

Molecular cloning of DNA complementary to bovine growth hormone mRNA.

We have cloned DNA complementary to mRNA coding for bovine growth hormone (bGH). Double-stranded DNA complementary to bovine pituitary mRNA was inserted into the Pst I site of plasmid pBR322 by the dC x dG tailing technique and amplified in E. coli x 1776. A recombinant plasmid containing bGH cDNA ws identified by hybridization to cloned rat growth hormone cDNA. It contains the entire coding and 3'-untranslated regions and 31 bases in the 5'-untranslated region. Nucleotide sequence analysis determined the sequence of the 26-amino acid signal peptide and confirmed the published amino acid sequence of the secreted hormone at all but 2 residues. Codon usage is nonrandom, with 81.7% of the codons ending in G or C. The nucleotide sequence of bGH mRNA is 83.9% homologous with rat GH mRNA and 76.5% homologous with human GH mRNA, while the respective amino acid sequence homologies are 83.5% and 66.8%.

Amino Acid Sequence↗

Molecular cloning of the chicken melanocortin 2 (ACTH)-receptor gene.

The chicken melanocortin 2-receptor (MC2-R) gene was isolated. It is found to be a single copy gene encoding a 357 amino acid protein, sharing 65.8-68.7% identity with mammalian counterparts. The chicken MC2-R mRNA is expressed in the adrenal and spleen, suggesting that the receptor mediates both endocrine and immunoregulatory functions of ACTH in the chicken. The amino acid sequence of the chicken MC2-R is collinear with those of other subtypes of MC-R, whereas all cloned mammalian MC2-Rs contain a gap in the third intracellular loop, suggesting that mammalian MC2-R molecules have evolved by lacking a part of the domain which determines the specificity of signal transduction in G-protein coupled receptors. Interestingly, the codon usage differs dramatically between MC1-R and MC2-R in the chicken; the GC-contents at the third codon position in MC1-R and MC2-R are 94.6 and 50.6%, respectively. It may reflect selective constraints on the usage of synonymous codons.

Adrenal Glands↗

The large mitochondrial genome of Syndiclis anlungensis (Lauraceae): Genome structure, comparative analysis, and phylogenetic relationships among Syndiclis species.

The complete mitochondrial genome (mitogenome) of Syndiclis anlungensis, a critically endangered tropical tree, was determined in this study. The mitogenome spans 2,368,454&#xa0;bp across four contigs and harbors 41 protein-coding genes, 22 tRNA genes, and three rRNA genes. Potential mutation regions, including 1317 repeat sequences and 698 simple sequence repeats (SSRs), were accurately located in the S. anlungensis mitogenome. Sixty-five transferred fragments of the repeats were found between its mitochondrial and chloroplast genomes. When compared to three other Laurales mitogenomes, extensive gene order shuffling is evident, leaving only five conserved gene clusters intact. Codon usage analysis reveals a pronounced A/T bias in both mitochondrial and chloroplast genes, and three mitochondrial genes (atp9, rps19, and sdh3) stand out for their high divergence across eleven Syndiclis taxa. Selection analyses indicate strong purifying pressure on rpl2, rpl16, and sdh3 (Ka/Ks&#xa0;<&#xa0;1), with no positive selection detected. Using 41 mitochondrial protein-coding gene sequences from sixteen and three individuals of Syndiclis and Beilschmiedia species, respectively, our phylogenetic tree recovers Syndiclis as monophyletic, with two well-supported clades: one includes S. anlungensis, S. chinensis, S. lotungensis, S. marlipoensis, and a putative new Syndiclis species from Yunnan; the other contains S. furfuracea, S. hongkongensis, S. kwangsiensis, and three putative new Syndiclis species from Guangdong and Vietnam.

Genome, Mitochondrial↗