Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “codon usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

The complete nucleotide sequence of region 1 of the CFA/I fimbrial operon of human enterotoxigenic Escherichia coli.

The production of the plasmid-encoded fimbrial antigen CFA/I of enterotoxigenic Escherichia coli requires two DNA regions: CFA/I region 1 and CFA/I region 2. These two regions are separated by about 40 kb on the wildtype plasmid. CFA/I region 1 contains the structural genes, whereas CFA/I region 2 contains a positive regulator. The first two genes (cfaA and cfaB) and the cfaD' sequence of region 1 have already been described. Here the total nucleotide sequence of region 1 is presented. Two new genes in region 1 are described, named cfaC and cfaE. The GC content of the genes in region 1 is 33.6% which is substantially lower than normally found in E. coli genes (50%). The codon usage also differs from the standard codons used in E. coli.

Amino Acid Sequence↗

Automated Machine Learning Tools to Build Regression Models for Schizosaccharomyces pombe Omics Data.

Machine learning is a powerful tool for analyzing biological data and making useful predictions. The surge of biological data from high-throughput omics technologies has raised the need for modeling approaches capable of tackling such amounts of data, which is pivotal to understanding the nature of complex molecular systems. Here, we show how to construct a simple model using automated machine learning (AutoML) to predict protein abundance in Schizosaccharomyces pombe, using data obtained from codon usage bias and quantitative proteomics.

Machine Learning↗

Hurdles to horizontal gene transfer: species-specific effects of synonymous variation and plasmid copy number determine antibiotic resistance phenotype.

Could codon composition condition the immediate success and the orientation of horizontal gene transfer? Horizontal gene transfer represents a change in the genome of expression of the transferred gene, and experimental evidence has accumulated indicating that the codon composition of a sequence is an important determinant of its compatibility with the translation machinery of the genome in which it is expressed. This suggests that codon composition influences the phenotype and the fitness conferred by a transferred gene and thus the immediate success of the transfer. To directly test this hypothesis, we characterized the resistance conferred by synonymous variants of a gentamicin resistance gene in three bacterial species: Escherichia coli, Acinetobacter baylyi and Pseudomonas aeruginosa. The strongest determinant of the resistance level conferred was the species in which the resistance gene was transferred, very likely because of important differences in the copy number of the plasmid carrying the gene. Significant differences in resistance were also found between synonymous variants within each of the three species, but more importantly, there was a strong interaction between species and variant: variants conferring high resistance in one species confer low resistance in another. However, the similarity in codon usage between the synonymous variants and the host genome only explained part of the phenotypic differences between variants in one species, P. aeruginosa. Further investigation of alternative explanations did not reveal common universal mechanisms across our three bacterial species. We conclude that codon composition can be a determinant of post-horizontal gene transfer success. However, there are multiple paths leading from synonymous sequence to phenotype, and sensitivity to these different paths is species-specific.

Gene Transfer, Horizontal↗

Negative effect of sequential serine codons on expression of foreign genes in Escherichia coli.

Herpes simplex virus encodes a 1298-residue protein designated ICP4 that regulates transcription of viral genes. Structural and functional analyses of ICP4 have been facilitated by production of portions of ICP4 in Escherichia coli. We previously observed that expression of most truncated forms of ICP4 in E. coli was relatively efficient, with the exception of portions of the ICP4 gene approximately between codons 160 and 220. We have now localized the portion of ICP4 that inhibits expression to a serine-rich region from position 176 to 199. Our experimental results suggest that codons within the serine-rich domain do not induce termination of transcription, do not alter the intrinsic stability of mRNA, and do not create a proteolytically sensitive site in this portion of ICP4. Silent mutations that alter codon usage of many of the 19 serine codons in this region had no effect on expression. However, we observed that the level of protein expression was inversely proportional to the number of serine codons in this region. The results are consistent with a model in which the serine-rich domain induces premature termination of translation. This effect is not due to any specific secondary structure in the mRNA or lack of sufficient seryl-tRNA synthetase. It remains to be determined whether premature termination can result from insufficient seryl-charged tRNAs. Our results suggest that foreign genes with more than 20 consecutive serine codons may be poorly expressed in E. coli.

Amino Acid Sequence↗

Expected frequencies of codon use as a function of mutation rates and codon fitnesses.

A method is shown to determine the expected pattern of codon use for any given set of mutation rates between nucleotides and any set of fitnesses for the codons. If it is assumed that mutations to stop codons are lethal then those codons which can mutate in one step to a stop codon tend to be used less frequently. This tendency is however, a very small one and is not likely to be observable within a single gene. Nor is it necessarily a general tendency. For example, the leucine pretermination codons may be used preferentially when mutations to proline are deleterious. It is shown that different mutation rates (eg: transitions occurring more frequently than transversions) may have as large an effect on codon usage as would strong selection for particular codons. For the model presented, an increase in the rate of transitions strongly decreases the expected frequency of UGG and CRR codons. Other codes are moderately affected by such a change in the mutation rates. Many other models can be examined using this method.

Amino Acids↗

scsB, a cDNA encoding the hydrogenosomal beta subunit of succinyl-CoA synthetase from the anaerobic fungus Neocallimastix frontalis.

A clone containing a Neocallimastix frontalis cDNA assumed to encode the beta subunit of succinyl-CoA synthetase (SCSB) was identified by sequence homology with prokaryotic and eukaryotic counter-parts. An open reading frame of 1311 bp was found. The deduced 437 amino acid sequence showed a high degree of identity to the beta-succinyl-CoA synthetase of Escherichia coli (46%), the mitochondrial beta-succinyl-CoA synthetase from pig (48%) and the hydrogenosomal beta-succinyl-CoA synthetase from Trichomonas vaginalis (49%). The G + C content of the succinyl-CoA synthetase coding sequence (43.8%) was considerably higher than that of the 5' (14.8%) and 3' (13.3%) non-translated flanking sequences, as has been observed for other genes from N. frontalis. The codon usage pattern was biased, with only 34 codons used and a strong preference for a pyrimidine (T) in the third positions of the codons. The coding sequence of the beta-succinyl-CoA synthetase cDNA was cloned in an E. coli expression vector encoding a 6(His) tag. The recombinant protein was purified by affinity binding and used to produce polyclonal antibodies. The anti-succinyl-CoA synthetase serum recognized a 45 kDa protein from a N. frontalis fraction enriched for hydrogenosomes and similar polypeptides in two related anaerobic fungi, Piromyces rhizinflata (45 kDa) and Caecomyces communis (47 kDa). Immunocytochemical experiments suggest that succinyl-CoA synthetase is located in the hydrogenosomal matrix. Staining for SCS activity in native electrophoretic gels revealed a band with an apparent molecular weight of approximately 330 kDa. The C-terminus of the succinyl-CoA synthetase sequence was devoid of the typical targeting signals identified so far in microbody proteins, indicating that N. frontalis uses a different signal for sorting SCSB into hydrogenosomes. Based on comparisons with other proteins we propose a putative N-terminal targeting signal for succinyl-CoA synthetase of N. frontalis that shows some of the features of mitochondrial targeting sequences.

Amino Acid Sequence↗

Nucleotide sequence and deduced amino acid sequence of a Plasmodium falciparum actin gene.

The nucleotide sequence of a Plasmodium falciparum actin gene has been established. The gene codes for a protein of 376 amino acids and is not interrupted by introns. The nucleotide sequence reveals an extreme bias in codon usage. Not less than 85% of the codons possess an A or T at the third position. As has been found for the actins in other unicellular eukaryotes, P. falciparum actin is related both to vertebrate cytoplasmic and vertebrate muscle specific actins. However, the malarial actin is one of the most alpha-like actins hitherto found in lower eukaryotes.

Actins↗

Nucleotide sequence of a gene from the Pseudomonas transposon Tn501 encoding mercuric reductase.

We have determined the nucleotide sequence of the merA gene from the mercury-resistance transposon Tn501 and have predicted the structure of the gene product, mercuric reductase. The DNA sequence predicts a polypeptide of Mr 58 660, the primary structure of which shows strong homologies to glutathione reductase and lipoamide dehydrogenase, but mercuric reductase contains as additional N-terminal region that may form a separate domain. The implications of these comparisons for the tertiary structure and mechanism of mercuric reductase are discussed. The DNA sequence presented here has an overall G+C content of 65.1 mol%, typical of the bulk DNA of Pseudomonas aeruginosa from which Tn501 was originally isolated. Analysis of the codon usage in the merA gene shows that codons with C or G at the third position are preferentially utilized.

Amino Acid Sequence↗

Occurrence of unmodified adenine and uracil at the first position of anticodon in threonine tRNAs in Mycoplasma capricolum.

Codon usage pattern in the threonine four-codon (ACN) box in Mycoplasma capricolum is strongly biased towards adenine and uracil for the third base of codons. Codons ending in uracil or adenine, especially ACU, predominate over ACC and ACG. This bacterium contains two isoacceptor threonine tRNAs having anticodon sequences AGU and UGU, both with unmodified first nucleotides. It would thus appear that ACN codons are translated in an unusual way; tRNA(Thr)(AGU) would translate the most abundantly used codon ACU exclusively, because adenine at the first anticodon position can, according to the wobble rule, pair only with uracil of the third codon position. The tRNA(Thr)(UGU) would mainly be responsible for translation of three other codons, ACA, ACG, and ACC. Anticodon UGU would also be used for reading codon ACU as a redundancy of tRNA(Thr)-(AGU), as deduced from the mitochondrial code where unmodified uracil at the first anticodon position can pair with adenine, cytosine, guanine, and uracil by four-way wobble. The tRNA(Thr)(AGU) has much higher sequence homology to tRNA(Thr)(UGU) from M. capricolum (88%), Bacillus subtilis (77%) and Escherichia coli (86%) than to tRNA(Thr)(GGU) from B. subtilis (66%) and E. coli (63%), suggesting that tRNA(Thr)-(AGU) has been derived from tRNA(Thr)(UGU), but not from tRNA(Thr)(GGU).

Adenine↗

Drosophila melanogaster mitochondrial DNA: gene organization and evolutionary considerations.

The sequence of a 8351-nucleotide mitochondrial DNA (mtDNA) fragment has been obtained extending the knowledge of the Drosophila melanogaster mitochondrial genome to 90% of its coding region. The sequence encodes seven polypeptides, 12 tRNAs and the 3' end of the 16S rRNA and CO III genes. The gene organization is strictly conserved with respect to the Drosophila yakuba mitochondrial genome, and different from that found in mammals and Xenopus. The high A + T content of D. melanogaster mitochondrial DNA is reflected in a reiterative codon usage, with more than 90% of the codons ending in T or A, G + C rich codons being practically absent. The average level of homology between the D. melanogaster and D. yakuba sequences is very high (roughly 94%), although insertion and deletions have been detected in protein, tRNA and large ribosomal genes. The analysis of nucleotide changes reveals a similar frequency for transitions and transversions, and reflects a strong bias against G + C on both strands. The predominant type of transition is strand specific.

Amino Acid Sequence↗

Ureaplasma urealyticum urease genes; use of a UGA tryptophan codon.

Nucleotide sequence analysis of a Ureaplasma urealyticum DNA fragment, homologous to cloned urease genes of other prokaryotes, revealed three consecutive open reading frames. The molecular weights of the three deduced polypeptides are 11.2 kD, 13.6 kD and 66.6 kD. These values are consistent with the size of the three subunits previously reported for purified native urease. A significant sequence homology was found between the three polypeptides of the ureaplasmal urease and the single polypeptide of jack bean (Canavalia ensiformis) urease. Codon usage indicates that UGA is a tryptophan codon in this mollicute. Use of polymerase chain reactions has disclosed the existence of genetic polymorphism among the urease genes of different serotypes of U. urealyticum.

Amino Acid Sequence↗

Identification, genetic analysis and DNA sequence of a 7.8-kb virulence region of the Salmonella typhimurium virulence plasmid.

The 90-kilobase (kb) virulence plasmid of Salmonella typhimurium is responsible for invasion from the intestines to mesenteric lymph nodes and spleens of orally inoculated mice. We used Tn5 and aminoglycoside phosphotransferase (aph) gene insertion mutagenesis and deletion mutagenesis of a previously identified 14-kb virulence region to reduce this virulence region to 7.8kb. The 7.8-kb virulence region subcloned into a low copy-number vector conferred a wild-type level of splenic infection to virulence plasmid-cured S. typhimurium and conferred essentially a wild-type oral LD50. Insertion mutagenesis identified five loci essential for virulence, and DNA sequence analysis of the virulence region identified six open reading frames. Expected protein products were identified from four of the six genes, with three of the proteins identified as doublet bands in Escherichia coli minicells. Three of the five mutated genes were able to be complemented by clones containing only the corresponding wild-type gene. Only one of the five deduced amino acid sequences, that of the positive regulatory element, SpvR, possessed significant homology to other proteins. The codon usage for the virulence genes showed no codon bias, which is consistent with the low levels of expression observed for the corresponding proteins. Consensus promoters for several different sigma factors were identified upstream of several of the genes, whereas only consensus Rho-dependent termination sequences were observed between certain of the genes. The operon structure of this virulence region therefore appears to be complex. The construction of the cloned 7.8-kb virulence region and the determination of the DNA sequence will aid in the further genetic analysis of the five plasmid-encoded virulence genes of S. typhimurium.

Amino Acid Sequence↗

Molecular cloning of DNA complementary to bovine growth hormone mRNA.

We have cloned DNA complementary to mRNA coding for bovine growth hormone (bGH). Double-stranded DNA complementary to bovine pituitary mRNA was inserted into the Pst I site of plasmid pBR322 by the dC x dG tailing technique and amplified in E. coli x 1776. A recombinant plasmid containing bGH cDNA ws identified by hybridization to cloned rat growth hormone cDNA. It contains the entire coding and 3'-untranslated regions and 31 bases in the 5'-untranslated region. Nucleotide sequence analysis determined the sequence of the 26-amino acid signal peptide and confirmed the published amino acid sequence of the secreted hormone at all but 2 residues. Codon usage is nonrandom, with 81.7% of the codons ending in G or C. The nucleotide sequence of bGH mRNA is 83.9% homologous with rat GH mRNA and 76.5% homologous with human GH mRNA, while the respective amino acid sequence homologies are 83.5% and 66.8%.

Amino Acid Sequence↗

Molecular cloning of the chicken melanocortin 2 (ACTH)-receptor gene.

The chicken melanocortin 2-receptor (MC2-R) gene was isolated. It is found to be a single copy gene encoding a 357 amino acid protein, sharing 65.8-68.7% identity with mammalian counterparts. The chicken MC2-R mRNA is expressed in the adrenal and spleen, suggesting that the receptor mediates both endocrine and immunoregulatory functions of ACTH in the chicken. The amino acid sequence of the chicken MC2-R is collinear with those of other subtypes of MC-R, whereas all cloned mammalian MC2-Rs contain a gap in the third intracellular loop, suggesting that mammalian MC2-R molecules have evolved by lacking a part of the domain which determines the specificity of signal transduction in G-protein coupled receptors. Interestingly, the codon usage differs dramatically between MC1-R and MC2-R in the chicken; the GC-contents at the third codon position in MC1-R and MC2-R are 94.6 and 50.6%, respectively. It may reflect selective constraints on the usage of synonymous codons.

Adrenal Glands↗

The large mitochondrial genome of Syndiclis anlungensis (Lauraceae): Genome structure, comparative analysis, and phylogenetic relationships among Syndiclis species.

The complete mitochondrial genome (mitogenome) of Syndiclis anlungensis, a critically endangered tropical tree, was determined in this study. The mitogenome spans 2,368,454&#xa0;bp across four contigs and harbors 41 protein-coding genes, 22 tRNA genes, and three rRNA genes. Potential mutation regions, including 1317 repeat sequences and 698 simple sequence repeats (SSRs), were accurately located in the S. anlungensis mitogenome. Sixty-five transferred fragments of the repeats were found between its mitochondrial and chloroplast genomes. When compared to three other Laurales mitogenomes, extensive gene order shuffling is evident, leaving only five conserved gene clusters intact. Codon usage analysis reveals a pronounced A/T bias in both mitochondrial and chloroplast genes, and three mitochondrial genes (atp9, rps19, and sdh3) stand out for their high divergence across eleven Syndiclis taxa. Selection analyses indicate strong purifying pressure on rpl2, rpl16, and sdh3 (Ka/Ks&#xa0;<&#xa0;1), with no positive selection detected. Using 41 mitochondrial protein-coding gene sequences from sixteen and three individuals of Syndiclis and Beilschmiedia species, respectively, our phylogenetic tree recovers Syndiclis as monophyletic, with two well-supported clades: one includes S. anlungensis, S. chinensis, S. lotungensis, S. marlipoensis, and a putative new Syndiclis species from Yunnan; the other contains S. furfuracea, S. hongkongensis, S. kwangsiensis, and three putative new Syndiclis species from Guangdong and Vietnam.

Genome, Mitochondrial↗

Influence of the codon following the AUG initiation codon on the expression of a modified lacZ gene in Escherichia coli.

In a lacZ expression vector (pMC1403Plac), all 64 codons were introduced immediately 3' from the AUG initiation codon. The expression of the second codon variants was measured by immunoprecipitation of the plasmid-coded fusion proteins. A 15-fold difference in expression was found among the codon variants. No distinct correlation could be made with the level of tRNA corresponding to the codons and large differences were observed between synonymous codons that use the same tRNA. Therefore the effect of the second codon is likely to be due to the influence of its composing nucleotides, presumably on the structure of the ribosomal binding site. An analysis of the known sequences of a large number of Escherichia coli genes shows that the use of codons in the second position deviates strongly from the overall codon usage in E. coli. It is proposed that codon selection at the second position is not based on requirements of the gene product (a protein) but is determined by factors governing gene regulation at the initiation step of translation.

Base Sequence↗

The close proximity of Escherichia coli genes: consequences for stop codon and synonymous codon use.

It is shown that synonymous codon usage is less biased in favor of those codons preferred by highly expressed genes at the end of Escherichia coli genes than in the middle. This appears to be due to the close proximity of many E. coli genes. It is shown that a substantial number of genes overlap either the Shine-Dalgarno sequence or the coding sequence of the next gene on the chromosome and that the codons that overlap have lower synonymous codon bias than those which do not. It is also shown that there is an increase in the frequency of A-ending codons, and a decrease in the frequency of G-ending codons at the end of E. coli genes that lie close to another gene. It is suggested that these trends in composition could be associated with selection against the formation of mRNA secondary structure near the start of the next gene on the chromosome. Stop codon use is also affected by the close proximity of genes; many genes are forced to use TGA and TAG stop codons because they terminate either within the Shine-Dalgarno or coding sequence of the next gene on the chromosome. The implications these results have for the evolution of synonymous codon use are discussed.

Base Sequence↗

Hydropathic characteristics of adenovirus hexons.

The complete nucleotide sequence and the predicted amino acid sequence of the adenovirus type 7 hexon gene were determined. The hydropathy of the hexon proteins from human adenovirus types 2, 3, 4, 5, 7, 12, 16, 40, 41, and 48, bovine adenovirus type 3, murine adenovirus type 1, and avian adenovirus types 1 and 10 was analysed. The presence of purines and pyrimidines in the second position of the codons was correlated to hydrophilicity and hydrophobicity, respectively. Comparison of the hydrophilicity plots of eight hexons showed seven hypervariable regions to be distributed on the surface. A large portion of the hypervariable regions manifests hydrophilicity. The strength of the surface charge accumulated on the hydrophilic and hydrophobic regions correlated to the tissue tropism of the different adenovirus types. Analysis of codon usage for adenovirus hexons showed that among synonymous codons those with cytidine in the third position were preferably used to a great extent. Analysis of the nucleotide and amino acid sequence pair distances and the phylogenetic tree of 14 hexon proteins showed members of subgenera B, D and E to be closely related, especially Ad4 and Ad16, and subgenus A to be closely related to subgenus F.

Adenoviruses, Human↗