Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codon Usage”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Isolation and characterization of a Ustilago maydis glyceraldehyde-3-phosphate dehydrogenase-encoding gene.

The complete nucleotide sequence of the glyceraldehyde-3-phosphate dehydrogenase gene from the corn smut fungus Ustilago maydis is reported. The gene encodes a 337-amino acid protein, parts of which show sequence identity to corresponding regions of GAPDH-encoding genes from other organisms. A single, putative 407-bp intron interrupts the tenth codon. Codon usage is highly biased for codons ending in cytosine.

Amino Acid Sequence↗

Enhanced readthrough of opal (UGA) stop codons and production of Mycoplasma pneumoniae P1 epitopes in Escherichia coli.

Expression of mycoplasma sequences in Escherichia coli is often hindered by an unusual mycoplasmal codon usage pattern: the UGA stop codon is utilized for tryptophan. This may result in the truncation of cloned proteins and may prevent the detection of products of many cloned genes. To circumvent this translation barrier, we have developed an expression system for the production of mycoplasma proteins in E. coli. The efficiency of an opal suppressor tRNA (trpT176) was augmented with other suppressor mutations (prfB3 or rrsB(SuUGA-delta C1054)) which influence termination events. System efficacy was analyzed by employing suppressor mutations in the expression of TGA-containing sequences from the P1 protein-encoding gene of Mycoplasma pneumoniae.

Adhesins, Bacterial↗

Molecular cloning of, and phylogenetic analysis of, an actin in Naegleria fowleri.

We cloned and sequenced an intronless actin gene from the amoebo-flagellate Naegleria fowleri, LEE strain, an opportunistic pathogen of man. Codon usage and third-position-codon nucleotide frequency were significantly different from Acanthamoeba, another amoeba genus which also includes opportunistic pathogens of man. Between the two amoebae, actin peptide sequences were 92.8% similar, while nucleotide sequences were only 70% similar. A phylogenetic reconstruction of actin amino acid sequences, using a distance method, placed Naegleria in a cluster with Plasmodium and Entamoeba.

Actins↗

Cloning and sequencing of a phospholipase C gene of Clostridium perfringens.

The gene encoding phospholipase C (alpha-toxin) of Clostridium perfringens was cloned into lambda gt10. The maximal size of the coding region was 1.4 kb and the minimum was 1.1 kb as determined by subcloning into the vector pBR322 and testing for activity. The nucleotide sequence of this region contained a single open reading frame of 1194 bp corresponding to a protein of Mr 45473 with a possible N-terminal signal sequence of 28 amino acids which when removed, would give a mature protein of Mr 42521. This is in good agreement with the reported size of 43 kDa. The coding region has a dG + dC content of 33.7%, and the codon usage displays a pronounced preference for codons with the lowest dG + dC content.

Amino Acid Sequence↗

Characterization of catalase transcripts and their differential expression in maize.

In maize, the three unlinked catalase (EC 1.11.1.6) structural genes (Cat1, Cat2 and Cat3) are differentially expressed temporally, spatially and in response to environmental signals in the developing seedling. In order to understand more fully the molecular mechanisms involved in catalase gene expression, full-length cDNA clones representing the maize Cat1, Cat2 and Cat3 transcripts were isolated and characterized. DNA sequence analysis confirmed that each cDNA encodes a unique catalase protein. Gene-specific probes for the three maize catalase cDNAs were isolated and used to probe blots of poly(A)+ RNA isolated from various maize tissues. Cat1 mRNA was found in scutella, milky endosperm of immature kernels, leaves and epicotyls. The Cat2 mRNA was present primarily in post-germinative scutella, with lower levels in leaves and epicotyls. Cat3 mRNA was detected primarily in epicotyls and, to a lesser extent, in leaves and scutella. The gene-specific probes hybridized with maize genomic DNA blots in simple, but unique patterns, indicating that there is one, or a very few copies of each catalase gene. The coding region of the Cat3 cDNA comprised 66% G + C, which led to a strong codon usage bias in this gene. This codon bias was also seen with the Cat2 transcripts, but not with those for Cat1. A high degree of similarity was found between the maize catalase nucleic acid and deduced amino-acid sequences and those of sweet potato and rat liver catalase.

Amino Acid Sequence↗

Structural proteins of mycobacteriophage I3: cloning, expression and sequence analysis of a gene encoding a 70-kDa structural protein.

The structural proteins of mycobacteriophage I3 have been analysed by sodium dodecyl sulfate-polyacrylamide-gel electrophoresis (SDS-PAGE), radioiodination and immunoblotting. Based on their abundance the 34- and 70-kDa bands appeared to represent the major structural proteins. Successful cloning and expression of the 70-kDa protein-encoding gene of phage I3 in Escherichia coli and its complete nucleotide sequence determination have been accomplished. A second (partial) open reading frame following the stop codon for the 70-kDa protein was also identified within the cloned fragment. The deduced amino-acid sequence of the 70-kDa protein and the codon usage patterns indicated the preponderance of codons, as predicted from the high G+C content of the genomic DNA of phage I3.

Amino Acid Sequence↗

A bacteriophage reagent for Salmonella: molecular studies on Felix 01.

Felix 01 (F01) is a bacteriophage originally isolated by Felix and Callow which lyses almost all Salmonella strains and has been widely used as a diagnostic test for this genus. Molecular information about this phage is entirely lacking. In the present study, the DNA of the phage was found to be a double-stranded linear molecule of about 80 kb. 11.5 kb has been sequenced and in this region A + T content is 60%. There are relatively few restriction endonuclease cleavage sites in the native genome and clones show this is due to their absence rather than modification. A restriction map of the genome has been constructed. The ends of the molecule cannot be ligated although they contain 5' phosphates. At least 60% of the genome must encode proteins. In the sequenced portion, many open reading frames exist and these are tightly packed together. These have been examined for homology to published proteins but only 1 to 17 shows similarity to known proteins. F01 is therefore the prototype of a new phage family. On the basis of restriction sites, codon usage and the distribution of nonsense codons in the unused reading frames, a strong case can be made for natural selection that reacts to mRNA structure and function.

Base Sequence↗

The complete mitochondrial DNA sequence of the horseshoe crab Limulus polyphemus.

We determined the complete 14,985-nt sequence of the mitochondrial DNA of the horseshoe crab Limulus polyphemus (Arthropoda: Xiphosura). This mtDNA encodes the 13 protein, 2 rRNA, and 22 tRNA genes typical for metazoans. The arrangement of these genes and about half of the sequence was reported previously; however, the sequence contained a large number of errors, which are corrected here. The two strands of Limulus mtDNA have significantly different nucleotide compositions. The strand encoding most mitochondrial proteins has 1. 25 times as many A's as T's and 2.33 times as many C's as G's. This nucleotide bias correlates with the biases in amino acid content and synonymous codon usage in proteins encoded by different strands and with the number of non-Watson-Crick base pairs in the stem regions of encoded tRNAs. The sizes of most mitochondrial protein genes in Limulus are either identical to or slightly smaller than those of their Drosophila counterparts. The usage of the initiation and termination codons in these genes seems to follow patterns that are conserved among most arthropod and some other metazoan mitochondrial genomes. The noncoding region of Limulus mtDNA contains a potential stem-loop structure, and we found a similar structure in the noncoding region of the published mtDNA of the prostriate tick Ixodes hexagonus. A simulation study was designed to evaluate the significance of these secondary structures; it revealed that they are statistically significant. No significant, comparable structure can be identified for the metastriate ticks Rhipicephalus sanguineus and Boophilus microplus. The latter two animals also share a mitochondrial gene rearrangement and an unusual structure of mt-tRNA(C) that is exactly the same association of changes as previously reported for a group of lizards. This suggests that the changes observed are not independent and that the stem-loop structure found in the noncoding regions of Limulus and Ixodes mtDNA may play the same role as that between trnN and trnC in vertebrates, i.e., the role of lagging strand origin of replication.

Animals↗

Amino acid translation program for full-length cDNA sequences with frameshift errors.

Here we present an amino acid translation program designed to suggest the position of experimental frameshift errors and predict amino acid sequences for full-length cDNA sequences having phred scores. Our program generates artificial insertions into artificial deletions from low-accuracy positions of the original sequence, thereby generating many candidate sequences. The validity of the most probable sequence (the likelihood that it represents the actual protein) is evaluated by using a score (V(a)) that is calculated in light of the Kozak consensus, preferred codon usage, and position of the initiation codon. To evaluate the software, we have used a database in which, out of 612 cDNA sequences, 524 (86%) carried 773 frameshift errors in the coding sequence. Our software detected and corrected 48% of the total frameshift errors in 62% of the total cDNA sequences with frameshift errors. The false positive rate of frameshift correction was 9%, and 91% of the suggested frameshifts were true.

Base Composition↗

Molecular evolutionary analysis of a histone gene repeating unit from Drosophila simulans.

A repeating unit of the histone gene cluster from Drosophila simulans containing the H1, H2A, H2B and H4 genes (the H3 gene region has already been analyzed) was cloned and analyzed. A nucleotide sequence of about 4.6 kbp was determined to study the nucleotide divergence and molecular evolution of the histone gene cluster. Comparison of the structure and nucleotide sequence with those of Drosophila melanogaster showed that the four histone genes were located at identical positions and in the same directions. The proportion of different nucleotide sites was 6.3% in total. The amino acid sequence of H1 was divergent, with a 5.1% difference. However, no amino acid change has been observed for the other three histone proteins. Analysis of the GC contents and the base substitution patterns in the two lineages, D. melanogaster and D. simulans, with a common ancestor showed the following. 1) A strong negative correlation was found between the GC content and the nucleotide divergence in the whole repeating unit. 2) The mode of molecular evolution previously found for the H3 gene was also observed for the whole repeating unit of histone genes; the nucleotide substitutions were stationary in the 3' and spacer regions, and there was a directional change of the codon usage to the AT-rich codons. 3) No distinct difference in the mode or pattern of molecular evolution was detected for the histone gene repeating unit in the D. melanogaster and D. simulans lineages. These results suggest that selectional pressure for the coding regions of histones, which eliminate A and T, is less effective in the D. melanogaster and D. simulans lineages than in the other GC-rich species.

Amino Acid Sequence↗

The complete nucleotide sequence of region 1 of the CFA/I fimbrial operon of human enterotoxigenic Escherichia coli.

The production of the plasmid-encoded fimbrial antigen CFA/I of enterotoxigenic Escherichia coli requires two DNA regions: CFA/I region 1 and CFA/I region 2. These two regions are separated by about 40 kb on the wildtype plasmid. CFA/I region 1 contains the structural genes, whereas CFA/I region 2 contains a positive regulator. The first two genes (cfaA and cfaB) and the cfaD' sequence of region 1 have already been described. Here the total nucleotide sequence of region 1 is presented. Two new genes in region 1 are described, named cfaC and cfaE. The GC content of the genes in region 1 is 33.6% which is substantially lower than normally found in E. coli genes (50%). The codon usage also differs from the standard codons used in E. coli.

Amino Acid Sequence↗

Characterization of the conserved region of the mxaF gene that encodes the large subunit of methanol dehydrogenase from a marine methylotrophic bacterium.

The highly conserved region of the mxaF gene that encodes the large subunit of methanol dehydrogenase (MDH) was cloned and sequenced from Methylophaga sp. strain MP cells. The calculated G + C content of the conserved region was found to be 44.9%. The nucleotide sequence homology of the region to those from methylotrophs was approximately 43.5%, while the identity of the deduced amino acid sequence to other MxaF peptides was approximately 76.8%. Analysis of the codon usage revealed that UUC and CGU codons seem to be used only for phenylalanine and arginine, respectively. The aligned amino acid sequences show that several key amino acids that are required for the MDH enzyme activity are located in the deduced MxaF peptide, together with tryptophan-docking motifs, called W4 and W5.

Alcohol Oxidoreductases↗

Automated Machine Learning Tools to Build Regression Models for Schizosaccharomyces pombe Omics Data.

Machine learning is a powerful tool for analyzing biological data and making useful predictions. The surge of biological data from high-throughput omics technologies has raised the need for modeling approaches capable of tackling such amounts of data, which is pivotal to understanding the nature of complex molecular systems. Here, we show how to construct a simple model using automated machine learning (AutoML) to predict protein abundance in Schizosaccharomyces pombe, using data obtained from codon usage bias and quantitative proteomics.

Machine Learning↗

Hurdles to horizontal gene transfer: species-specific effects of synonymous variation and plasmid copy number determine antibiotic resistance phenotype.

Could codon composition condition the immediate success and the orientation of horizontal gene transfer? Horizontal gene transfer represents a change in the genome of expression of the transferred gene, and experimental evidence has accumulated indicating that the codon composition of a sequence is an important determinant of its compatibility with the translation machinery of the genome in which it is expressed. This suggests that codon composition influences the phenotype and the fitness conferred by a transferred gene and thus the immediate success of the transfer. To directly test this hypothesis, we characterized the resistance conferred by synonymous variants of a gentamicin resistance gene in three bacterial species: Escherichia coli, Acinetobacter baylyi and Pseudomonas aeruginosa. The strongest determinant of the resistance level conferred was the species in which the resistance gene was transferred, very likely because of important differences in the copy number of the plasmid carrying the gene. Significant differences in resistance were also found between synonymous variants within each of the three species, but more importantly, there was a strong interaction between species and variant: variants conferring high resistance in one species confer low resistance in another. However, the similarity in codon usage between the synonymous variants and the host genome only explained part of the phenotypic differences between variants in one species, P. aeruginosa. Further investigation of alternative explanations did not reveal common universal mechanisms across our three bacterial species. We conclude that codon composition can be a determinant of post-horizontal gene transfer success. However, there are multiple paths leading from synonymous sequence to phenotype, and sensitivity to these different paths is species-specific.

Gene Transfer, Horizontal↗

Negative effect of sequential serine codons on expression of foreign genes in Escherichia coli.

Herpes simplex virus encodes a 1298-residue protein designated ICP4 that regulates transcription of viral genes. Structural and functional analyses of ICP4 have been facilitated by production of portions of ICP4 in Escherichia coli. We previously observed that expression of most truncated forms of ICP4 in E. coli was relatively efficient, with the exception of portions of the ICP4 gene approximately between codons 160 and 220. We have now localized the portion of ICP4 that inhibits expression to a serine-rich region from position 176 to 199. Our experimental results suggest that codons within the serine-rich domain do not induce termination of transcription, do not alter the intrinsic stability of mRNA, and do not create a proteolytically sensitive site in this portion of ICP4. Silent mutations that alter codon usage of many of the 19 serine codons in this region had no effect on expression. However, we observed that the level of protein expression was inversely proportional to the number of serine codons in this region. The results are consistent with a model in which the serine-rich domain induces premature termination of translation. This effect is not due to any specific secondary structure in the mRNA or lack of sufficient seryl-tRNA synthetase. It remains to be determined whether premature termination can result from insufficient seryl-charged tRNAs. Our results suggest that foreign genes with more than 20 consecutive serine codons may be poorly expressed in E. coli.

Amino Acid Sequence↗

Expected frequencies of codon use as a function of mutation rates and codon fitnesses.

A method is shown to determine the expected pattern of codon use for any given set of mutation rates between nucleotides and any set of fitnesses for the codons. If it is assumed that mutations to stop codons are lethal then those codons which can mutate in one step to a stop codon tend to be used less frequently. This tendency is however, a very small one and is not likely to be observable within a single gene. Nor is it necessarily a general tendency. For example, the leucine pretermination codons may be used preferentially when mutations to proline are deleterious. It is shown that different mutation rates (eg: transitions occurring more frequently than transversions) may have as large an effect on codon usage as would strong selection for particular codons. For the model presented, an increase in the rate of transitions strongly decreases the expected frequency of UGG and CRR codons. Other codes are moderately affected by such a change in the mutation rates. Many other models can be examined using this method.

Amino Acids↗

scsB, a cDNA encoding the hydrogenosomal beta subunit of succinyl-CoA synthetase from the anaerobic fungus Neocallimastix frontalis.

A clone containing a Neocallimastix frontalis cDNA assumed to encode the beta subunit of succinyl-CoA synthetase (SCSB) was identified by sequence homology with prokaryotic and eukaryotic counter-parts. An open reading frame of 1311 bp was found. The deduced 437 amino acid sequence showed a high degree of identity to the beta-succinyl-CoA synthetase of Escherichia coli (46%), the mitochondrial beta-succinyl-CoA synthetase from pig (48%) and the hydrogenosomal beta-succinyl-CoA synthetase from Trichomonas vaginalis (49%). The G + C content of the succinyl-CoA synthetase coding sequence (43.8%) was considerably higher than that of the 5' (14.8%) and 3' (13.3%) non-translated flanking sequences, as has been observed for other genes from N. frontalis. The codon usage pattern was biased, with only 34 codons used and a strong preference for a pyrimidine (T) in the third positions of the codons. The coding sequence of the beta-succinyl-CoA synthetase cDNA was cloned in an E. coli expression vector encoding a 6(His) tag. The recombinant protein was purified by affinity binding and used to produce polyclonal antibodies. The anti-succinyl-CoA synthetase serum recognized a 45 kDa protein from a N. frontalis fraction enriched for hydrogenosomes and similar polypeptides in two related anaerobic fungi, Piromyces rhizinflata (45 kDa) and Caecomyces communis (47 kDa). Immunocytochemical experiments suggest that succinyl-CoA synthetase is located in the hydrogenosomal matrix. Staining for SCS activity in native electrophoretic gels revealed a band with an apparent molecular weight of approximately 330 kDa. The C-terminus of the succinyl-CoA synthetase sequence was devoid of the typical targeting signals identified so far in microbody proteins, indicating that N. frontalis uses a different signal for sorting SCSB into hydrogenosomes. Based on comparisons with other proteins we propose a putative N-terminal targeting signal for succinyl-CoA synthetase of N. frontalis that shows some of the features of mitochondrial targeting sequences.

Amino Acid Sequence↗

A phylogenomic study of the OCTase genes in Pseudomonas syringae pathovars: the horizontal transfer of the argK-tox cluster and the evolutionary history of OCTase genes on their genomes.

Phytopathogenic Pseudomonas syringae is subdivided into about 50 pathovars due to their conspicuous differentiation with regard to pathogenicity. Based on the results of a phylogenetic analysis of four genes (gyrB, rpoD, hrpL, and hrpS), Sawada et al. (1999) showed that the ancestor of P. syringae had diverged into at least three monophyletic groups during its evolution. Physical maps of the genomes of representative strains of these three groups were constructed, which revealed that each strain had five rrn operons which existed on one circular genome. The fact that the structure and size of genomes vary greatly depending on the pathovar shows that P. syringae genomes are quite rich in plasticity and that they have undergone large-scale genomic rearrangements. Analyses of the codon usage and the GC content at the codon third position, in conjunction with phylogenomic analyses, showed that the gene cluster involved in phaseolotoxin synthesis (argK-tox cluster) expanded its distribution by conducting horizontal transfer onto the genomes of two P. syringae pathovars (pv. actinidiae and pv. phaseolicola) from bacterial species distantly related to P. syringae and that its acquisition was quite recent (i.e., after the ancestor of P. syringae diverged into the respective pathovars). Furthermore, the results of a detailed analysis of argK [an anabolic ornithine carbamoyltransferase (anabolic OCTase) gene], which is present within the argK-tox cluster, revealed the plausible process of generation of an unusual composition of the OCTase genes on the genomes of these two phaseolotoxin-producing pathovars: a catabolic OCTase gene (equivalent to the orthologue of arcB of P. aeruginosa) and an anabolic OCTase gene (argF), which must have been formed by gene duplication, have first been present on the genome of the ancestor of P. syringae; the catabolic OCTase gene has been deleted; the ancestor has diverged into the respective pathovars; the foreign-originated argK-tox cluster has horizontally transferred onto the genomes of pv. actinidiae and pv. phaseolicola; and hence two copies of only the anabolic OCTase genes (argK and argF) came to exist on the genomes of these two pathovars. Thus, the horizontal gene transfer and the genomic rearrangement were proven to have played an important role in the pathogenic differentiation and diversification of P. syringae.

Evolution, Molecular↗