Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codons”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

CaLMPhosKAN: prediction of general phosphorylation sites in proteins via fusion of codon aware embeddings with amino acid aware embeddings and wavelet-based Kolmogorov-Arnold network.

MOTIVATION: The mapping from codon to amino acid is surjective due to codon degeneracy, suggesting that codon space might harbor higher information content. Embeddings from the codon language model have recently demonstrated success in various protein downstream tasks. However, predictive models for residue-level tasks such as phosphorylation sites, arguably the most studied Post-Translational Modification (PTM), and PTM sites prediction in general, have predominantly relied on representations in amino acid space. RESULTS: We introduce a novel approach for predicting phosphorylation sites by utilizing codon-level information through embeddings from the codon adaptation language model (CaLM), trained on protein-coding DNA sequences. Protein sequences are first reverse-translated into reliable coding sequences by mapping UniProt sequences to their corresponding NCBI reference sequences and extracting the exact coding sequences from their GenBank format using a dynamic programming-based global pairwise alignment. The resulting coding sequences are encoded using the CaLM encoder to generate codon-aware embeddings, which are subsequently integrated with amino acid-aware embeddings obtained from a protein language model, through an early fusion strategy. Next, a window-level representation of the site of interest, retaining the full sequence context, is constructed from the fused embeddings. A ConvBiGRU network extracts feature maps that capture spatiotemporal correlations between proximal residues within the window. This is followed by a prediction head based on a Kolmogorov-Arnold network (KAN) using the derivative of gaussian wavelet transform to generate the inference for the site. The overall model, dubbed CaLMPhosKAN, performs better than the existing approaches across multiple datasets. AVAILABILITY AND IMPLEMENTATION: CaLMPhosKAN is publicly available at https://github.com/KCLabMTU/CaLMPhosKAN.

Codon↗

Absence of p53 mutation at codon 249 in duck hepatocellular carcinomas from the high incidence area of Qidong (China).

Dietary aflatoxin and hepatitis B virus infection may play a role in generating the p53 tumor suppressor gene codon 249 hotspot mutation found in human hepatocellular carcinomas (HCCs) from Qidong (China) and southern Africa. No data are available on the HCC site-specific mutation of the p53 gene in hepadnavirus-infected animals exposed to AFB1. We have searched for the presence of p53 gene codon 249 mutations in both duck hepatitis B virus (DHBV) positive and negative HCCs of domestic ducks from Qidong, where the human p53 hotspot is so prevalent, as well as in duck HCCs experimentally induced by AFB1. Direct sequencing of DNA amplification products encompassing p53 codon 249 did not reveal any mutations in 11 HCCs from Qidong ducks, regardless of the status of DHBV infection. In addition no mutation was detected in four HCCs from AFB1-treated ducks. This contrasts with the human data; however, in humans, the mutation and the preferential binding of AFB1 to codon 249 occurs at the third nucleotide G, while in duck, the codon 249 lacks this G residue. The DNA sequence of adjacent codons is also different in the two species even though the amino acid sequence is identical. This may explain the low frequency of mutation we have observed. In addition, species differences in metabolism and DNA repair could influence the occurrence of codon 249 mutations.

Aflatoxin B1↗

The effects of mutation and natural selection on codon bias in the genes of Drosophila.

Codon bias varies widely among the loci of Drosophila melanogaster, and some of this diversity has been explained by variation in the strength of natural selection. A study of correlations between intron and coding region base composition shows that variation in mutation pattern also contributes to codon bias variation. This finding is corroborated by an analysis of variance (ANOVA), which shows a tendency for introns from the same gene to be similar in base composition. The strength of base composition correlations between introns and codon third positions is greater for genes with low codon bias than for genes with high codon bias. This pattern can be explained by an overwhelming effect of natural selection, relative to mutation, in highly biased loci. In particular, this correlation is absent when examining fourfold degenerate sites of highly biased genes. In general, it appears that selection acts more strongly in choosing among fourfold degenerate codons than among twofold degenerate codons. Although the results indicate regional variation in mutational bias, no evidence is found for large scale regions of compositional homogeneity.

Animals↗

Maximizing transcription efficiency causes codon usage bias.

The rate of protein synthesis depends on both the rate of initiation of translation and the rate of elongation of the peptide chain. The rate of initiation depends on the encountering rate between ribosomes and mRNA; this rate in turn depends on the concentration of ribosomes and mRNA. Thus, patterns of codon usage that increase transcriptional efficiency should increase mRNA concentration, which in turn would increase the initiation rate and the rate of protein synthesis. An optimality model of the transcriptional process is presented with the prediction that the most frequently used ribonucleotide at the third codon sites in mRNA molecules should be the same as the most abundant ribonucleotide at the third codon sites in mRNA molecules should be the same as the most abundant ribonucleotide in the cellular matrix where mRNA is transcribed. This prediction is supported by four kinds of evidence. First, A-ending codons are the most frequently used synonymous codons in mitochondria, where ATP is much more abundant than that of the three other ribonucleotides. Second, A-ending codons are more frequently used in mitochondrial genes than in nuclear genes. Third, protein genes from organisms with a high metabolic rate use more A-ending codons and have higher A content in their introns than those from organisms with a low metabolic rate.

Animals↗

Codon usage and gene expression.

The hypothesis that codon usage regulates gene expression at the level of translation is tested. Codon usage of Escherichia coli and phage lambda is compared by correspondence analysis, and the basis of this hypothesis is examined by connecting codon and tRNA distributions to polypeptide elongation kinetics. Both approaches indicate that if codon usage was random tRNA limitation would only affect the rarest tRNA species. General discrimination against their cognate codons indicates that polypeptide elongation rates are maintained constant. Thus, differences in expression of E. coli genes are not a consequence of their variable codon usage. The preference of codons recognized by the most abundant tRNAs in E. coli genes encoding abundant proteins is explained by a constraint on the cost of proof-reading.

Bacteriophage lambda↗

Synonymous codon usage in Bacillus subtilis reflects both translational selection and mutational biases.

Codon usage data for 56 Bacillus subtilis genes show that synonymous codon usage in B. subtilis is less biased than in Escherichia coli, or in Saccharomyces cerevisiae. Nevertheless, certain genes with a high codon bias can be identified by correspondence analysis, and also by various indices of codon bias. These genes are very highly expressed, and a general trend (a decrease) in codon bias across genes seems to correspond to decreasing expression level. This, then, may be a general phenomenon in unicellular organisms. The unusually small effect of translational selection on the pattern of codon usage in lowly expressed genes in B. subtilis yields similar dinucleotide frequencies among different codon positions, and on complementary strands. These patterns could arise through selection on DNA structure, but more probably are largely determined by mutation. This prevalence of mutational bias could lead to difficulties in assessing whether open reading frames encode proteins.

Bacillus subtilis↗

Codon usage in Pseudomonas aeruginosa.

We have generated a codon usage table for Pseudomonas aeruginosa. Codon usage in P. aeruginosa is extremely biased. In contrast to E. coli and yeast, P. aeruginosa preferentially uses those codons within a synonymous codon group with the strongest predicted codon-anticodon interaction. We were unable to correlate a particular codon usage pattern with predicted levels of mRNA expressivity. The choice of a third base reflects the high guanine plus cytosine content of the P. aeruginosa genome (67.2%) and cytosine is the preferred nucleotide for the third codon position.

Bacteriophages↗

Initiation of translation at an AUA codon for an archaebacterial protein gene expressed in E.coli.

Overexpression of the Sulfolobus solfataricus L12 ribosomal protein gene in E.coli cells yielded two products of different size. If the E.coli cells carrying the overexpression plasmid were induced in the early stage of bacterial growth, the smaller of the two products was almost exclusively produced. However, induction in a late stage of bacterial growth yielded the larger product in significant excess. The larger protein was identified as the translation product of the entire SsoL12 gene, while the smaller product was a N-terminally shortened version of the L12 protein (sh-SsoL12), starting with a N-terminal methionine at position 22 of the coded protein and continuing with the predicted protein sequence. Position 22 is an isoleucine in the complete SsoL12 protein sequence, coded by an AUA codon. A subclone (SsoL12**) of the SsoL12 gene containing overexpression plasmid, lacking the regular AUG start codon and the putative Shine Dalgarno sequence, was constructed to determine if E.coli ribosomes could initiate at this AUA codon. During overexpression the SsoL12** construct yielded exclusively the sh-SsoL12 product in significant amounts. An AUA start codon has never been found before in a natural message. However, experiments utilizing site directed mutagenesis to generate AUA start codons showed that this codon can be functional for initiation in prokaryotes and eukaryotes. The findings presented in this paper show that AUA acts as an initiation codon in a natural message expressed in a heterologous organism.

Amino Acid Sequence↗

Evolution of codon usage patterns: the extent and nature of divergence between Candida albicans and Saccharomyces cerevisiae.

Codon usage in a sample of 28 genes from the pathogenic yeast Candida albicans has been analysed using multivariate statistical analysis. A major trend among genes, correlated with gene expression level, was identified. We have focussed on the extent and nature of divergence between C.albicans and the closely related yeast Saccharomyces cerevisiae. It was recently suggested that significant differences exist between the subsets of preferred codons in these two species [Brown et al. (1991) Nucleic Acids Res. 19, 4293]. Overall, the genes of C.albicans are more A + T-rich, reflecting the lower genomic G + C content of that species, and presumably resulting from a different pattern of mutational bias. However, in both species highly expressed genes preferentially use the same subset of 'optimal' codons. A suggestion that the low frequency of NCG codons in both yeast species results from selection against the presence of codons that are potentially highly mutable is discounted. Codon usage in C.albicans, as in other unicellular species, can be interpreted as the result of a balance between the processes of mutational bias and translational selection. Codon usage in two related Candida species, C.maltosa and C.tropicalis, is briefly discussed.

Biological Evolution↗

Codon usage in Caenorhabditis elegans: delineation of translational selection and mutational biases.

Synonymous codon usage varies considerably among Caenorhabditis elegans genes. Multivariate statistical analyses reveal a single major trend among genes. At one end of the trend lie genes with relatively unbiased codon usage. These genes appear to be lowly expressed, and their patterns of codon usage are consistent with mutational biases influenced by the neighbouring nucleotide. At the other extreme lie genes with extremely biased codon usage. These genes appear to be highly expressed, and their codon usage seems to have been shaped by selection favouring a limited number of translationally optimal codons. Thus, the frequency of these optimal codons in a gene appears to be correlated with the level of gene expression, and may be a useful indicator in the case of genes (or open reading frames) whose expression levels (or even function) are unknown. A second, relatively minor trend among genes is correlated with the frequency of G at synonymously variable sites. It is not yet clear whether this trend reflects variation in base composition (or mutational biases) among regions of the C.elegans genome, or some other factor. Sequence divergence between C.elegans and C.briggsae has also been studied.

Animals↗

Mutations to nonsense codons in human genetic disease: implications for gene therapy by nonsense suppressor tRNAs.

Nonsense suppressor tRNAs have been suggested as potential agents for human somatic gene therapy. Recent work from this laboratory has described significant effects of 3' codon context on the efficiency of human nonsense suppressors. A rapid increase in the number of reports of human diseases caused by nonsense codons, prompted us to determine how the spectrum of mutation to either UAG, UAA or UGA codons and their respective 3' contexts, might effect the efficiency of human suppressor tRNAs employed for purposes of gene therapy. This paper presents a survey of 179 events of mutations to nonsense codons which cause human germline or somatic disease. The analysis revealed a ratio of approximately 1:2:3 for mutation to UAA, UAG and UGA respectively. This pattern is similar, but not identical, to that of naturally occurring stop codons. The 3' contexts of new mutations to stop were also analysed. Once again, the pattern was similar to the contexts surrounding natural termination signals. These results imply there will be little difference in the sensitivity of nonsense mutations and natural stop codons to suppression by nonsense suppressor tRNAs. Analysis of the codons altered by nonsense mutations suggests that efforts to design human UAG suppressor tRNAs charged with Trp, Gln, and Glu; UAA suppressors charged with Gln and Glu, and UGA suppressors which insert Arg, would be an essential step in the development of suppressor tRNAs as agents of human somatic gene therapy.

Codon↗

The CUG codon is decoded in vivo as serine and not leucine in Candida albicans.

Previous studies have shown that the yeast Candida albicans encodes a unique seryl-tRNA(CAG) that should decode the leucine codon CUG as serine. However, in vitro translation of several different CUG-containing mRNAs in the presence of this unusual seryl-tRNA(CAG) result in an apparent increase in the molecular weight of the encoded polypeptides as judged by SDS-PAGE even though the molecular weight of serine is lower than that of leucine. A possible explanation for this altered electrophoretic mobility is that the CUG codon is decoded as modified serine in vitro. To elucidate the nature of CUG decoding in vivo, a reporter system based on the C. albicans gene (RBP1) encoding rapamycin-binding protein (RBP), coupled to the promoter of the C. albicans TEF3 gene, was utilized. Sequencing and mass-spectrometry analysis of the recombinant RBP expressed in C. albicans demonstrated that the CUG codon was decoded exclusively as serine while the related CUU codon was translated as leucine. A database search revealed that 32 out of the 65 C. albicans gene sequences available have CUG codons in their open reading frames. The CUG-containing genes do not belong to any particular gene family. Thus the amino acid specified by the CUG codon has been reassigned within the mRNAs of C. albicans. We argue here that this unique genetic code change in cellular mRNAs cannot be explained by the 'Codon Reassignment Theory'.

Amino Acid Sequence↗

Exploiting unassigned codons in Micrococcus luteus for tRNA-based amino acid mutagenesis.

An alternative to suppression of stop codons for the biosynthetic insertion of non-natural amino acids has been developed. Micrococcus luteus , a Gram-positive bacterium, is incapable of translating at least two codons. One of these unused codons was inserted in a gene to act as a nonsense site. An aminoacylated tRNA was synthesized which was complementary to this codon. The gene containing the missing codon was expressed in vitro in a M.luteus transcription/translation system. Read-through of the missing codon occurred only when the complementary tRNA was included. The results demonstrate that M.luteus can be used for incorporation of amino acids via synthetically prepared aminoacylated tRNAs. The use of a M. luteus translation system provides a method for incorporation of non-natural amino acids which avoids the use of stop codons.

Amino Acid Sequence↗

Rates of synonymous substitution do not indicate selective constraints on the codon use of the plant psbA gene.

The psbA gene of the flowering plant chloroplast genome has a pattern of codon bias that differs from all other angiosperm chloroplast genes. In psbA, unlike all other chloroplast genes, the third-codon-position composition does not reflect the general genome compositional bias of a high A+T content. Instead, in specific synonymous groups, the codon use of psbA more closely corresponds to the tRNA population available for translation. Since it requires a composition unlike the genome composition bias, this pattern of codon use is likely to be the result of selection. Selective constraints on codon use are expected to result in decreased rates of synonymous substitution, and it has been observed that psbA has the lowest rate of synonymous substitution among plant chloroplast genes. In the present study, this is examined further by testing whether or not those synonymous groups that specifically have an atypical codon use in psbA have correspondingly low rates of silent substitution. An analysis of synonymous substitution rate, performed separately for different degeneracy classes of amino acids, shows that, contrary to the expectation, those sites that are presumed, based on their codon use, to be under selective constraint in psbA do not show low rates of substitution and are not responsible for the overall low rate of synonymous substitution in this gene. Instead, they actually show increased rates of substitution relative to other chloroplast genes. Two hypotheses concerning the role of selection in psbA are advanced to explain the results.

Chloroplasts↗

A codon-based model designed to describe lentiviral evolution.

A codon-based model designed to describe lentiviral evolution is developed. The model incorporates unequal base compositions in the three codon positions and selection against the CpG dinucleotide within codons to account for a deficit of this dinucleotide exhibited by lentiviral genes. The model is, to a large extent, able to account for the pattern of codon usage exhibited by the HIV1 genes gag, pol, and env, in spite of its parameter paucity. The model is extended to a similar model which operates on pentets (codons and their neighboring bases). The results obtained by the pentet model establish the importance of depression of CpGs across codon boundaries as well as within codons. The goodness of fit of the CpG depression model to the observed evolution in pairwise alignments of HIV1 sequences is assessed. The model provides a significantly better description of the observed evolution than the simpler models examined. The parameter estimates indicate that part of the unusually large biases in nucleotide frequencies observed in HIV1 genes is caused by selection against CpGs. We find that the estimates of expected numbers of substitutions, of transitions to transversions, and of synonymous to nonsynonymous substitution rates are robust to CpG depression, whereas the ratio of CpG-generating substitutions to other substitutions is strongly influenced by the choice of model.

Base Composition↗

Selection at the wobble position of codons read by the same tRNA in Saccharomyces cerevisiae.

The transfer RNA gene complement of Saccharomyces cerevisiae was utilized for a whole-genome analysis of the deviation from a neutral usage of pyrimidine-ending cognate codons, that is, codons read by a single tRNA species having either inosine or guanosine as the first anticodon base. Mutational pressure at the wobble position was estimated from the base composition of the noncoding portion of the yeast genome. The selective pressure for translational efficiency was inferred from the degree of codon adaptation to tRNA gene redundancy and from mRNA abundance data derived from yeast transcriptome analysis. Amino acid conservation in orthologous comparisons with wholly sequenced microbial genomes was used to estimate translational accuracy requirements. A close correspondence was observed between the usage of wobble position pyrimidines and the frequency predicted by mutational bias. However, in the case of four cognate pairs (Gly: ggu/ggc; Asn: aau/aac; Phe: uuu/uuc; Tyr: uau/ uac) all read by guanosine-starting anticodons, we found evidence for a strong selective pressure driven by translational efficiency. Only for the glycine pair, wobble pyrimidine choice also appears to fulfill a translational accuracy requirement. Wobble pyrimidine selection is strictly related to the number of hydrogen bonds formed by alternative cognate codons: whenever a different number of hydrogen bonds can be formed at the wobble position, there is selection against six- or nine-hydrogen-bonded codon-anticodon pairs. Our results indicate that an intrinsic codon preference, critically dependent on the stability of codon-anticodon interaction and mainly reflecting selection for the optimization of translational efficiency, is built into the translational apparatus.

Codon↗

Codon usage and tRNA content in unicellular and multicellular organisms.

Choices of synonymous codons in unicellular organisms are here reviewed, and differences in synonymous codon usages between Escherichia coli and the yeast Saccharomyces cerevisiae are attributed to differences in the actual populations of isoaccepting tRNAs. There exists a strong positive correlation between codon usage and tRNA content in both organisms, and the extent of this correlation relates to the protein production levels of individual genes. Codon-choice patterns are believed to have been well conserved during the course of evolution. Examination of silent substitutions and tRNA populations in Enterobacteriaceae revealed that the evolutionary constraint imposed by tRNA content on codon usage decelerated rather than accelerated the silent-substitution rate, at least insofar as pairs of taxonomically related organisms were examined. Codon-choice patterns of multicellular organisms are briefly reviewed, and diversity in G+C percentage at the third position of codons in vertebrate genes--as well as a possible causative factor in the production of this diversity--is discussed.

Animals↗

Natural selection versus primitive gene structure as determinant of codon usage.

Different codons are not utilized equally in known gene sequences. One of the important biases of codon usage is observed in the form of an enrichment of RNY codons, especially within RNN codon families. Such biases could represent the residue of a primitive repeating-RNY gene structure, or the outcome of natural selection, or both. Analyses based on the rates of silent substitutions, the frequencies of base doublets, and synonymous codon ratios for Escherichia coli, yeast, Drosophila and Xenopus proteins have been performed. The results rule out any significant support for a primitive repeating-RNY or repeating-RRY gene structure, and establish the important role of natural selection in determining the choice of codons. With strong intervention by natural selection, the relationship between primitive gene structure and codon usage necessarily becomes minimal.

Animals↗