Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codons”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Nucleotides of 18S rRNA surrounding mRNA codons at the human ribosomal A, P, and E sites: a crosslinking study with mRNA analogs carrying an aryl azide group at either the uracil or the guanine residue.

The 18S rRNA environment of the mRNA at the decoding site of human 80S ribosomes has been studied by cross-linking with derivatives of hexaribonucleotide UUUGUU (comprising Phe and Val codons) that carried a perfluorophenylazide group either at the N7 atom of the guanine or at the C5 atom of the 5'-terminal uracil residue. The location of the codons on the ribosome at A, P, or E sites has been adjusted by the cognate tRNAs. Three types of complexes have been obtained for each type derivative, namely, (1) codon UUU and Phe-tRNAPhe at the P site (codon GUU at the A site), (2) codon UUU and tRNAPhe at the P site and PheVal-tRNAVal at the A site, and (3) codon GUU and Val-tRNAVal at the P site (codon UUU at the E site). This allowed the placement of modified nucleotides of the mRNA analog at positions -3, +1, or +4 on the ribosome. Mild UV irradiation resulted in tRNA-dependent crosslinking of the mRNA analogs to the 18S rRNA. Nucleotide G961 crosslinked to mRNA position -3, nucleotide G1207 to position +1, and A1823 together with A1824 to position +4. All of these nucleotides are located in the most strongly conserved regions of the small subunit RNA structure, and correspond to nucleotides G693, G926, G1491, and A1492 of bacterial 16S rRNA. Three of them (with the exception of G1491) had been found earlier at the 70S ribosomal decoding site. The similarities and differences between the 16S and 18S rRNA decoding sites are discussed.

Anticodon↗

Codon 201Arg/Gly polymorphism of DCC (deleted in colorectal carcinoma) gene in flat- and polypoid-type colorectal tumors.

Recent studies have identified the distinct existence of flat-type colorectal tumors. The low incidence of ras gene mutations in these tumors suggests that their genetic pathways of tumor progression may be different from those of the polypoid type. To elucidate further genetic alterations in flat-type colorectal tumors, codon 201Arg/Gly polymorphism in the DCC (deleted in colorectal carcinoma) gene was analyzed in normal tissue (normal colonic mucosa or peripheral lymphocytes) and in tumor tissue from 191 patients with colorectal tumors (36 patients with flat-type colorectal tumors, 81 patients with polypoid-type colorectal tumors, and 74 patients with advanced carcinomas). For normal controls, 30 samples obtained from patients who had neither colorectal tumors (confirmed by total colonoscopy) nor a family history of colorectal carcinoma were analyzed. DCC gene codon 201Arg/Gly polymorphism was investigated by polymerase chain reaction-based restriction fragment length polymorphism analysis, fluorescence-based dideoxy sequencing, or both. For the flat type, the frequency of codon 201Gly of the DCC gene was 64% and 54% in the normal tissue of patients with adenoma with high-grade dysplasia and submucosal carcinoma, respectively. It was 49%, 52%, and 49% in the normal tissue of patients with polypoid-type adenoma with high-grade dysplasia, submucosal carcinoma, and advanced carcinoma, respectively. In the normal tissue, codon 201Gly of the DCC gene was more frequently observed in patients with flat-type adenoma with low-grade dysplasia (67%) than in those with polypoid-type adenoma with low-grade dysplasia (18%) or in normal controls (17%, P < 0.05, chi2 test). Codon 201Arg/Gly polymorphism in tumor tissues did not differ from that in the corresponding normal tissues, except for 10 cases of carcinoma with loss of heterozygosity (LOH). In carcinomas with LOH, preferential loss of the codon 201Arg allele was noted (9/10 cases). These results suggest that codon 201Gly of the DCC gene is not only associated with flat-type colorectal tumors, but that it may serve as a useful genetic marker for identifying groups at higher risk for colorectal cancer.

Carcinoma↗

Detection of the DCC gene product in normal and malignant colorectal tissues and its relation to a codon 201 mutation.

Protein expression of the putative tumour-suppressor gene DCC on chromosome 18q was evaluated in a panel of 16 matched colorectal cancer and normal colonic tissue samples together with DCC mRNA expression and allelic deletions (loss of heterozygosity, LOH). Determined by a polymerase chain reaction (PCR)-LOH assay, 12 of the 16 (75%) cases were informative with LOH occurring in 2 of the 12 cases. For DCC mRNA, transcripts could be detected in all analysed normal tissues (eight out of eight) by RT-PCR, whereas 6 of the 15 tumours were negative. DCC protein expression, investigated by immunohistochemistry using the monoclonal antibody 15041 A directed against the intracellular domain, was homogeneously positive in all normal tissue samples. In tumour tissues, no DCC protein was seen in 11 out of 16 samples (69%). For the DCC codon 201, we found a loss of a wild-type codon sequence caused by mutation or LOH in at least 8 out of 15 cases (53%) compared with the corresponding normal tissue. DCC protein expression was undetectable in eight of the nine tumours missing both wild-type codons. Only one of the five tumours with retained DCC protein expression had no detectable wild-type codon 201. In addition, 9 out of 15 normal tissue specimens were mutated in codon 201. In two out of three cases with homozygous wild-type codons in peripheral blood lymphocyte (PBL) DNA, mutations were already observed in the tumour adjacent normal colonic mucosa. We conclude that DCC immunostaining should be introduced in the clinicopathological routine because of its strong correlation with the known prognostic markers 18q LOH and mutation of codon 201.

Blotting, Southern↗

Expanding the amino acid repertoire of ribosomal polypeptide synthesis via the artificial division of codon boxes.

In ribosomal polypeptide synthesis the library of amino acid building blocks is limited by the manner in which codons are used. Of the proteinogenic amino acids, 18 are coded for by multiple codons and therefore many of the 61 sense codons can be considered redundant. Here we report a method to reduce the redundancy of codons by artificially dividing codon boxes to create vacant codons that can then be reassigned to non-proteinogenic amino acids and thereby expand the library of genetically encoded amino acids. To achieve this, we reconstituted a cell-free translation system with 32 in vitro transcripts of transfer RNASNN (tRNASNN) (S = G or C), assigning the initiator and 20 elongator amino acids. Reassignment of three redundant codons was achieved by replacing redundant tRNASNNs with tRNASNNs pre-charged with non-proteinogenic amino acids. As a demonstration, we expressed a 32-mer linear peptide that consists of 20 proteinogenic and three non-proteinogenic amino acids, and a 14-mer macrocyclic peptide that contains more than four non-proteinogenic amino acids.

Amino Acid Sequence↗

Does recombination improve selection on codon usage? Lessons from nematode and fly complete genomes.

Understanding the factors responsible for variations in mutation patterns and selection efficacy along chromosomes is a prerequisite for deciphering genome sequences. Population genetics models predict a positive correlation between the efficacy of selection at a given locus and the local rate of recombination because of Hill-Robertson effects. Codon usage is considered one of the most striking examples that support this prediction at the molecular level. In a wide range of species including Caenorhabditis elegans and Drosophila melanogaster, codon usage is essentially shaped by selection acting for translational efficiency. Codon usage bias correlates positively with recombination rate in Drosophila, apparently supporting the hypothesis that selection on codon usage is improved by recombination. Here we present an exhaustive analysis of codon usage in C. elegans and D. melanogaster complete genomes. We show that in both genomes there is a positive correlation between recombination rate and the frequency of optimal codons. However, we demonstrate that in both species, this effect is due to a mutational bias toward G and C bases in regions of high recombination rate, possibly as a direct consequence of the recombination process. The correlation between codon usage bias and recombination rate in these species appears to be essentially determined by recombination-dependent mutational patterns, rather than selective effects. This result highlights that it is necessary to take into account the mutagenic effect of recombination to understand the evolutionary role and impact of recombination.

Animals↗

Allosteric mechanism for codon-dependent tRNA selection on ribosomes.

We suggest that the interaction between a codon and its cognate tRNA induces conformational changes in the tRNA. We further suggest that sites on the ribosome preferentially bind tRNA in those conformations which require proper matching of codon and anticodon. According to this model, the codon functions as an allosteric effector which influences the conformation at various sites in the tRNA. This is made possible by the ribosome, which we suggest traps tRNA molecules in those conformation states that maximize the energy difference between cognate and noncognate codon-anticodon interactions. Studies of the interactions between tRNA molecules and their cognate codons in the absence of the ribosome have suggested that triplet-triplet interaction between codon and anticodon is far too weak to account for the specificity of the tRNA selection mechanism during protein synthesis. In contrast, we suggest that such affinity measurements do not adequately describe the interaction between a codon and its cognate tRNA. Thus, such experiments can not detect conformational changes in the tRNA, and, in particular, those stabilized by the ribosome.

Allosteric Regulation↗

Direct mapping of adeno-associated virus capsid proteins B and C: a possible ACG initiation codon.

The three major capsid proteins of adeno-associated virus type 2 (AAV2) virions are designated A, B, and C and have molecular sizes of 90, 72, and 60 kDa, respectively. These proteins are related, and genetic studies have shown they are encoded by a long open reading frame located in the right half of the genome. The coding capacity distal to the first ATG in this reading frame is only 503 amino acids (i.e., a protein about the size of protein C), but an open frame sequence devoid of ATG codons extends upstream for an additional 184 codons. Although the amino terminus of the C capsid protein is blocked, partial amino acid sequence analyses of peptides from C have confirmed that it is encoded within the portion of the reading frame distal to the first ATG at nucleotide (nt) location 2810. The amino terminus of the B capsid protein is not blocked, and its sequence begins with alanine. The triplet encoding this alanine lies 64 codons upstream from the initiation site for protein C and is immediately preceded by the threonine codon, ACG, at nt 2615. This ACG codon lies in the most favorable sequence context for protein synthesis initiation. All three AAV2 capsid proteins are labeled in vitro with formyl[35S]methionyl-tRNAf, indicating that synthesis of each protein is initiated independently. Our data suggest that the nt 2615 ACG codon directs the methionyl-tRNA-dependent initiation of the AAV2 B capsid protein. Proteins B and C may be synthesized from the same mRNA species and their relative abundance could be determined by the efficiencies of their respective initiation codons.

Amino Acid Sequence↗

Transcription attenuation in Salmonella typhimurium: the significance of rare leucine codons in the leu leader.

The leucine operon of Salmonella typhimurium is controlled by a transcription attenuation mechanism. Four adjacent leucine codons within a 160-nucleotide leu leader RNA are thought to play a central role in this mechanism. Three of the four codons are CUA, a rarely used leucine codon within enteric bacteria. To determine whether the nature of the leucine codon affects the regulation of the leucine operon, we used oligonucleotide-directed mutagenesis to first convert one CUA of the leader to CUG and then convert all three CUA codons to CUG. CUG is the most frequently used leucine codon in enteric bacteria. A mutant having (CUA)2CUGCUC in place of (CUA)3CUC has an altered response to leucine limitation, requiring a slightly higher degree of limitation to effect derepression. Changing (CUA)3CUC to (CUG)3CUC has more dramatic effects upon operon expression. First, the basal level of expression is lowered to the point that the mutant grows more slowly than the parent in a minimal medium lacking leucine. Second, the response of the mutant to a leucine limitation is dramatically altered such that even a strong limitation elicits only a modest degree of derepression. If the mutant is grown under conditions of leucyl-tRNA limitation rather than leucine limitation, complete derepression can be achieved, but only at a much higher degree of limitation than for the wild-type operon. These results provide a clear-cut example of codon usage having a dramatic effect upon gene expression.

Codon↗

Nonrandom utilization of codon pairs in Escherichia coli.

We have analyzed protein-coding sequences of Escherichia coli and find that codon-pair utilization is highly biased, reflecting overrepresentation or underrepresentation of many pairs compared with their random expectations. This effect is over and above that contributed by nonrandomness in the use of amino acid pairs, which itself is highly evident; it is much weaker when nonadjacent codon pairs are examined and virtually disappears when pairs separated by two or three intervening codons are evaluated. There appears to be a high degree of directionality in this bias: any codon that participates in many nonrandom pairs tends to make both over- and underrepresented pairs, but preferentially as a left- or right-hand member. We show a relationship between codon-pair utilization patterns and levels of gene expression: genes encoding proteins expressed at high levels tend to contain more abundant, but more highly underrepresented, codon pairs, relative to genes expressed at low levels. The nonrandom utilization of codon pairs may be a consequence of their effects on translational efficiency, which in turn may be related to the compatibility of adjacent aminoacyl-tRNA isoacceptors at the A and P sites of a translating ribosome.

Bacterial Proteins↗

Codon discrimination and anticodon structural context.

Site-directed mutagenesis has been used to change the nucleotide C in the wobble position of tRNA(1Gly) (CCC) to U. The mutated tRNA was tested for its ability to read glycine codons in an in vitro protein-synthesizing system programmed with the phage message MS2-RNA that had been modified by site-directed mutagenesis so as to make it possible to monitor conveniently the reading of all four glycine codons. The results showed that while the efficiency of tRNA(1Gly) (UCC) was comparable to that of mycoplasma tRNA(Gly) (UCC) in the reading of the codon GGA, the mycoplasma tRNA(Gly) was far more efficient than the tRNA(1Gly) (UCC) in the reading of the codons GGU and GGC. Thus, the anticodon UCC, when present in the structural context of the tRNA(1Gly) molecule, behaved as predicted by the wobble rules while in the structural context of the mycoplasma tRNA(Gly) it read without discrimination between the nucleotides in the third codon position, in violation of the wobble restrictions. The result with the codon GGG showed that the anticodon UCC, when present in tRNA(1Gly), was considerably less efficient in reading this codon than it was in the structural context of the mycoplasma tRNA(Gly). It would therefore seem that the anticodon UCC, when present in a certain tRNA, can be an efficient wobbler, while in the molecular environment of another tRNA it is markedly restricted in its ability to wobble.

Amino Acid Sequence↗

Features of the formate dehydrogenase mRNA necessary for decoding of the UGA codon as selenocysteine.

The fdhF gene encoding the 80-kDa selenopolypeptide subunit of formate dehydrogenase H from Escherichia coli contains an in-frame TGA codon at amino acid position 140, which encodes selenocysteine. We have analyzed how this UGA "sense codon" is discriminated from a UGA codon signaling polypeptide chain termination. Deletions were introduced from the 3' side into the fdhF gene and the truncated 5' segments were fused in-frame to the lacZ reporter gene. Efficient read-through of the UGA codon, as measured by beta-galactosidase activity and incorporation of selenium, was dependent on the presence of at least 40 bases of fdhF mRNA downstream of the UGA codon. There was excellent correlation between the results of the deletion studies and the existence of a putative stem-loop structure lying immediately downstream of the UGA in that deletions extending into the helix drastically reduced UGA translation. Similar secondary structures can be formed in the mRNAs coding for other selenoproteins. Selenocysteine insertion cartridges were synthesized that contained this hairpin structure and variable portions of the fdhF gene upstream of the UGA codon and inserted into the lacZ gene. Expression studies showed that upstream sequences were not required for selenocysteine insertion but that they may be involved in modulating the efficiency of read-through. Translation of the UGA codon was found to occur with high fidelity since it was refractory to ribosomal mutations affecting proofreading and to suppression by the sup-9 gene product.

Aldehyde Oxidoreductases↗

Ribosomal decoding processes at codons in the A or P sites depend differently on 2'-OH groups.

The importance of 2'-OH groups of codons for binding of cognate tRNAs to ribosomal P and A sites was analyzed applying the following strategy. An mRNA of 41 nucleotides was synthesized with the structure C16-GAA-UUC-GUC-C16 coding for glutamic acid (E), phenylalanine (F) and valine (V), respectively, in the middle (EFV-mRNA). A second template, the E(dF)V-mRNA, was identical except that it carried a deoxyribo-codon-dUdUdC- for phenylalanine. tRNA binding to the P site is totally insensitive to the presence or absence of the 2'-OH group of the P-site codon, and tRNA binding to the P site is also not affected if the A-site codon lacks the 2'-OH groups. However, binding is impaired if the deoxy-codon is present at the E site. In sharp contrast, the A-site binding of Ac-aminoacyl-tRNA was severely reduced in the presence of the deoxy-codon at the A site as well as at the P site. The results demonstrate that the correctness of base pairing is also "sensed" via a correct sugar structure of the codon, e.g. positioning of the sugar pucker (2'-OH), during the decoding process at the A site (elongation) but not during the decoding at the P site (initiation).

Base Sequence↗

Analysis of the codon bias in E. coli sequences.

Fifty-three gene sequences from E. coli containing 18,288 reading frame triplets have been characterized according to the nature and level of average codon preference. The distribution of average preferences is bimodal, with approximately half the genes using an average of only 36 codons, and the remainder just 42 codons. There is a high correlation between the level of codon bias, the tRNA population and the abundance of protein product, indicating biased patterns are exploited by the cell for the production of widely different levels of gene product. This relationship is especially striking in genes involved in the production of components for transcription and translation. Overall, the genes for these processes generate some five-fold more protein than the average in the genome, and use about five fewer codons. The very high codon bias found in the RNA polymerase gene thus provides a simple, autogenous mechanism for the coordinate synthesis of these components and RNA polymerase. A surprisingly high level of codon probability is also found in triplets of the complement of coding sequences. This is apparently due to the evolutionary dispersion of coding sequences and/or the requirement for increased levels of secondary structure in messenger RNAs.

Base Sequence↗

CaLMPhosKAN: prediction of general phosphorylation sites in proteins via fusion of codon aware embeddings with amino acid aware embeddings and wavelet-based Kolmogorov-Arnold network.

MOTIVATION: The mapping from codon to amino acid is surjective due to codon degeneracy, suggesting that codon space might harbor higher information content. Embeddings from the codon language model have recently demonstrated success in various protein downstream tasks. However, predictive models for residue-level tasks such as phosphorylation sites, arguably the most studied Post-Translational Modification (PTM), and PTM sites prediction in general, have predominantly relied on representations in amino acid space. RESULTS: We introduce a novel approach for predicting phosphorylation sites by utilizing codon-level information through embeddings from the codon adaptation language model (CaLM), trained on protein-coding DNA sequences. Protein sequences are first reverse-translated into reliable coding sequences by mapping UniProt sequences to their corresponding NCBI reference sequences and extracting the exact coding sequences from their GenBank format using a dynamic programming-based global pairwise alignment. The resulting coding sequences are encoded using the CaLM encoder to generate codon-aware embeddings, which are subsequently integrated with amino acid-aware embeddings obtained from a protein language model, through an early fusion strategy. Next, a window-level representation of the site of interest, retaining the full sequence context, is constructed from the fused embeddings. A ConvBiGRU network extracts feature maps that capture spatiotemporal correlations between proximal residues within the window. This is followed by a prediction head based on a Kolmogorov-Arnold network (KAN) using the derivative of gaussian wavelet transform to generate the inference for the site. The overall model, dubbed CaLMPhosKAN, performs better than the existing approaches across multiple datasets. AVAILABILITY AND IMPLEMENTATION: CaLMPhosKAN is publicly available at https://github.com/KCLabMTU/CaLMPhosKAN.

Codon↗

Absence of p53 mutation at codon 249 in duck hepatocellular carcinomas from the high incidence area of Qidong (China).

Dietary aflatoxin and hepatitis B virus infection may play a role in generating the p53 tumor suppressor gene codon 249 hotspot mutation found in human hepatocellular carcinomas (HCCs) from Qidong (China) and southern Africa. No data are available on the HCC site-specific mutation of the p53 gene in hepadnavirus-infected animals exposed to AFB1. We have searched for the presence of p53 gene codon 249 mutations in both duck hepatitis B virus (DHBV) positive and negative HCCs of domestic ducks from Qidong, where the human p53 hotspot is so prevalent, as well as in duck HCCs experimentally induced by AFB1. Direct sequencing of DNA amplification products encompassing p53 codon 249 did not reveal any mutations in 11 HCCs from Qidong ducks, regardless of the status of DHBV infection. In addition no mutation was detected in four HCCs from AFB1-treated ducks. This contrasts with the human data; however, in humans, the mutation and the preferential binding of AFB1 to codon 249 occurs at the third nucleotide G, while in duck, the codon 249 lacks this G residue. The DNA sequence of adjacent codons is also different in the two species even though the amino acid sequence is identical. This may explain the low frequency of mutation we have observed. In addition, species differences in metabolism and DNA repair could influence the occurrence of codon 249 mutations.

Aflatoxin B1↗

The effects of mutation and natural selection on codon bias in the genes of Drosophila.

Codon bias varies widely among the loci of Drosophila melanogaster, and some of this diversity has been explained by variation in the strength of natural selection. A study of correlations between intron and coding region base composition shows that variation in mutation pattern also contributes to codon bias variation. This finding is corroborated by an analysis of variance (ANOVA), which shows a tendency for introns from the same gene to be similar in base composition. The strength of base composition correlations between introns and codon third positions is greater for genes with low codon bias than for genes with high codon bias. This pattern can be explained by an overwhelming effect of natural selection, relative to mutation, in highly biased loci. In particular, this correlation is absent when examining fourfold degenerate sites of highly biased genes. In general, it appears that selection acts more strongly in choosing among fourfold degenerate codons than among twofold degenerate codons. Although the results indicate regional variation in mutational bias, no evidence is found for large scale regions of compositional homogeneity.

Animals↗

Maximizing transcription efficiency causes codon usage bias.

The rate of protein synthesis depends on both the rate of initiation of translation and the rate of elongation of the peptide chain. The rate of initiation depends on the encountering rate between ribosomes and mRNA; this rate in turn depends on the concentration of ribosomes and mRNA. Thus, patterns of codon usage that increase transcriptional efficiency should increase mRNA concentration, which in turn would increase the initiation rate and the rate of protein synthesis. An optimality model of the transcriptional process is presented with the prediction that the most frequently used ribonucleotide at the third codon sites in mRNA molecules should be the same as the most abundant ribonucleotide at the third codon sites in mRNA molecules should be the same as the most abundant ribonucleotide in the cellular matrix where mRNA is transcribed. This prediction is supported by four kinds of evidence. First, A-ending codons are the most frequently used synonymous codons in mitochondria, where ATP is much more abundant than that of the three other ribonucleotides. Second, A-ending codons are more frequently used in mitochondrial genes than in nuclear genes. Third, protein genes from organisms with a high metabolic rate use more A-ending codons and have higher A content in their introns than those from organisms with a low metabolic rate.

Animals↗

Codon usage and gene expression.

The hypothesis that codon usage regulates gene expression at the level of translation is tested. Codon usage of Escherichia coli and phage lambda is compared by correspondence analysis, and the basis of this hypothesis is examined by connecting codon and tRNA distributions to polypeptide elongation kinetics. Both approaches indicate that if codon usage was random tRNA limitation would only affect the rarest tRNA species. General discrimination against their cognate codons indicates that polypeptide elongation rates are maintained constant. Thus, differences in expression of E. coli genes are not a consequence of their variable codon usage. The preference of codons recognized by the most abundant tRNAs in E. coli genes encoding abundant proteins is explained by a constraint on the cost of proof-reading.

Bacteriophage lambda↗