Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Codons”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Naturally occurring splicing variants of the hMSH2 gene containing nonsense codons identify possible mRNA instability motifs within the gene coding region.

We have identified certain unusually spliced cDNA species following PCR amplification of peripheral blood lymphocyte (PBL) mRNA from the hMSH2 gene. A naturally occurring transcript containing a nonsense codon due to the skipping of 5 exons was amplified from PBLs of several healthy individuals. A feature of this and another unusual splicing product was the presence of sequence motifs which bore significant similarity to mRNA instability determinants in the region immediately downstream of the stop codon. In particular, the rare tetranucleotide GAUG, previously identified in yeast as being of critical importance to the rapid degradation of nonsense-containing mRNAs was situated 23 base pairs downstream of the stop codon. Furthermore the region downstream of the stop codon was A:U rich and contained 2 copies of the AUUUA motif. As other forms of alternative splicing would not result in the same juxtaposition of stop codons and instability motifs, we suggest that the stop codons may have been deliberately introduced by the splicing process for their proximity to these destabilising motifs, and that splicing may play a role in channeling mRNAs into degradative pathways. These results are consistent with the hypothesis that nuclear factors may scan pre-mRNAs prior to splicing.

Adenosine Triphosphatases↗

Codon usage in selected AT-rich bacteria.

The relationship between DNA base composition and codon bias in very AT-rich bacteria was analyzed. Five clostridial genes, five mycoplasmal genes and three rickettsial genes constituted the data base. In the genes of these three organisms, the rule for codon bias was very simple: use U or A in the first and third positions of the codon when possible. This was contrasted with the bias found in Bacillus subtilis and Escherichia coli. The rule for Bacillus subtilis was equally straightforward: use all codons without bias. Only in E. coli, amongst the species examined, did the codon bias appear to be a complicated codon 'choice'.

Adenine↗

Preferential codon usage in genes.

We present a method which permits comparison of the preferential use of degenerate codons within any gene. The method makes use of the triplet frequencies in the noncoding frames to assess whether a preference is specific to the reading frame. Preference is given a statistical meaning by use of the analysis of variance coupled to Duncan's multiple range test. Preferential use of degenerate codons is gene-specific and independent of gene size. The data suggest that any correlation between codon frequency distribution and tRNA levels is unreliable. In those animal genes examined, codons ending in C or G are preferred; in animal viruses tested, codons ending in U or A are preferred. Similarly, the bacterial genes and the genes of single-stranded DNA phages that we analyzed differed from each other as well as from eukaryotic genes in the third base of the codon.

Animals↗

Codon usage in the G+C-rich Streptomyces genome.

The codon usage (CU) patterns of 64 genes from the Gram+ prokaryotic genus Streptomyces were analysed. Despite the extremely high overall G+C content of the Streptomyces genome (estimated at 0.74), individual genes varied in G+C content from 0.610 to 0.797, and had third codon position G+C contents (GC3s) that varied from 0.764 to 0.983. The variation in GC3s explains a significant proportion of the variation in CU patterns. This is consistent with an evolutionary model of the Streptomyces genome where biased mutation pressure has led to a high average G+C content with random variation about the mean, although the variation observed is greater than that expected from a simple binomial model. The only gene in the sample that can be confidently predicted to be highly expressed, EF-Tu of Streptomyces coelicolor A3(2) (GC3s = 0.927), shows a preference for a third position C in several of the four codon families, and for CGY and GGY for Arg and Gly codons, respectively (Y = pyrimidine); similar CU patterns are found in highly expressed genes of the G+C-rich Micrococcus luteus genome. It thus appears that codon usage in Streptomyces is determined predominantly by mutation bias, with weak translational selection operating only in highly expressed genes. We discuss the possible consequences of the extreme codon bias of Streptomyces and consider how it may have evolved. A set of CU tables is provided for use with computer programs that locate protein-coding regions.

Base Composition↗

Effects of rare codon clusters on high-level expression of heterologous proteins in Escherichia coli.

Within Escherichia coli and other species, a clear codon bias exists among the 61 amino acid codons found within the population of mRNA molecules, and the level of cognate tRNA appears directly proportional to the frequency of codon usage. Given this situation, one would predict translational problems with an abundant mRNA species containing an excess of rare low tRNA codons. Such a situation might arise after the initiation of transcription of a cloned heterologous gene in the E. coli host. Recent studies suggest clusters of AGG/AGA, CUA, AUA, CGA or CCC codons can reduce both the quantity and quality of the synthesized protein. In addition, it is likely that an excess of any of these codons, even without clusters, could create translational problems.

Arginine↗

A common periodic table of codons and amino acids.

A periodic table of codons has been designed where the codons are in regular locations. The table has four fields (16 places in each) one with each of the four nucleotides (A, U, G, C) in the central codon position. Thus, AAA (lysine), UUU (phenylalanine), GGG (glycine), and CCC (proline) were placed into the corners of the fields as the main codons (and amino acids) of the fields. They were connected to each other by six axes. The resulting nucleic acid periodic table showed perfect axial symmetry for codons. The corresponding amino acid table also displaced periodicity regarding the biochemical properties (charge and hydropathy) of the 20 amino acids and the position of the stop signals. The table emphasizes the importance of the central nucleotide in the codons and predicts that purines control the charge while pyrimidines determine the polarity of the amino acids. This prediction was experimentally tested.

Algorithms↗

Mistranslation in IGF-1 during over-expression of the protein in Escherichia coli using a synthetic gene containing low frequency codons.

Partial misincorporation of Lys for Arg has been observed for the Arg residues of IGF-1 when the molecule is expressed in Escherichia coli using a synthetic gene with the low frequency AGA codon encoding all six Arg residues and yeast preferred codons encoding the remaining residues. The Lys for Arg substitution at these residues could not be detected when a gene containing E. coli preferred codons, with the codon CGT coding for all Arg residues, was used for the expression of the protein. Similarly, no misincorporation of Lys for Arg could be detected when a gene containing Escherichia coli preferred codons at all positions, except for an AGA codon at Arg (36), was utilized.

Amino Acid Sequence↗

Codon discrimination due to presence of abundant non-cognate competitive tRNA.

It has been thought that preferential use of synonymous codons provides high efficiency and fidelity of protein synthesis through specific codon-anticodon interactions. In yeast genes, some codon boxes seem to prefer a codon which is unsuited for its cognate anticodon. Now, we propose that codon usage biases may arise due to presence of abundant non-cognate competitive tRNA capable of misreading a codon by C-U or G-U pairing in the middle position.

Codon↗

Effect of distribution of unfavourable codons on the maximum rate of gene expression by an heterologous organism.

We have analysed theoretically the effect of the relative position of unfavourable codons on the maximum level of synthesis of foreign proteins in E. coli. We predict that the occurrence of such codons scattered in the corresponding genes has little effect. In contrast, clustering (in our terminology indicating directly adjacent codons) of unfavourable codons is predicted to dramatically reduce the maximum level of protein synthesis. The context effect would explain the reduction of expression level for a chloramphenicol acetyl transferase gene modified by Robinson et al. (1984), which contains 4 contiguous unfavourable codons. As an example, we predict that due to the different downstream contexts of unfavourable codons in the alpha 1 and beta interferon genes, the maximum level of synthesis in E. coli for these proteins will be different.

Acetyltransferases↗

Effects of codon-optimization on protein expression by the human herpesvirus 6 and 7 U51 open reading frame.

Codon-optimization refers to the alteration of gene sequences, to make codon usage match the available tRNA pool within the cell/species of interest. Codon-optimization has emerged as a powerful tool to increase protein expression by genes from small RNA and DNA viruses, which commonly contain overlapping reading frames as well as structural elements that are embedded within coding regions; these features are not widespread among large DNA viruses. We therefore examined whether codon-optimization might influence protein expression from a herpesvirus gene. We focused on the U51 gene from human herpesviruses-6 and -7, which was cloned in both native and codon-optimized form, with an N-terminal HA epitope tag to allow protein detection. Codon-optimization was associated with a profound (10-100 fold) increase in U51 expression in human (293A, HSG, K562) or hamster (CHO) cell lines, suggesting this may represent a valuable tool to facilitate functional studies on recalcitrant herpesvirus genes. Finally, it is postulated that the suboptimal expression of native U51 may reflect a regulatory mechanism that controls viral gene expression.

Animals↗

Positioning of mRNA codons with respect to 18S rRNA at the P and E sites of human ribosome.

Positioning of each nucleotide of the E site and the P site bound codons with respect to the 18S rRNA on the human ribosome was studied by cross-linking with mRNA analogs, derivatives of the hexaribonucleotide UUUGUU (comprising Phe and Val codons) that carried a perfluorophenylazide group on the second or the third uracil, and a derivative of the dodecaribonucleotide UUAGUAUUUAUU with a similar group on the guanine residue. The location of the modified nucleotides at any mRNA position from -3 to +3 (position +1 corresponds to the 5' nucleotide of the P site bound codon) was adjusted by the cognate tRNAs. A modified uridine at positions from -1 to +3 cross-linked to nucleotide G1207 of the 18S rRNA, and to nucleotide G961 when it was in position -2. A modified guanosine cross-linked to nucleotide G1207 if it was in position -3 of the mRNA. These data indicate that nucleotide G961 of the 18S rRNA is close only to mRNA positions -3 and -2, while G1207 is in the vicinity of positions from -3 to +3. The latter suggests that there is a sharp turn between the P and E site bound codons that brings nucleotide G1207 of the 18S rRNA close to each nucleotide of these codons. This correlates well with X-ray crystallographic data on bacterial ribosomes, indicating existence of a sharp turn between the P site and E site bound codons near a conserved nucleotide G926 of the 16S rRNA (corresponding to G1207 in 18S rRNA) close to helix 23b containing the conserved nucleotide 693 of the 16S rRNA (corresponding exactly to G961 of the 18S rRNA).

Codon↗

Generation of protein isoform diversity by alternative initiation of translation at non-AUG codons.

The use of several translation initiation codons in a single mRNA, by expressing several proteins from a single gene, contributes to the generation of protein diversity. A small, yet growing, number of mammalian mRNAs initiate translation from a non-AUG codon, in addition to initiating at a downstream in-frame AUG codon. Translation initiation on such mRNAs results in the synthesis of proteins harbouring different amino terminal domains potentially conferring on these isoforms distinct functions. Use of non-AUG codons appears to be governed by several features, including the sequence context and the secondary structure surrounding the codon. Selection of the downstream initiation codon can occur by leaky scanning of the 43S ribosomal subunit, internal entry of ribosome or ribosomal shunting. The biological significance of non-AUG alternative initiation is demonstrated by the different subcellular localisations and/or distinct biological functions of the isoforms translated from the single mRNA as illustrated by the two main angiogenic factor genes encoding the fibroblast growth factor 2 (FGF2) and the vascular endothelial growth factor (VEGF). Consequently, the regulation of alternative initiation of translation might have a crucial role for the biological function of the gene product.

Animals↗

The translational stop signal: codon with a context, or extended factor recognition element?

Wide ranging studies of the readthrough of translational stop codons within the last 25 years have suggested that the stop codon might be only part of the molecular signature for recognition of the termination signal. Such studies do not distinguish between effects on suppression and effects on termination, and so we have used a number of different approaches to deduce whether the stop signal is a codon with a context or an extended factor recognition element. A data base of natural termination sites from a wide range of organisms (148 organisms, approximately 40,000 sequences) shows a very marked bias in the bases surrounding the stop codon in the genes for all organisms examined, with the most dramatic bias in the base following the codon (+4). The nature of this base determines the efficiency of the stop signal in vivo, and in Escherichia coli this is reinforced by overexpressing the stimulatory factor, release factor 3. Strong signals, defined by their high relative rates of selecting the decoding release factors, are enhanced whereas weak signals respond relatively poorly. Site-directed cross-linking from the +1, and bases up to +6 but not beyond make close contact with the bacterial release factor-2. The translational stop signal is deduced to be an extended factor recognition sequence with a core element, rather than simply a factor recognition triplet codon influenced by context.

Base Sequence↗

Hypothetical clues suggesting that 'AU' and 'UA' were the first 'beginning' and 'end' codons: is their ancient polymerase extant present as a protein domain fossil in reverse transcriptase and telomerase?

Tryptophan and cysteine are excluded from the histones, except histone H3, presumably to decrease premature terminations, since the first two letters of their codons are the same as that of the termination codon UGA. Though tyrosine codon also begins with 'UG', any mutation in tyrosine codon causing premature termination is corrected more professionally by the suppressor tRNA system. Given this background, a hypothetical model for the first replication is proposed. The first termination codon is 'UA' and the first initiation codon is 'AU', in 'UAU' centred RNA rings, which are replicated both clockwise and counter-clockwise. Replication of these would result in the formation of palindromic strands, which may have led to the formation of palindromic sequences in the genome, such as those forming hairpin loops. Their polymerase may be extant in reverse transcriptase (RT), because of tendencies to 'turn the corner' and to form 'hairpin loops', when RT is used for DNA synthesis in vitro. The catalytic subunit of telomerase hTERT in close contact to its RNA template may be another site to find this ancient 'polymerase'.

Codon↗

DNA sequence analysis of the complete mitochondrial genome of the green alga Scenedesmus obliquus: evidence for UAG being a leucine and UCA being a non-sense codon.

The complete DNA sequence of the mitochondrial genome of the chlorophyceen alga Scenedesmus obliquus was determined. The circular genome of 42781bp contains a basic set of 13 mitochondrial genes, which are conserved among plant or algal chondriomes. In addition, two scrambled rRNA and 27 tRNA genes are present, together with four intronic sequences (group I and II) and five open reading frames (ORFs), which show no significant homology to other ORFs from organellar genomes. The comparison with deduced amino acid sequences from 13 conserved mitochondrial genes gives rise to the conclusion that two deviations from the standard genetic code must be present in S. obliquus mitochondria: (i) UAG codes for leucine as was already found in some other algal mitochondria; (ii) UCA is a stop codon, which seems unique for mitochondrial genomes. This was supported by our finding that a tRNA-Leu gene possesses a UCA anticodon and by a missing tRNA-serine, able to decode the UCA codon. Consistent with these data is the absence of any UCA codon from conserved mitochondrial ORFs. This codon occurs only close to the end of all ORFs, while UAA or UGA codons are found at some distance from any conserved ORF. Codon changes by RNA editing can be excluded, since RT-PCR analysis does not reveal any evidence for post-transcriptional RNA modifications of the primary transcript.

Algal Proteins↗

Gene expressivity is the main factor in dictating the codon usage variation among the genes in Pseudomonas aeruginosa.

Codon usage biases of all DNA sequences (length greater than or equal to 300 bp) from the complete genome of Pseudomonas aeruginosa have been analyzed. As P. aeruginosa is a GC-rich organism, G and/or C are expected to predominate in their codons. Overall codon usage data analysis indicates that indeed codons ending in G and/or C are predominant in this organism. But multivariate statistical analysis indicates that there is a single major trend in the codon usage variation among the genes in this organism, which has a strong negative correlation with the expressivities of the genes. The majority of the lowly expressed genes are scattered towards the positive end of the major axis whereas the highly expressed genes are clustered towards the negative end. This is the first report where the prokaryotic organism having highly skewed base composition is dictated mainly by translational selection, though some other factors such as the lengths of the genes as well as the hydrophobicity of genes also influence the codon usage variation among the genes in this organism in a minor way.

Codon↗

The involucrin gene of the tree shrew: recent repeat additions and the relocation of cysteine codons.

The coding region of the involucrin gene of Tupaia glis has been cloned and sequenced. It resembles the involucrin coding region of other non-anthropoid mammals in possessing a segment of related, short tandem repeats at a defined location, but in Tupaia, there has been recent serial duplication of a repeat into which a cysteine codon had earlier been introduced. As a result of the duplication, there is a total of as many as six cysteine codons in the segment of repeats, a number larger than for any other species yet examined. In Ratttus there has been a comparable but independent addition of cysteine codons, and both Tupaia and Rattus have eliminated an otherwise conserved cysteine codon 75 located close to but outside the segment of repeats. In Tupaia, this elimination probably occurred by gene conversion. Also independently, the gene of Canis has added cysteine codons to the segment of repeats but has not yet lost cysteine 75. It is proposed that the gain and the loss of cysteine codons are parts of a multi-stage program of cysteine relocation.

Animals↗

Translation-coupled violation of Parity Rule 2 in human genes is not the cause of heterogeneity of the DNA G+C content of third codon position.

The genome of higher eukaryotes consists of genes having a widely heterogeneous base composition at the third codon position. Ubiquitous variability of the DNA base composition has the following two aspects: intragenomic heterogeneity of the G+C content and the amino-acid-specific translation-coupled biases from the Parity Rule 2 (PR2). PR2 is an intrastrand rule where A = T and G = C are expected if there is no bias in mutation and selection between the two complementary strands of DNA. To examine whether or not the biases from PR2 are responsible for the wide heterogeneity of the DNA G+C content in human, the third codon position of 846 human genes was analyzed. Genes were separated into six groups according to their G+C content of the third codon position, and each group was examined for the translation-coupled PR2 biases in the nucleotide composition of the third codon position for two- and four-codon amino acids. The results show that genes in the different G+C content groups have similar PR2 biases, indicating that the intragenomic heterogeneity of the G+C content is not correlated with translation-coupled biases from the PR2. Therefore, the heterogeneity of the G+C content is likely to be determined by some other mechanism (e.g. locally variable directional mutation pressures) than amino-acid-specific selections for the codon preference.

Base Composition↗