On the origin and evolution of the genetic code. II. Origin of the genetic code as a primordial collector language. The pairing-release hypothesis.
Explore the source record for details and available documents.
SEARCH · Search PubMed
Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
It is known that in the transcription of genetic information there are ambiguities, i.e. the fact that a triplet codes for several aminoacids. This has generally been taken as due to errors in the transcription mechanisms. However, it has been postulated that instead of accidental miscoding, ambiguities are part of the expression of a generalized genetic code which depends on biological context. Group theory has been used to find the generalized genetic code. Here we present a new group theory approach which we think removes some weaknesses of previous works. The generalized genetic code presented here is different from that previously reported. We compare our results with experimental evidence and discuss the predictions presented.
The genetic code doublets can be divided into two octets of completely degenerate and ambiguous coding dinucleotides. These two octets have the algebraic property of lying on continuously connected planes on the group graph (a tesseract) of the Cartesian product of two Klein 4-groups of nucleotide exchange operators. The K X K group can also be broken into four cosets, one of which has completely degenerate coding elements, and another that has completely ambiguous coding elements. The two octets of coding doublets have the further algebraic property that the product of their internal exchange operators naturally divide into two exactly equivalent sets. These properties of the genetic code are relevant to unraveling error-detecting and error-correcting (proof-reading) aspects of the genetic code and may be helpful in understanding the context-sensitive grammar of genetic language.
The genetic code has been influenced by directional mutation pressure affecting the base composition of DNA, sometimes in the direction of increased GC content and at other times, in the direction of AT. Such pressure led to changes in species-specific usages of codons and tRNA anticodons, and also in amino acid assignments of codons in mitochondria and in several intact organisms. These code changes are probably recent evolutionary events. The genetic code is not 'frozen', but instead it is still evolving.
The genetic code, which directs the protein biosynthesis, is an information system. Although all its details are not known at present, its essential characteristics are elucidated, as well for the replication or transcription as for the translation of the genetic message. A coherent picture now appears, which reveals the existence of an universal structure, the most fundamental features of which seem to obey some logic. A systematic approach has been devised, which aims to their integration in a theorectical scheme: many features of the code table can thus be interpreted as resulting from a unique principle of best resistance against the effects of mutations. Any group of triplets or amino-acids can be considered along this line. It is more difficult however, to analyse the coexistence of two (or more) different groups. In this work, we propose to extend our optimization principle into a more general one, which includes the notion of information as defined by Shannon. We explore some consequences of this new principle in the most simple models that one can build for the origin and evolution of the genetic code.
The genetic code is an instantaneous code i.e., each codon is deciphered out of ambiguities without knowing other symbols than the constituting nucleotides. Moreover entropy of the genetic source of information has a maximal value.
The genetic code is characterized by hidden symmetry. Amino acids possessing common antiamino acids are located symmetrically in the graphic models of the code. There is only one exception--apolar amino acids V, M, I, L and F are asymmetrically arranged. Asymmetric disposition of these amino acids is apparently due to divergence in the course of structural evolution of amino acid families as a result of inclusion of new members into the coding system.
The genetic code determines not only the amino acid sequences of proteins but also mRNA stability. How is this hidden message read? Hia and colleagues have now identified human DHX29 as a reader of the mRNA stability code carried by codons, providing new mechanistic insights into translation-coupled gene regulation.
The genetic code, formerly thought to be frozen, is now known to be in a state of evolution. This was first shown in 1979 by Barrell et al. (G. Barrell, A. T. Bankier, and J. Drouin, Nature [London] 282:189-194, 1979), who found that the universal codons AUA (isoleucine) and UGA (stop) coded for methionine and tryptophan, respectively, in human mitochondria. Subsequent studies have shown that UGA codes for tryptophan in Mycoplasma spp. and in all nonplant mitochondria that have been examined. Universal stop codons UAA and UAG code for glutamine in ciliated protozoa (except Euplotes octacarinatus) and in a green alga, Acetabularia. E. octacarinatus uses UAA for stop and UGA for cysteine. Candida species, which are yeasts, use CUG (leucine) for serine. Other departures from the universal code, all in nonplant mitochondria, are CUN (leucine) for threonine (in yeasts), AAA (lysine) for asparagine (in platyhelminths and echinoderms), UAA (stop) for tyrosine (in planaria), and AGR (arginine) for serine (in several animal orders) and for stop (in vertebrates). We propose that the changes are typically preceded by loss of a codon from all coding sequences in an organism or organelle, often as a result of directional mutation pressure, accompanied by loss of the tRNA that translates the codon. The codon reappears later by conversion of another codon and emergence of a tRNA that translates the reappeared codon with a different assignment. Changes in release factors also contribute to these revised assignments. We also discuss the use of UGA (stop) as a selenocysteine codon and the early history of the code.
The universally valid genetic code is the final result of a multi-stage course of development. Degeneracy, as an important property of the genetic code, was possibly not yet present in the earliest code, first appearing at a later stage of development (Code III). Possibly this step in development is coupled with the presence of a total of four amino acid groups (L, I, E, F). Each group contains a specific number of amino acid (AL, AI, AE, AF). Amino acid groups: - (L) hydrophobic - (I) weakly hydrophobic or polar but uncharged - (E) hydrophilic, acidic - (F) hydrophilic, basic - (D) hydrophobic, aromatic (only in Code IV and Code M. This group is not considered in the calculations below.) In a subsequent stage of development the number of amino acids increases further. At the same time the code becomes more degenerate. The universal genetic code is characterized by three constants of being degenerate. Its immediate predecessor has linear degeneration with two constants. The mitochondrial code represents a transitional form between these two codes.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The chemical language of genetic code is proposed. As a result of chemical language application for the analysis of the modern genetic code, the existence of an unambiguous correspondence between the chemical properties of amino acids and their coding triplets (codons and anticodons) is shown. This confirms the hypothesis of the code chemical determination. The complementarity between the chemical properties of amino acids and their anticodons (but not the codons) has been found also to exist. This observation supports the hypothesis of the genetic code determination by the direct recognition and also underlines the primary role of anticodon in the origin of genetic code in comparison with codons.
The prokaryotic genetic code has been influenced by directional mutation pressure (GC/AT pressure) that has been exerted on the entire genome. This pressure affects the synonymous codon choice, the amino acid composition of proteins and tRNA anticodons. Unassigned codons would have been produced in bacteria with extremely high GC or AT genomes by deleting certain codons and the corresponding tRNAs. A high AT pressure together with genomic economization led to a change in assignment of the UGA codon, from stop to tryptophan, in Mycoplasma.
The structure of the genetic code suggests that amino acid biosynthesis and hydrophobicity were important factors in shaping the genetic code, as the primitive code coevolved with new varieties of amino acids generated by the expanding pathways of biosynthesis. The current code is exceptionally stable. Deviant codes nonetheless have been observed in a number of mitochondrial and cellular genomes. Even the membership of encoded amino acids is undergoing expansion to include phosphoserine and selenocysteine. Experimental mutation of the code also has proven feasible, in a replacement of tryptophan by 4-fluorotryptophan as a component constituent of proteins. Such mutations, introducing novel varieties of encoded amino acids, will open up a new dimension in protein engineering and design.
The genetic code of mammalian mitochondria presents some symmetric in the sequences of certain codons and in the relationships existing between the codons and aminoacids. These symmetries suggest the classification of codons in four classes and indicate that the evolution of the genetic code developed in four stages.
Chemical language of the genetic code is suggested in which elementary information code units are presented by functional groups of amino acids and nucleotides. Using this language, the existence of correspondence and conformity of chemical parameters of amino acids and of central nucleotides of their anticodons was demonstrated. These findings confirm the idea that the genetic code is determined by chemical properties of amino acids and nucleotides and that this determination is the result of direct specific interactions between amino acids and nucleotide triplets at the stage of the origin of the code. The data obtained reveal primary role of anticodon triplets in the origin of the code. Key role of the central nucleotide in triplets for amino acid coding is confirmed.
The genetic code is evolving as shown by 9 departures from the universal code: 6 of them are in mitochondria and 3 are in nuclear codes. We propose that these changes are preceded by disappearance of a codon from coding sequences in mRNA of an organism or organelle. The function of the codon that disappears is taken by other, synonymous codons, so that there is no change in amino acid sequences of proteins. The deleted codon then reappears with a new function. Wobble pairing between anticodons and codons has evolved, starting with a single UNN anticodon pairing with 4 codons. Directional mutation pressure affects codon usage and may produce codon reassignments, especially of stop codons. Selenocysteine is coded by UGA, which is also a stop codon, and this anomaly is discussed. The outlook for discovery of more changes in the code is favorable, and open reading frames should be compared with actual sequential analyses of protein molecules in this search.