Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genetic code”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Relationships between genomic base content and distribution of mass in coded proteins.

The aim of this research was to examine the possible significance of genome/protein relationships in terms of effects on distribution of mass, especially in proteins. Amino acid residues in proteins have side-chains and polypeptide segments. We use "SCM" (side-chain mass), "MCM" (main-chain mass), and "deltaM" (SCM-MCM) as the deviation from "mass balance." Total MCM of the 61 amino acids in the standard code, 3412, equals total SCM: they form a mass balanced set (mean deltaM = 0). Of 14 natural variants of the code, seven have slightly positive mean deltaM values and seven have slightly negative values. Codes with the standard amino acids assigned randomly to the 20 codon sets of the standard code have about one chance in 3,300 of producing a mass balanced set. In natural proteins, as %A + T increases, the proportion of the mass in the side-chains also increases, by about half the amount calculated for standard genes with various AT/GC ratios, partly due to selection of codons with greater variability in composition at synonymous sites. For 203 representative species (including organelles), the total protein mass is distributed approximately equally between SCM and MCM (overall mean deltaM/amino acid residue, -0.06). The attainment of some overall macromolecular mass balance may have been a criterion for selecting the codon/amino acid pairs. When both structural and dynamic requirements are considered, a genetic code based on hydrophobicity and mass balance as key properties seems likely.

AT Rich Sequence↗

Inadequacy of prebiotic synthesis as origin of proteinous amino acids.

The production of some nonproteinous, and lack of production of other proteinous, amino acids in model prebiotic synthesis, along with the instability of glutamine and asparagine, suggest that not all of the 20 present day proteinous amino acids gained entry into proteins directly from the primordial soup. Instead, a process of active co-evolution of the genetic code and its constituent amino acids would have to precede the final selection of these proteinous amono acids.

Amino Acids↗

A conformational rationale for the wobble behaviour of the first base of the anticodon triplet in tRNA.

We present a conformational rationale for wobble behaviour of the first base in the anticodon triplet of tRNA and hence for the well-known degeneracy of the genetic code. The U-turn hydrogen bond plays an important role in the structure of the anticodon arm and particularly for the anticodon triplet to be in a geometry suitable for the process of recognition in the adaptor-mediated synthesis of proteins. This hydrogen bond in turn precludes a hydrogen bond between the first two sugars of the anticodon triplet, allowing the first base to wobble, while it facilitates one between the second and third sugars of the triplet, positioning these bases for the standard base-pairing with the codon. This neatly explains why there is a degeneracy in the code and why a RNA happens to be the adaptor for protein synthesis. Relevent conformational calculations are presented in support of the theory.

Anticodon↗

Unbiased estimation of symmetrical directional mutation pressure from protein-coding DNA.

The most generally applicable procedure for obtaining estimates of the symmetrical, or strandnonspecific, directional mutation pressure (microD) on protein-coding DNA sequences is to determine the G+C content at synonymous codon sites (Psyn), and to divide Psyn by twice the arithmetic mean of the G+C content at synonymous codon sites of a large number of randomly generated, synonymously coding DNA sequences (Psyn). Unfortunately, the original procedure yields biased estimates of Psyn and microD and is computationally expensive. We here present a fast procedure for estimating unbiased microD values. The procedure employs direct calculation of Psyn (approximately Psyn) and two normalization procedures, one for Psyn < or = Psyn and another for Psyn > or = Psyn. The normalization removes a bias sometimes caused by codons specifying arginine, asparagine, isoleucine, and leucine. Consequently, comparison of protein-coding genes that are translated using different genetic codes is facilitated.

Animals↗

The role of self-assembled monolayers of the purine and pyrimidine bases in the emergence of life.

The experimental evidence for the spontaneous formation and structure determination of two-dimensional monolayers of the purine and pyrimidine bases is examined. The plausibility of such structures forming spontaneously at the solid-liquid interface following their prebiotic synthesis suggests a functional role for them in the emergence of life. It is proposed that prebiotic interactions of enantiomorphic monolayers of mixed base composition with racemic amino acids might be implicated in a simultaneous origin of a primitive genetic coding mechanism and biomolecular homochirality. The interactions of these monolayers with carbohydrates and other derivatives is also discussed.

Adsorption↗

Codon-acticodon recognition in the valine codon family.

An in vitro protein-synthesizing system completely dependent on added valine tRNA (valyl-tRNAval) and programmed with RNA from the phage MS2 has been used to investigate the incorporation into MS2 coat protein of valine from isoaccepting valyl-tRNAsval with the anticodons U AC (U represents 5-oxyacetic acid uridine monophosphate), GAC, and IAC in response to the four valine codons GUU, GUC, GUA, and GUG. By examining the incorporation of valine into NH2-terminal and internal positions of three tryptic peptides from the MS2 coat protein it has been established that these anticodons each recognize all four valine codons. We therefore conclude that under our conditions of in vitro protein synthesis the genetic code, as far as the valine codons are concerned, is operationally a two letter code, i.e. the third codon nucleotide has no absolute discriminating function.

Amino Acid Sequence↗

The evolution of biased codon and amino acid usage in nematode genomes.

Despite the degeneracy of the genetic code, whereby different codons encode the same amino acid, alternative codons and amino acids are utilized nonrandomly within and between genomes. Such biases in codon and amino acid usage have been demonstrated extensively in prokaryote genomes and likely reflect a balance between the action of mutation, selection, and genetic drift. Here, we quantify the effects of selection and mutation drift as causes of codon and amino acid-usage bias in a large collection of nematode partial genomes from 37 species spanning approximately 700 Myr of evolution, as inferred from expressed sequence tag (EST) measures of gene expression and from base composition variation. Average G + C content at silent sites among these taxa ranges from 10% to 63%, and EST counts range more than 100-fold, underlying marked differences between the identities of major codons and optimal codons for a given species as well as influencing patterns of amino acid abundance among taxa. Few species in our sample demonstrate a dominant role of selection in shaping intragenomic codon-usage biases, and these are principally free living rather than parasitic nematodes. This suggests that deviations in effective population size among species, with small effective sizes among parasites, are partly responsible for species differences in the extent to which selection shapes patterns of codon usage. Nevertheless, a consensus set of optimal codons emerges that is common to most taxa, indicating that, with some notable exceptions, selection for translational efficiency and accuracy favors similar sets of codons regardless of the major codon-usage trends defined by base compositional properties of individual nematode genomes.

Amino Acids↗

Direct charging of tRNA(CUA) with pyrrolysine in vitro and in vivo.

Pyrrolysine is the 22nd amino acid. An unresolved question has been how this atypical genetically encoded residue is inserted into proteins, because all previously described naturally occurring aminoacyl-tRNA synthetases are specific for one of the 20 universally distributed amino acids. Here we establish that synthetic L-pyrrolysine is attached as a free molecule to tRNA(CUA) by PylS, an archaeal class II aminoacyl-tRNA synthetase. PylS activates pyrrolysine with ATP and ligates pyrrolysine to tRNA(CUA) in vitro in reactions specific for pyrrolysine. The addition of pyrrolysine to Escherichia coli cells expressing pylT (encoding tRNA(CUA)) and pylS results in the translation of UAG in vivo as a sense codon. This is the first example from nature of direct aminoacylation of a tRNA with a non-canonical amino acid and shows that the genetic code of E. coli can be expanded to include UAG-directed pyrrolysine incorporation into proteins.

Acylation↗

A code in the protein coding genes.

A statistical analysis with 12,288 autocorrelation functions applied in protein (coding) genes of prokaryotes and eukaryotes identifies three subsets of trinucleotides in their three frames: T0 = X0 [symbol: see text] {AAA, TTT} with X0 = {AAC, AAT, ACC, ATC, ATT, CAG, CTC, CTG, GAA, GAC, GAG, GAT, GCC, GGC, GGT, GTA, GTC, GTT, TAC, TTC} in frame 0 (the reading frame established by the ATG start trinucleotide), T1 = X1 [symbol: see text] {CCC} in frame 1 and T2 = X2 [symbol: see text] {GGG} in frame 2 (the frames 1 and 2 being the frame 0 shifted by one and two nucleotides, respectively, to the right). These three subsets are identical in these two gene populations and have five important properties: (i) the property of maximal (20 trinucleotides) circular code for X0 (resp. X1, X2) allowing to retrieve automatically the frame 0 (resp. 1, 2) in any region of the gene without start codon; (ii) the DNA complementarity property C (e.g. C(AAC) = GTT): C(T0) = T0, C(T1) = T2 and C(T2) = T1 allowing the two paired reading frames of a DNA double helix simultaneously to code for amino acids; (iii) the circular permutation property P (e.g. P(AAC) = ACA): P(X0) = X1 and P(X1) = X2 implying that the two subsets X1 and X2 can be deduced from X0; (iv) the rarity property with an occurrence probability of X0 = 6 x 10(-8); and (v) the concatenation properties in favour of an evolutionary code: a high frequency (27.5%) of misplaced trinucleotides in the shifted frames, a maximum (13 nucleotides) length of the minimal window to retrieve automatically the frame and an occurrence of the four types of nucleotides in the three trinucleotide sites. In Discussion, a simulation based on an independent mixing of the trinucleotides of T0 allows to retrieve the two subsets T1 and T2. Then, the identified subsets T0, T1 and T2 replaced in the 2-letter genetic alphabet {R, Y} (R = purine = A or G, Y = pyrimidine = C or T) allow to retrieve the RNY model (N = R or Y) and to explain previous works in the alphabet {R, Y}. Then, these three subsets are related to the genetic code. The trinucleotides of T0 code for 13 amino acids: Ala, Asn, Asp, Gln, Glu, Gly, Ile, Leu, Lys, Phe, Thr, Tyr and Val. Finally, a strong correlation between the usage of the trinucleotides of T0 in protein genes and the amino acid frequencies in proteins is observed as six among seven amino acids not coded by T0, have as expected the lowest frequencies in proteins of both prokaryotes and eukaryotes.

Animals↗

Studies on order in prebiological systems at the Laboratory of Chemical Evolution.

The basic tenet of investigations in the Laboratory of Chemical Evolution (LCE) under Cyril Ponnamperuma was that biology is a recapitulation of prebiology, and that protobiology is an outcome of simple molecular interactions, engendered by the physics and chemistry of the molecules themselves. Studies were undertaken to continue research into understanding the determining physical and chemical parameters of molecular interactions leading to increasing complexity in pre- and proto-biological systems. Among other, related work, research was performed on the origin of the genetic code, the origin of order in prebiotic polymers, and related studies on the origins of optical activity in biological macromolecules. Highlights of some these studies are presented here.

Amino Acids↗

On the appearance of function and organisation in the origin of life.

Models for the steps of organisation in the origin of life are discussed with an emphasis on stability, and the possibilities of acquiring a diversity of functions. In particular, two basic models are described: that of simple self-replicating molecules, and that of autocatalytic self-reproduction, which is accomplished by a hypercyclic organisation. The latter may be exemplified by the RNA world. The view of a step-wise development with new functions successively incorporated and a high accuracy of the reproduction from the onset is criticised. Instead, we suggest that no clear systematic information is continued to the first cell before the start of protein synthesis. A non-selective manifold of self-replicating molecules and unsystematic protein production from the beginning could have caused a very large diversity from which functions that could stabilise the system by feedback loops could be selected. The only way to stabilise the protein synthesis and the genetic code would be to have feedback mechanisms so that the code actually produced the proteins that supported that very code. The code would then become frozen. As DNA would require control functions, it would not be used as a single information-carrier until protein synthesis had been established and the functions were available.

Animals↗

Transfer RNA recognition by aminoacyl-tRNA synthetases.

The aminoacyl-tRNA synthetases are an ancient group of enzymes that catalyze the covalent attachment of an amino acid to its cognate transfer RNA. The question of specificity, that is, how each synthetase selects the correct individual or isoacceptor set of tRNAs for each amino acid, has been referred to as the second genetic code. A wealth of structural, biochemical, and genetic data on this subject has accumulated over the past 40 years. Although there are now crystal structures of sixteen of the twenty synthetases from various species, there are only a few high resolution structures of synthetases complexed with cognate tRNAs. Here we review briefly the structural information available for synthetases, and focus on the structural features of tRNA that may be used for recognition. Finally, we explore in detail the insights into specific recognition gained from classical and atomic group mutagenesis experiments performed with tRNAs, tRNA fragments, and small RNAs mimicking portions of tRNAs.

Amino Acyl-tRNA Synthetases↗

The cytochrome oxidase subunit I gene of Tetrahymena: a 57 amino acid NH2-terminal extension and a 108 amino acid insert.

The gene sequence for cytochrome oxidase subunit I (COI) in the ciliate Tetrahymena mitochondrial DNA has been determined and shown to be coded by the same strand as codes the genes (in order) for 14S rRNA, tRNA(trp), tRNA(glu), 21S rRNA, tRNA(leu) and tRNA(met). The predicted protein has 698 amino acids, including an NH2-terminal 57 amino acid extension and a 108 amino acid insert originally found in Paramecium COI. These extension and insert segments are not highly hydrophobic but are relatively rich in lysine, arginine and serine. In analogy with the presequence of nuclear-encoded mitochondrial proteins, they might function as a transmembrane signal. The remaining polypeptide segments show a hydrophobicity characteristic of membrane spanning proteins. TCOI shows a 64% amino acid identity with Paramecium COI but less than a 38% amino acid conservation with human COI. The Tetrahymena mitochondrial code is analogous with the mammalian mitochondrial code; but differs from the Tetrahymena nuclear genetic code; TGA is exclusively translated as tryptophan; ATA is used as an initiation codon probably for methionine, and TAA as a stop codon; the arginine codons (CGN) are not used. The use of the leucine codon TTA in TCOI is contradictory to the codon recognition pattern previously obtained from the isolated tRNA(leu) isoacceptors recognizing only the CUN codons, but consistent with the tRNA(leu) (anticodon UAA) gene encoded in the genome. The reason for this inconsistency has not been resolved.

Amino Acid Sequence↗

Selection of the simplest RNA that binds isoleucine.

We have identified the simplest RNA binding site for isoleucine using selection-amplification (SELEX), by shrinking the size of the randomized region until affinity selection is extinguished. Such a protocol can be useful because selection does not necessarily make the simplest active motif most prominent, as is often assumed. We find an isoleucine binding site that behaves exactly as predicted for the site that requires fewest nucleotides. This UAUU motif (16 highly conserved positions; 27 total), is also the most abundant site in successful selections on short random tracts. The UAUU site, now isolated independently at least 63 times, is a small asymmetric internal loop. Conserved loop sequences include isoleucine codon and anticodon triplets, whose nucleotides are required for amino acid binding. This reproducible association between isoleucine and its coding sequences supports the idea that the genetic code is, at least in part, a stereochemical residue of the most easily isolated RNA-amino acid binding structures.

Base Sequence↗

Amino acid substitution during functionally constrained divergent evolution of protein sequences.

In aligning homologous protein sequences, it is generally assumed that amino acid substitutions subsequent in time occur independently of amino acid substitutions previous in time, i.e. that patterns of mutation are similar at low and high sequence divergence. This assumption is examined here and shown to be incorrect in an interesting way. Separate mutation matrices were constructed for aligned protein sequence pairs at divergences ranging from 5 to 100 PAM units (point accepted mutations per 100 aligned positions). From these, the corresponding log-odds (Dayhoff) matrices, normalized to 250 PAM units, were constructed. The matrices show that the genetic code influences accepted point mutations strongly at early stages of divergence, while the chemical properties of the side chains dominate at more advanced stages.

Amino Acid Sequence↗

Fine structure of RNA codewords recognized by bacterial, amphibian, and mammalian transfer RNA.

Nucleotide sequences of 50 RNA codons recognized by amphibian and mammalian liver transfer RNA preparations were determined and compared with those recognized by Escherichia coli transfer RNA. Almost identical translations were obtained with transfer RNA from guinea pig liver, Xenopus laevis liver (South African clawed toad), and E. coli. However, guinea pig and Xenopus transfer RNA differ markedly from E. coli transfer RNA in relative response to certain trinucleotides. Transfer RNA from mammalian liver, amphibian liver, and amphibian muscle respond similarly to trinucleotide codons. Thus the genetic code is essentially universal, but transfer RNA from one organism may differ from that of another in relative response to some codons.

Amino Acid Sequence↗

A theory of the origin of life.

Life on Earth is essentially nucleic acids (NAs) influencing peptide synthesis such that NA replication is favored. It is proposed that the ability to synthesize polypeptides evolved gradually - one peptide bond at a time. The proposed evolution of the peptide synthesis apparatus begins with a 'transfer NA' (tNA) which catalyzes the transfer of activated amino acids to accessible amino groups in its environment. The resulting 'capped molecules' (with single amino acid 'caps') in turn favor NA replication. The proposed evolution of the peptide synthesis apparatus from the tNA onward is characterized by a progressive increase in the number of amino acids per cap: two tNAs jointly produce a 'dipeptide cap', three tNAs jointly produce a 'tripeptide cap', etc. Messenger NAs evolve because they can specify the composition and sequence order of the peptide caps. Lastly, ribosomal NAs evolve. The origin, expansion, and standardization of the genetic code are discussed. It is proposed that the presence triplet code evolved by a process of codon length refinement, and the originally codons of varying lengths were allowable, as were unassigned bases between codons. An environmental supply of activated compounds for early evolving entities is proposed. An 'environmental retention and redistribution process' is proposed to have acted as a functional substitute for the cell wall and cell division of early evolving entities.

Amino Acids↗

The code within the codons.

For the first time it is shown that each of the three codon bases has a general correlation with a different, predictable amino acid property, depending on position within the codon. In addition to the previously recognized link between the mid-base and the hydrophobic-hydrophilic spectrum, we show that, with the exception of G, the first base is generally invariant within a synthetic pathway. G--coded amino acids show a different order, being found only at the head of the synthetic pathways. The redundancy of the nature of the third base has a previously unrecognised relationship with molecular weight. The bases U and A (transversions) are associated with the most sharply defined or opposite states in both the first and second position, C somewhat less so or intermediate, anf G neutral. The apparently systematic nature of these relationships has profound implications for the origin of the genetic code. It appears to be the remains of the first language of the cell, predating the tRNA/ribosome system, persisting with remarkably little change at a deeper level of organisation than the codon language.

Amino Acid Sequence↗