Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genetic code”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Mechanisms of specificity in mRNA degradation: autoregulation and cognate interactions.

Autoregulation of gene expression is a common control mechanism for a large number of transcriptional units. Cases of self-regulation of the stability of various messenger RNAs (mRNAs) are re-evaluated here, and a general hypothesis for the origins and the mechanism of this process is presented. It is proposed that post-transcriptional autoregulation and mRNA stability are closely associated processes that might represent a general class of gene regulation mechanisms, with special regard to mRNA-protein cognate interactions. Generalizing from known examples, autoregulation is here considered to induce the decay of certain messenger RNAs through a yet undiscovered mechanism. Autoregulation via cognate interactions might be the vestigial process of a primitive world, where protein-nucleic acid interactions originated. The model can therefore serve as a framework to study the origins of the genetic code in particular, and gene expression in general.

Animals↗

Scalable approaches for functional analyses of whole-genome sequencing non-coding variants.

Non-coding genetic variants outside of protein-coding genome regions play an important role in genetic and epigenetic regulation. It has become increasingly important to understand their roles, as non-coding variants often make up the majority of top findings of genome-wide association studies (GWAS). In addition, the growing popularity of disease-specific whole-genome sequencing (WGS) efforts expands the library of and offers unique opportunities for investigating both common and rare non-coding variants, which are typically not detected in more limited GWAS approaches. However, the sheer size and breadth of WGS data introduce additional challenges to predicting functional impacts in terms of data analysis and interpretation. This review focuses on the recent approaches developed for efficient, at-scale annotation and prioritization of non-coding variants uncovered in WGS analyses. In particular, we review the latest scalable annotation tools, databases and functional genomic resources for interpreting the variant findings from WGS based on both experimental data and in silico predictive annotations. We also review machine learning-based predictive models for variant scoring and prioritization. We conclude with a discussion of future research directions which will enhance the data and tools necessary for the effective functional analyses of variants identified by WGS to improve our understanding of disease etiology.

Genome-Wide Association Study↗

A complementary circular code in the protein coding genes.

Recently, shifted periodicities 1 modulo 3 and 2 modulo 3 have been identified in protein (coding) genes of both prokaryotes and eukaryotes with autocorrelation functions analysing eight of 64 trinucleotides (Arquès et al., 1995). This observation suggests that the trinucleotides are associated with frames in protein genes. In order to verify this hypothesis, a distribution of the 64 trinucleotides AAA,..., TTT is studied in both gene populations by using a simple method based on the trinucleotide frequencies per frame. In protein genes, the trinucleotides can be read in three frames: the reading frame 0 established by the ATG start trinucleotide and frame 1 (resp. 2) which is the frame 0 shifted by 1 (resp. 2) nucleotide in the 5'-3' direction. Then, the occurrence frequencies of the 64 trinucleotides are computed in the three frames. By classifying each of the 64 trinucleotides in its preferential occurrence frame, i.e. the frame associated with its highest frequency, three subsets of trinucleotides can be identified in the three frames. This approach is applied in the two gene populations. Unexpectedly, the same three subsets of trinucleotides are identified in these two gene populations: Tzero = Xzero [symbol: see text] {AAA,TTT} with Xzero = {AAC,AAT,ACC,ATC,ATT, CAG,CTC,CTG,GAA,GAC,GAG, GAT,GCC,GGC,GGT,GTA,GTC,GTT,TAC,TTC} in frame 0, T1 = X1 [symbol: see text] {CCC} in frame 1 and T2 = X2 [symbol: see text] {GGG} in frame 2, each subset Xzero, X1 and X2 having 20 trinucleotides. Surprisingly, these three subsets have five important properties: (i) the property of maximal circular code for Xzero (resp. X1, X2) allowing the automatical retrieval of frame 0 (resp. 1, 2) in any region of a protein gene model (formed by a series of trinucleotides of Xzero) without using a start codon; (ii) the DNA complementarity property C (e.g. C(AAC) = GTT): C(T0) = T0, C(T1) = T2 and C(T2) = T1 allowing the two paired reading frames of a DNA double helix simultaneously to code for amino acids; (iii) the circular permutation property P (e.g. P(AAC) = ACA): P(Xzero) = X1 and P(X1) = X2 implying that the two subsets X1 and X2 can be deduced from Xzero; (iv) the rarity property with an occurrence probability of Xzero equal to 6 x 10(-8); and (v) the concatenation property with: a high frequency (27.5%) of misplaced trinucleotides in the shifted frames, a maximum (13 nucleotides) length of the minimal window to automatically retrieve the frame and an occurrence of the four types of nucleotides in the three trinucleotides sites, in favour of an evolutionary code. In the Discussion, the identified subsets Tzero, T1 and T2 replaced in the three two-letter genetic alphabets purine/pyrimidine, amino/ceto and strong/weak interaction, allow us to deduce that the RNY model (R = purine = A or G, Y = pyrimidine = C or T, N = R or Y) (Eigen & Schuster, 1978) is the closest two-letter codon model to the trinucleotides of Tzero. Then, these three subsets are related to the genetic code. The trinucleotides of Tzero code for 13 amino acids: Ala, Asn, Asp, Gln, Glu, Gly, Ile, Leu, Lys, Phe, Thr, Tyr and Val. Finally, a strong correlation between the usage of the trinucleotides of Tzero in protein genes and the amino acid frequencies in proteins is observed as six among seven amino acids not coded by Tzero, have as expected the lowest frequencies in proteins of both prokaryotes and eukaryotes.

Amino Acids↗

Synthetic oligonucleotide probes deduced from amino acid sequence data. Theoretical and practical considerations.

Synthetic probes deduced from amino acid sequence data are widely used to detect cognate coding sequences in libraries of cloned DNA segments. The redundancy of the genetic code dictates that a choice must be made between (1) a mixture of probes reflecting all codon combinations, and (2) a single longer "optimal" probe. The second strategy is examined in detail. The frequency of sequences matching a given probe by chance alone can be determined and also the frequency of sequences closely resembling the probe and contributing to the hybridization background. Gene banks cannot be treated as random associations of the four nucleotides, and probe sequences deduced from amino acid sequence data occur more often than predicted by chance alone. Probe lengths must be increased to confer the necessary specificity. Examination of hybrids formed between unique homologous probes and their cognate targets reveals that short stretches of perfect homology occurring by chance make a significant contribution to the hybridization background. Statistical methods for improving homology are examined, taking human coding sequences as an example, and considerations of codon utilization and dinucleotide frequencies yield an overall homology of greater than 82%. Recommendations for probe design and hybridization are presented, and the choice between using multiple probes reflecting all codon possibilities and a unique optimal probe is discussed.

Amino Acid Sequence↗

A view of early cellular evolution.

Some recent puzzling data on mitochondria put in question their place on the phylogenetic tree. A hypothesis, the archigenetic hypothesis, is presented, which generally agrees with Woese-Fox's concept of the common origin of eubacteria, archaebacteria, and eukaryotic hosts. However, for the first time, a case is made for the evolution of mitochondria from the ancient predecessors of pro- and eukaryotes (protobionts), not from eubacteria. Animal, fungal, and plant mitochondria are considered to be endosymbionts derived from independent free-living cells (mitobionts), which, having arisen at different developmental stages of protobionts, retained some of their ancient primitive features of the genetic code and the transcription-translation systems. The molecular-biological, bioenergetic, and paleontological aspects of this new concept of cellular evolution are discussed.

Animals↗

A proposed model for interaction of polypeptides with RNA.

Pairs of antiparallel beta polypeptide-chain segments in known protein structures are usually observed to form right-handed double helixes with helix parameters in the same range as those of nucleic acids. We have constructed a model containing only standard bond lengths, bond angles, and dihedral angles in which such a polypeptide double helix fits precisely into the minor groove of an RNA double helix with identical helix parameters. The geometry of the RNA portion is essentially a hybrid between those of the A and A' forms. Hydrogen bonds can be made between the ribose 2'-hydroxyls and polypeptide carbonyl oxygens. Since such precise complementarity between the stable conformations of RNA and polypeptides is unlikely to be merely coincidental, we propose that it played a fundamental role in the initiation of precellular evolution. Specificially, we propose that the two double-helical structures are mutually catalytic for assembly of one another from activated precursors in the prebiotic soup, and moreover that they provide some degree of genetic coding.

Biological Evolution↗

The future is noisy: the role of spatial fluctuations in genetic switching.

A genetic switch may be realized by a certain operator sector on the DNA strand from which either genetic code, to the left or to the right of this operator sector, can be transcribed and the corresponding information processed. This switch is controlled by messenger molecules, i.e., they determine to which side the switch is flipped. Recently, it has been realized that noise plays an elementary role in genetic switching, and the effect of number fluctuations of the messenger molecules have been explored. Here we argue that the assumption of well stirredness taken in the previous models may not be sufficient to characterize the influence of noise: spatial fluctuations play a non-negligible part in cellular genetic switching processes.

Bacteriophage T4↗

Quadruplet codons: implications for code expansion and the specification of translation step size.

One of the requirements for engineering expansion of the genetic code is a unique codon which is available for specifying the new amino acid. The potential of the quadruplet UAGA in Escherichia coli to specify a single amino acid residue in the presence of a mutant tRNA(Leu) molecule containing the extra nucleotide, U, at position 33.5 of its anticodon loop has been examined. With this mRNA-tRNA combination and at least partial inactivation of release factor 1, the UAGA quadruplet specifies a leucine residue with an efficiency of 13 to 26 %. The decoding properties of tRNA(Leu) with U at position 33.5 of its eight-membered anticodon loop, and a counterpart with A at position 33.5, strongly suggest that in both cases their anticodon loop bases stack in alternative conformations. The identity of the codon immediately 5' of the UAGA quadruplet influences the efficiency of quadruplet translation via the properties of its cognate tRNA. When there is the potential for the anticodon of this tRNA to dissociate from pairing with its codon and to re-pair to mRNA at a nearby 3' closely matched codon, the efficiency of quadruplet translation at UAGA is reduced. Evidence is presented which suggests that when there is a purine base at position 32 of this 5' flanking tRNA, it influences decoding of the UAGA quadruplet.

Amino Acid Sequence↗

On concerted origin of transfer RNAs with complementary anticodons.

Pairs of antiparallely oriented consensus tRNAs with complementary anticodons show surprisingly small numbers of mispairings within the 17-bp- long anticodon stem and loop region. Even smaller such complementary distances are shown by illegitimately complementary anticodons, i.e. those with allowed pairing between G and U bases. Accordingly, we suppose that transfer RNAs have emerged concertedly as complementary strands of primordial double helix-like RNA molecules. Replication of such molecules with illegitimately complementary anticodons might generate new synonymous codons for the same pair of amino acids. Logically, the idea of tRNA concerted origin dictates very ancient establishment of direct links between anticodons and the type of amino acids with which pre-tRNAs were to be charged. More specifically, anticodons (first of all, the 2nd base) could selectively target 'their' amino acids, reaction of acylating itself being performed by another non-specific site of pre-tRNA or even by another ribozyme. In all, the above findings and speculations are consistent to the hypercyclic concept (Eigen and Schuster, 1979), and throw new light on the genetic code origin and associated problems. Also favoring this idea are data on complementary codon usage patterns in different genomes.

Amino Acids↗

Increased frequency of cysteine, tyrosine, and phenylalanine residues since the last universal ancestor.

Analysis of extant proteomes has the potential of revealing how amino acid frequencies within proteins have evolved over biological time. Evidence is presented here that cysteine, tyrosine, and phenylalanine residues have substantially increased in frequency since the three primary lineages diverged more than three billion years ago. This inference was derived from a comparison of amino acid frequencies within conserved and non-conserved residues of a set of proteins dating to the last universal ancestor in the face of empirical knowledge of the relative mutability of these amino acids. The under-representation of these amino acids within last universal ancestor proteins relative to their modern descendants suggests their late introduction into the genetic code. Thus, it appears that extant ancient proteins contain evidence pertaining to early events in the formation of biological systems.

Amino Acid Sequence↗

Genetically Trained Cellular Neural Networks.

Real-coded genetic algorithms on a parallel architecture are applied to optimize the synaptic couplings of a Cellular Neural Network for specific greyscale image processing tasks. Using supervised learning information in the fitness function, we propose the Genetic Algorithm as a general training method for Cellular Neural Networks. Copyright 1997 Elsevier Science Ltd.

Journal Article↗

Emergence of adaptable systems and evolution of a translation device.

An over-all organizational framework for the origin of life is outlined and attempts for realization are given. Evolution can be described as a process resulting in an increase of "knowledge" where knowledge is the number of carriers of genetic information discarded, on the average, until the evolutionary state under consideration is reached. A model for the evolution of a translation device, a crucial event in the origin of life, is described in detail. Aggregates of short polynucleotide strands in a hairpin conformation play a major role in this model. Experimental evidence for the selectivity of aggregation supports the idea of aggregates as error filters. Chromatographic separation as selection process during chemical evolution supports the model of the early translation device leading to the origin of the genetic code.

Amino Acids↗

The scene of a frozen accident.

It has been suggested that in vitro selection experiments can provide information not only on what might have occurred during the evolution of the RNA world, but can in fact yield insights into particular features of the RNA world. In particular, it has been suggested that the sequences of anti-amino acid aptamers can provide clues to the origin of the genetic code, and that there is a statistically significant association between motifs found in aptamers and codons. We argue that the suggested connections between modern motifs and ancient sequences are logically tenuous, and show that there is no statistically meaningful association between motifs found in aptamers and codons.

Amino Acid Motifs↗

A thermodynamic theory of codon bias in viral genes.

The relationship between degeneracy in the genetic code and the occurrence of a strong codon bias is examined, with particular reference to a group of viral genomes. The present paper shows how codon bias may have been imposed by thermodynamic considerations at the time the primitive DNA first formed in the primordial soup. Using a four-state Ising-like model with stacking interactions between successive base pairs, we show how primeval periodic DNA polymers could have arisen the remnants of which are still observed in codon biases today.

Base Composition↗

The mathematical logic of life.

Protein synthesis can be likened to a particular coded information storage, transmission and execution system. Noise, error or mutations are the essential phenomena to which a living organism is subjected. Genetic coding aims at preserving the integrity of a structure under aggression from the surroundings. It can be shown that the different amino acids translated in the proteins, except the particular case of SER, obey a logical code for optimization of resistance to mutation effects. The study of the structure of this code allows a better comprehension of the logic of life.

Escherichia coli↗

Method to determine the reading frame of a protein from the purine/pyrimidine genome sequence and its possible evolutionary justification.

The periodic variations obtained by correlating the relative positions of purines and pyrimidines (and of the four bases thymine, cytosine, adenine, and guanine) in a wide variety of genomes of wholly or partly known sequence suggest that there may be enough of an earlier comma-free coding system (i.e., only readable in one frame) still present to permit determination of the reading frame and approximate extent of the present protein coding stretches. The characteristics of these variations support the hypothesis that these primitive messages were formed of coding triplets having the form RNY (R = purine; Y = pyrimidine; and N = purine or pyrimidine). The base sequences and reading frames that have a minimal deviation from such a message are still good predictors of actual coding regions and reading frames in spite of the many mutations that have occurred since such a genetic code was last in use. In fact, the right frame for almost all the proteins in a number of viruses and various prokaryotes and eukaryotes is deduced purely from purine/pyrimidine information and not by using the normal start and stop signals.

Biological Evolution↗

tRNA structure from a graph and quantum theoretical perspective.

One of the objectives of theoretical biochemistry is to find a suitable representation of molecules allowing us to encode what we know about their structures, interactions and reactivity. Particularly, tRNA structure is involved in some processes like aminoacylation and genetic code translation, and for this reason these molecules represent a biochemical object of the utmost importance requiring characterization. We propose here two fundamental aspects for characterizing and modeling them. The first takes into consideration the connectivity patterns, i.e. the set of linkages between atoms or molecular fragments (a key tool for this purpose is the use of graph theory), and the second one requires the knowledge of some properties related to the interactions taking place within the molecule, at least in an approximate way, and perhaps of its reactivity in certain means. We used quantum mechanics to achieve this goal; specifically, we have used partial charges as a manifestation of the reply to structural changes. These charges were appropriately modified to be used as weighted factors for elements constituting the molecular graph. This new graph-tRNA context allow us to detect some structure-function relationships.

Base Sequence↗

A computer program to display codon changes caused by mutagenesis.

A FORTRAN program for displaying the correspondence between codon changes and different possible base changes is presented. Changes of both single bases and dimers are considered. The user can specify the mutagenesis spectrum. Additionally, the user can choose whether or not to consider single or double events in a codon and whether or not to consider the possibility that the change of two bases (a dimer) can overlap a codon boundary. Furthermore, a variety of ways may be chosen to display and summarize the codon changes that can result from the specified mutagenesis. A user-supplied sequence or the genetic code table can be analyzed.

Codon↗