Search PubMed⌕ Search

SEARCH · Search PubMed

Results for “Genetic code”

Search indexed PubMed citations on genomics, clinical trials, systematic reviews and public health. Explore titles, authors and supplied subject terms, then open the PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,315 records · Page 73Linked to original sources

Identification of the tubulin gene family and sequence determination of one beta-tubulin gene in a cold-poikilotherm protozoan, the antarctic ciliate Euplotes focardii.

Four different tubulin genes were identified in the somatic nucleus (macronucleus) of Euplotes focardii, a strictly cold-adapted, Antarctic ciliate: one of 1,800 bp for alpha-tubulin and three of 2,150, 1,900, and 1,600 bp, respectively, for beta-tubulin. Preliminarily analysed for restriction fragment length polymorphisms, these genes showed remarkable differences in organisation from tubulin genes of other ciliates which live in temperate areas and were analysed in parallel with E. focardii. The complete coding sequence of the 1,600 bp beta-tubulin gene was then determined and shown to contain unique structural features of potential importance for E. focardii microtubule organization and activity. Of eight unique substitutions detected, seven were concentrated in the large amino terminal domain of the molecule that directly interacts with the carboxy terminal region of alpha-tubulin for heterodimer formation. Sequence analysis of the cloned gene revealed, in addition, a potential new exception in the use of the genetic code by ciliates. A TAG codon was aligned in correspondence with Trp-21 which is strictly conserved in every tubulin sequence so far determined.

Amino Acid Sequence↗

Unfinished business.

The past 40 years have witnessed the development to maturity of molecular biology from its beginnings and determined attempts to apply this discipline in medicine, particularly to genetics, haematology, immunology and cancer research. However, although we have accumulated a detailed understanding of the genetic code and the mechanisms of the central dogma of molecular biology and have developed exquisitely sensitive and precise technologies, most recently represented by DNA recombinant techniques, many of the fundamental biological problems of 40 years ago still remain unsolved today. Some of these are listed and discussed.

Allergy and Immunology↗

The origin of the protein synthesis mechanism.

The origin and development of the protein synthesis mechanism is considered in four successive steps. The genetic code is supposed to be controlled by the relative amount (availability) of various amino acids and nucleotides on the one hand, and utility on each amino acid in the polypeptide. on the other hand. Thus, more simple (inutile) and abundant amino acids tended to correspond to codons which were rich in the less frequent base species, G and C. Features of primitive tRNA in the discrimination of amino acid are discussed. Primitive tRNA is proposed to have a discriminator site for amino acid and, separated from it, an anticodon site for interaction with nucleotides. A hypothetical course of subdivision of various nucleic acid species is proposed. In the scheme, mRNA and ribosomal RNA (rRNA) were derived from more primitive insoluble RNA. DNA appeared in the late, not first, step of the development. Several other aspects of evolutionary development of the whole protein synthesis mechanism, e.g., role of the discriminator site on primitive tRNA, modification and subdivision of code catalogue into a more precise specification of amino acids, and possible primordial interactions between tRNA and tRNA-binding sites on insoluble rRNA, are discussed.

Amino Acyl-tRNA Synthetases↗

RecA protein promotes strand exchange with DNA substrates containing isoguanine and 5-methyl isocytosine.

The Escherichia coli RecA protein pairs homologous DNA molecules and promotes DNA strand exchange in vitro. We have examined DNA strand exchange between a 70 nucleotide ssDNA fragment and a 40 bp duplex, in which all G and C residues (at 18 positions distributed throughout the 40 bp exchanged region) were replaced with the nonstandard nucleosides 2'-deoxyisoguanosine (iG) and 2'-deoxy-5-methylisocytidine (MiC), respectively. We demonstrate that the nonstandard oligonucleotides are substrates for the RecA protein, permitting DNA strand exchange in vitro at a rate and efficiency comparable to exchange with normal DNA substrates. This observation provides an expanded experimental basis for discussions of potential roles for iG and MiC in a genetic code. Experiments of this type also provide another avenue for exploring RecA-facilitated DNA pairing mechanisms.

5-Methylcytosine↗

NEXUS: an extensible file format for systematic information.

NEXUS is a file format designed to contain systematic data for use by computer programs. The goals of the format are to allow future expansion, to include diverse kinds of information, to be independent of particular computer operating systems, and to be easily processed by a program. To this end, the format is modular, with a file consisting of separate blocks, each containing one particular kind of information, and consisting of standardized commands. Public blocks (those containing information utilized by several programs) house information about taxa, morphological and molecular characters, distances, genetic codes, assumptions, sets, trees, etc.; private blocks contain information of relevance to single programs. A detailed description of commands in public blocks is given. Guidelines are provided for reading and writing NEXUS files and for extending the format.

Animals↗

A basis for new approaches to the chemotherapy of AIDS: novel genes in HIV-1 potentially encode selenoproteins expressed by ribosomal frameshifting and termination suppression.

Several previously unnoticed genes in the human immunodeficiency virus type 1 (HIV-1), potentially encoding selenoproteins, have been discovered by analyzing the genomic RNA structure and its relation to novel open reading frames. We have found a number of new potential RNA pseudoknots, including one in the long terminal repeat, several that coincide with highly conserved enzyme active site sequences in the pol coding region, and one in the env coding region. These pseudoknots can potentially direct the synthesis of selenocysteine (SeC) containing--1 frameshift fusion proteins. This is possible because we have found potential SeC insertion sequences (SECIS) in the RNA of HIV and other retroviruses; such structures are known to be necessary and sufficient for the incorporation of SeC at UGA "stop" codons anywhere in a eukaryotic mRNA. In several locations, UGA codons in the -1 reading frame are highly conserved across a broad spectrum of primate immunodeficiency viruses. Due to the degeneracy of the genetic code, this conservation cannot be explained by evolutionary selection of the pol gene protein sequence alone. Such observations, combined with the conservation of the associated reading frames, strongly suggest that these are real genes, and thus that the pseudoknots are also real. A protease pseudoknot-directed -1 frameshift fusion protein contains a highly conserved SeC codon and has significant similarities to a number of DNA binding proteins, including papillomavirus E2 proteins, suggesting it may be a virally encoded repressor of HIV transcription when cleaved by protease from the rest of the gag-pol gene product. A reverse transcriptase (RT) frameshift fusion protein replaces the RT active site with a highly conserved SeC-containing module. An integrase frameshift fusion protein contains the N-terminal integrase DNA-binding domain and a potential ATP-binding "GKS" motif; it has significant similarities to several helicases, but no SeC codons. A potential frameshift fusion protein from env has one SeC codon, but not in a highly conserved position. SeC incorporation could extend the nef gene product by 33 residues through the C-terminal UGA codon without frameshifting, potentially leading to substantial SeC utilization in infected cells.(ABSTRACT TRUNCATED AT 400 WORDS)

Acquired Immunodeficiency Syndrome↗

Multiple LTR-retrotransposon families in the asexual yeast Candida albicans.

We have begun a characterization of the long terminal repeat (LTR) retrotransposons in the asexual yeast Candida albicans. A database of assembled C. albicans genomic sequence at Stanford University, which represents 14.9 Mb of the 16-Mb haploid genome, was screened and >350 distinct retrotransposon insertions were identified. The majority of these insertions represent previously unrecognized retrotransposons. The various elements were classified into 34 distinct families, each family being similar, in terms of the range of sequences that it represents, to a typical Ty element family of the related yeast Saccharomyces cerevisiae. These C. albicans retrotransposon families are generally of low copy number and vary widely in coding capacity. For only three families, was a full-length and apparently intact retrotransposon identified. For many families, only solo LTRs and LTR fragments remain. Several families of highly degenerate elements appear to be still capable of transposition, presumably via trans-activation. The overall structure of the retrotransposon population in C. albicans differs considerably from that of S. cerevisiae. In that species, retrotransposon insertions can be assigned to just five families. Most of these families still retain functional examples, and they generally appear at higher copy numbers than the C. albicans families. The possibility that these differences between the two species are attributable to the nonstandard genetic code of C. albicans or the asexual nature of its genome is discussed. A region rich in retrotransposon fragments, that lies adjacent to many of the CARE-2/Rel-2 sub-telomeric repeats, and which appears to have arisen through multiple rounds of duplication and recombination, is also described.

Amino Acid Sequence↗

Real-time measurement of multiple intramolecular distances during protein folding reactions: a multisite stopped-flow fluorescence energy-transfer study of yeast phosphoglycerate kinase.

Understanding the set of rules which dictate how the primary amino acid sequence determines tertiary structure is an unsolved problem in biophysics. If it were possible to simultaneously measure all of the intramolecular distances in a protein (in real time) during a folding reaction, the "second" genetic code problem would be solved. Regrettably, no such technique currently exists. As a first step toward this goal, an optical distance assay system has been developed for a two-domain protein, yeast phosphoglycerate kinase (PGK), using Förster resonance energy transfer [Lillo, M. P., et al. (1997) Biochemistry 36, 11261-11272]. In this study, real-time stopped-flow distance changes are measured using six unique pairs of donor/acceptor fluorescent labels strategically placed throughout the tertiary structure of PGK. These multiple donor/acceptor sites were genetically engineered into PGK by cysteine substitution mutagenesis followed by extrinsic labeling with fluorescent probes, 5-[[[(2-iodoacetyl)amino]ethyl]amino]naphthalenesulfonic acid (as a donor) and 5-iodoacetamidofluorescein (acceptor). The unfolding of PGK is found to be a sequential multistep process (native --> I1 --> I2 --> unfolded) with rate constants of 0.30, 0.16, and 0.052 s-1, respectively (from native to unfolded). Unique to this unfolding study, six intramolecular distance vectors have been resolved for both the I1 and I2 states. With this distance information, it is shown that the transition from the native to I1 state can be modeled as a large hinge-bending motion, in which both domains "swing away" from each other by about 15 A. As the domains move apart, the carboxyl-terminal domain rotates almost 90 degrees about the hinge region connecting the two domains. It is also shown that the amino-terminal domain remains intact during the native --> I1 transition, consistent with our previous site-specific tryptophan fluorescence anisotropy stopped-flow study [Beechem, J. M., et al. (1995) Biochemistry 34, 13943-13948]. Future experiments are proposed which will attempt to resolve in detail the unfolding/refolding transitions in this protein with a resolution of approximately 5-10 A.

Energy Transfer↗

[Natural history of molecular pathology].

The author recalls the main stages of biosynthesis of proteins based on molecular pathology : transcription of DNA into messenger RNA, the genetic code, translation of protein molecules and regulation of their synthesis. Examples are then given of qualitative and quantitative disturbances. Qualitative disturbances with synthesis of abnormal proteins, or protein diseases of protein structure, together with their consequences. Quantitative disorders, with modified synthesis of normal proteins, result very often from abnormalities of structural genes, but also from abnormalities of transcription or translation. The author considers, in conclusion, pathological situations based on molecular abnormalities for mutations occur at random and it may be possible to correct them by genetic manipulation.

Amino Acid Sequence↗

transAlign: using amino acids to facilitate the multiple alignment of protein-coding DNA sequences.

BACKGROUND: Alignments of homologous DNA sequences are crucial for comparative genomics and phylogenetic analysis. However, multiple alignment represents a computationally difficult problem. For protein-coding DNA sequences, it is more advantageous in terms of both speed and accuracy to align the amino-acid sequences specified by the DNA sequences rather than the DNA sequences themselves. Many implementations making use of this concept of "translated alignments" are incomplete in the sense that they require the user to manually translate the DNA sequences and to perform the amino-acid alignment. As such, they are not well suited to large-scale automated alignments of large and/or numerous DNA data sets. RESULTS: transAlign is an open-source Perl script that aligns protein-coding DNA sequences via their amino-acid translations to take advantage of the superior multiple-alignment capabilities and speed of an amino-acid alignment. It operates by translating each DNA sequence into its corresponding amino-acid sequence, passing the entire matrix to ClustalW for alignment, and then back-translating the resulting amino-acid alignment to derive the aligned DNA sequences. In the translation step, transAlign determines the optimal orientation and reading frame for each DNA sequence according to the desired genetic code. It also checks for apparent frame shifts in the DNA sequences and can handle frame-shifted sequences in one of three ways (delete, align as amino acids regardless, or profile align as DNA). As a set of comparative benchmarks derived from six protein-coding genes for mammals shows, the strategy implemented in transAlign always improves the speed and usually the apparent accuracy of the alignment of protein-coding DNA sequences. CONCLUSION: transAlign represents one of few full and cross-platform implementations of the concept of translated alignments. Both the advantages accruing from performing a translated alignment and the suite of user-definable options available in the program mean that transAlign is ideally suited for large-scale automated alignments of very large and/or very numerous protein-coding DNA data sets. However, the good performance offered by the program also translates to the alignment of any set of protein-coding sequences. transAlign, including the source code, is freely available at http://www.tierzucht.tum.de/Bininda-Emonds/ (under "Programs").

Algorithms↗

The mitochondrial genome of the fission yeast Schizosaccharomyces pombe. The cytochrome b gene has an intron closely related to the first two introns in the Saccharomyces cerevisiae cox1 gene.

The DNA sequence of the cob region of the Schizosaccharomyces pombe mitochondrial DNA has been determined. The cytochrome b structural gene is interrupted by an intron of 2526 base-pairs, which has an open reading frame of 2421 base-pairs in phase with the upstream exon. The position of the intron differs from those found in the cob genes of Saccharomyces cerevisiae, Aspergillus nidulans or Neurospora crassa. The Sch. pombe cob intron has the potential of assuming an RNA secondary structure almost identical to that proposed for the first two cox1 introns (group II) in S. cerevisiae and the p1-cox1 intron in Podospora anserina. It has most of the consensus nucleotides in the central core structure described for this group of introns and its comparison with other group II introns allows the identification of an additional conserved nucleotide stretch. A comparison of the predicted protein sequences of group II intronic coding regions reveals three highly conserved blocks showing pairwise amino acid identities of 34 to 53%. These regions comprise over 50% of the coding length of the intron but do not include the 5' region, which has strong secondary structural features. In addition to the potential intron folding, long helical structures involving repetitive sequences can be formed in the flanking cob exon regions. A comparison of the Sch. pombe cytochrome b sequence with those available from other organisms indicates that Sch. pombe is evolutionarily distant from both budding yeasts and filamentous fungi. As was seen for the Sch. pombe cox1 gene (Lang, 1984), the cob exons are translated using the universal genetic code and this distinguishes Sch. pombe mitochondria from all other fungal and animal mitochondrial systems.

Amino Acid Sequence↗

The advent of DNA databanks: implications for information privacy.

Genetic identification tests -- better known as DNA profiling -- currently allow criminal investigators to connect suspects to physical samples retrieved from a victim or the scene of a crime. A controversial yet acclaimed expansion of DNA analysis is the creation of a massive databank of genetic codes. This Note explores the privacy concerns arising out of the collection and retention of extremely personal information in a central database. The potential for unauthorized access by those not investigating a particular crime compels the implementation of national standards and stringent security measures.

Confidentiality↗

Cerebellar malformations: some pathogenetic considerations.

1) Destructive processes are responsible for most cases of cerebellar microgyria of the trabecular pattern. Erosion and subsequent fusion of the folia produce the disorganized pattern in which the various cellular elements retain their noraml relationship and are capable of normal maturation. Intrauterine infection is responsible for most cases; the evidence is conclusive in some cases, presumptive in others. 2) Faulty genetic coding, as illustrated by the trisomies, may lead to formation of heterotopias. The primitive cells aggregating around the dentate nucleus should be interpreted as matrix cells and not as cells of the external granular layer. Cortical heterotopias with attempted internal organisation also occur; their origin is obscure. The unusual, possibly unique, transposition of the internal granular and Purkinje cell layers observed in one case may be ascribed to faulty formation of the Bergmann glia by analogy with the weaver mouse. 3) It is impossible at present to disentangle the role of genetic and environmental factors in the pathogenesis of the hysraphic malformations. It is possible, however, that defective fusion of the intraventricular cerebellar primordium plays a part in the development of the Dandy-Walker malformation, of midine cerebellar clefts in some cases of occipital encephalocele, and of extra-axial ependymal cysts of the posterior fossa.

Brain Diseases↗

A class of edit kernels for SVMs to predict translation initiation sites in eukaryotic mRNAs.

The prediction of translation initiation sites (TISs) in eukaryotic mRNAs has been a challenging problem in computational molecular biology. In this paper, we present a new algorithm to recognize TISs with a very high accuracy. Our algorithm includes two novel ideas. First, we introduce a class of new sequence-similarity kernels based on string editing, called edit kernels, for use with support vector machines (SVMs) in a discriminative approach to predict TISs. The edit kernels are simple and have significant biological and probabilistic interpretations. Although the edit kernels are not positive definite, it is easy to make the kernel matrix positive definite by adjusting the parameters. Second, we convert the region of an input mRNA sequence downstream to a putative TIS into an amino acid sequence before applying SVMs to avoid the high redundancy in the genetic code. The algorithm has been implemented and tested on previously published data. Our experimental results on real mRNA data show that both ideas improve the prediction accuracy greatly and that our method performs significantly better than those based on neural networks and SVMs with polynomial kernels or Salzberg kernels.

Algorithms↗

Polymorphism, recombination and alternative unscrambling in the DNA polymerase alpha gene of the ciliate Stylonychia lemnae (Alveolata; class Spirotrichea).

DNA polymerase alpha is the most highly scrambled gene known in stichotrichous ciliates. In its hereditary micronuclear form, it is broken into >40 pieces on two loci at least 3 kb apart. Scrambled genes must be reassembled through developmental DNA rearrangements to yield functioning macronuclear genes, but the mechanism and accuracy of this process are unknown. We describe the first analysis of DNA polymorphism in the macronuclear version of any scrambled gene. Six functional haplotypes obtained from five Eurasian strains of Stylonychia lemnae were highly polymorphic compared to Drosophila genes. Another incompletely unscrambled haplotype was interrupted by frameshift and nonsense mutations but contained more silent mutations than expected by allelic inactivation. In our sample, nucleotide diversity and recombination signals were unexpectedly high within a region encompassing the boundary of the two micronuclear loci. From this and other evidence we infer that both members of a long repeat at the ends of the loci provide alternative substrates for unscrambling in this region. Incongruent genealogies and recombination patterns were also consistent with separation of the two loci by a large genetic distance. Our results suggest that ciliate developmental DNA rearrangements may be more probabilistic and error prone than previously appreciated and constitute a potential source of macronuclear variation. From this perspective we introduce the nonsense-suppression hypothesis for the evolution of ciliate altered genetic codes. We also introduce methods and software to calculate the likelihood of hemizygosity in ciliate haplotype samples and to correct for multiple comparisons in sliding-window analyses of Tajima's D.

Animals↗

A block coding method that leads to significantly lower entropy values for the proteins and coding sections of Haemophilus influenzae.

A simple statistical block code in combination with the LZW-based compression utilities gzip and compress has been found to increase by a significant amount the level of compression possible for the proteins encoded in Haemophilus influenzae, the first fully sequenced genome. The method yields an entropy value of 3.665 bits per symbol (bps), which is 0.657 bps below the maximum of 4.322 bps and an improvement of 0.452 bps over the best known to date of 4.118 bps using Matsumoto, Sadakane, and Imai's lza-CTW algorithm. Calculations based on a compact inverse genetic code show that the genome has a maximum entropy of 1.757 bps for the coding regions, with a possibly lower actual entropy. These results hint at the existence of hitherto unexplored redundancies that do not show up in Markov models and are indicative of more internal structure than suspected in both the protein and the genome.

Algorithms↗

Life before RNA.

The hypothesis that life originated and evolved from linear informational molecules capable of facilitating their own catalytic replication is deeply entrenched. However, widespread acceptance of this paradigm seems oblivious to a lack of direct experimental support. Here, we outline the fundamental objections to the de novo appearance of linear, self-replicating polymers and examine an alternative hypothesis of template-directed coding of peptide catalysts by adsorbed purine bases. The bases (which encode biological information in modern nucleic acids) spontaneously self-organize into two-dimensional molecular solids adsorbed to the uncharged surfaces of crystalline minerals; their molecular arrangement is specified by hydrogen bonding rules between adjacent molecules and can possess the aperiodic complexity to encode putative protobiological information. The persistence of such information through self-reproduction, together with the capacity of adsorbed bases to exhibit enantiomorphism and effect amino acid discrimination, would seem to provide the necessary machinery for a primitive genetic coding mechanism.

Exobiology↗

Genetic polymorphism of C3 and Bf in IgA nephropathy.

C3 and Bf alleles were examined in the general population, in 67 patients with biopsy-confirmed mesangial IgA nephropathy and 81 patients with other types of glomerulonephritis, from the Heidelberg and Leiden renal programmes respectively. In both populations, a significant excess of homozygous phenotype C3FF (3.4% in controls; 10.4% in IgA nephropathy) and a deficit of C3FS heterozygous phenotype (35.8% in controls; 19.4% in IgA nephropathy) were observed in patients with IgA nephropathy, but not in other types of glomerulonephritis. No difference of C3 gene frequencies was found. C3FF was associated with an adverse clinical outcome (a higher prevalence of renal failure and hypertension). A significant excess of Bf-F gene frequency was noted (0.20 in controls; 0.33 in IgA nephropathy). In addition, an excess of phenotype BfFF was found (none in controls; 10.4% in IgA nephropathy). BfFF homozygotes also carried a higher risk of an adverse outcome (renal failure and hypertension). The data suggest a role for genetically coded (presumably) immunological factors in the genesis and course of IgA nephropathy.

Adult↗