Search PubMedSearch

Biomedical subjects

E N Trifonov

Publications and source records attributed to E N Trifonov.

At least 19 recordsLinked to original sources

Mammalian retroposons integrate at kinkable DNA sites.

Integration of retroposed RNA in mammals occurs at staggered breaks resulting from an enzyme-generated pair of nicks at opposite DNA strands, preferably within 15-16 bp. Although consensus sequences associated with the two nicks appear somewhat different from one another, both nicking sites are rich in TA, CA and TG dinucleotide steps which are known as specific DNA sites where kinks may occur under bending constraints. This suggests that during interaction with the endonucleolytic enzyme, or enzymes, DNA undergoes bending at the integration sites and kinks are formed, as initial steps in generating the nicks. Nicking at kinkable sites, particularly at TA steps, may also play a role in integration of other insertion elements.

Base Composition

Protective nucleosome centering at splice sites as suggested by sequence-directed mapping of the nucleosomes.

The characteristic AA(TT) sequence pattern of the nucleosome DNA derived earlier is used for prediction of nucleosome positions around splice junctions of eukaryotic genes. Two large datasets (2000 sequences each) were collected consisting of DNA segments with the exon/intron and intron/exon splice junctions, from various eukaryotic species. Positions of predicted nucleosomes near the junction sites were calculated. Those junctions which are found to belong to the nucleosomes, are located preferentially within a few base pairs from the midpoint of the nucleosome DNA. That is, obligatory GT- and AG-ends of the introns are more frequently located near the nucleosome dyad axis, within the best protected middle 10-15 base pairs of the nucleosome DNA. In addition, a tendency is observed for the strongest nucleosomes to form more often in the introns, in accordance with the hypothesis on the chromatin-organizing role of introns.

DNA

Sequence fossils, triplet expansion, and reconstruction of earliest codons.

mRNA sequences are known to carry a hidden periodical pattern (GCU)n, which may be considered a remnant of sequence organization of mRNA early in its evolution, dominated by codons for alanine and their point mutation derivatives. A similar pattern is characteristic of the master (consensus) tRNA sequence derived in 1981 by Eigen and Winkler-Oswatitsch. The master tRNA sequence is thought to represent one of the earliest mRNA. From analysis of literature and from our own calculations presented in this work, the (GCU)n pattern appears to be the most expandable in the norm and in disease. The speculation is put forward that (GCU)n and polyalanine have been key players at the beginning of the triplet code, and the first codons, apart from the GCU triplet, were point change derivatives of the generic triplet GCU, coding for amino acids present in the early prebiotic-biotic environment. The set of the earliest amino acids is derived on the basis of structural simplicity, presence in imitated prebiotic conditions and involvement with class II aminoacyl-tRNA synthetases. The set consists of six amino acids: Ala, Asp, Gly, Pro, Ser and Thr. All these amino acids are, indeed, encoded by the GCU triplet and its derivatives, as predicted. Thus, the pairs GCN (Ala), GAU (Asp), GGU (Gly), CCU (Pro), UCU (Ser) and ACU (Thr) can be viewed as an early triplet code.

Codon

Enhancement of the nucleosomal pattern in sequences of lower complexity.

Intuitively, the complexity of a given DNA sequence is related to the number of various superimposed biological messages it contains. Here we assess the expectation that in nucleosome DNA sequences of lower linguistic complexity, the nucleosome DNA positioning pattern would be more pronounced than in those of higher linguistic complexity. The nucleosome DNA positioning pattern is one of the weakest (highly degenerate) sequence patterns. It has been extracted recently by specially designed multiple alignment procedures. We applied the most sensitive of these procedures to nearly equal subsets of a nucleosome database separated according to linguistic complexity. The pattern extracted from the subset of the simpler nucleosome sequences not only possesses all major attributes of the known nucleosomal pattern, but is substantially stronger with respect to amplitude in comparison with the total database. This result constitutes the first demonstration that a weak pattern can be significantly enhanced by selective treatment of a lower complexity subset of the sequence ensemble under consideration.

Base Composition

Taxonomy of 5 S ribosomal RNA by the linguistic technique: probing with mitochondrial and mammalian sequences.

Linguistic similarities and dissimilarities between 5 S rRNA sequences allowed taxonomical separation of species and classes. Comparisons with the molecule from mammals distinguished fungi and plants from protists and animals. Similarities to mammalians progressively increased from protists to invertebrates and to somatic-type molecules of the vertebrates lineage. In this, deviations were detected in avian, oocyte type, and pseudogene sequences. Among bacteria, actinobacteria were most similar to the mammalians, which could be related to the high frequency of associations among members of these groups. Some archaebacterial species most similar to the mammalians belonged to the Thermoproteales and Halobacteria groups. Comparisons with the soybean mitochondrial molecule revealed high internal homogeneity among plant mitochondria. The eubacterial groups most similar to it were Thermus and Rhodobacteria gamma-1 and alpha-2. Other procedures have already indicated similarities of Rhodobacteria alpha to mitochondria but the linguistic similarities were on the average higher with the first two groups.

Animals

Segmented structure of separate and transposable DNA and RNA elements as suggested by their size distributions.

A collection of about 1000 different eukaryotic and prokaryotic DNA mobile and separate elements is compiled from literature-transposons, plasmids, extrachromosomal circular DNA, insertion sequences, as well as viral genomes and separate genome segments. Only small elements are collected, upto 2000 base pairs. Analysis of the sequence length distributions of the elements reveals that certain sizes are clearly preferred, namely those which correspond to multiples of about 345 bp in eukaryotes and multiples of about 210 bp in prokaryotes. This provides additional evidence in support of the theory (1) that segmented structure is characteristic of not only protein-coding sequences (2) but rather of genomes in general. In particular, it confirms the prediction (1) that mobile and separate elements would also be segmented.

Animals

Nucleosome DNA sequence pattern revealed by multiple alignment of experimentally mapped sequences.

Five different algorithms have been applied for detecting DNA sequence pattern hidden in 204 DNA sequences collected from the literature which are experimentally found to be involved in nucleosome formation. Each algorithm was used to perform a multiple alignment of the nucleosome DNA sequences within the window 145 nt, the size of a nucleosome core DNA. From these alignments five pairs of AA and TT dinucleotide positional frequency distributions have been computed. The frequency profiles calculated by different algorithms are rather different due to substantial noise. They, however, share several important features. Both AA and TT dinucleotide positional frequencies display periodicity with the period of 10.3(+/- 0.2) bases. TT dinucleotides appear to be distributed symmetrically relative to AA dinucleotides of the same DNA strand, with the center of symmetry at the midpoint of the nucleosome core DNA. The phase shift between the AA and TT patterns is about 6 bp. Superposition of the five pairs of the AA (TT) positional frequency profiles has produced the refined pattern, with the above features well pronounced. An interesting novel feature of the pattern is an absence of central peaks in the periodical AA and TT distributions. This may indicate that the central section of nucleosome DNA, 15 bp around the dyad axis of the nucleosome, is not bent. Positional distributions of other dinucleotides were not found in this study to be as informative as the ones for AA and TT.

Algorithms

Linguistic complexity of protein sequences as compared to texts of human languages.

A notion and a measure of linguistic complexity introduced earlier (Trifonov, 1990) were originally used for analysis of nucleotide sequences. This measure was shown to reflect multiplicity of codes (messages) of different natures superimposed in the sequences. Unlike human language texts, genetic texts are 'read' by cellular mechanisms in several different ways, each time using a different selection of the characters of the same text while skipping others (Trifonov, 1989). Human texts are read in one way only, sequentially and involving all characters (one code). The conceptual significance and essence of the idea on the multiplicity of overlapping codes in genetic sequences, as opposed to human languages, is discussed. The linguistic complexity technique allows a calculation to be made of the structural complexity of any linear sequence of characters irrespective of whether the text is cognized or presently undeciphered. The texts (sequences) are compared exclusively from the point of view of their structural complexity with no reference to the meaning of the texts which is beyond the scope of this article. Results of such a comparison of protein sequences with various texts, written in English, Italian and Welsh are presented. The human texts are found to be structurally simpler than genetic (protein) texts, reflecting, apparently, a difference in the reading modes: single code versus many codes.

Amino Acid Sequence

Applicability of the multiple alignment algorithm for detection of weak patterns: periodically distributed DNA pattern as a study case.

MOTIVATION: A nucleosome DNA positioning pattern is known to be one of the weakest (highly degenerated) patterns. The alignment procedure that has been developed recently for the extraction of such a pattern is based on a statistical matching of the sequences, and its success depends on the pattern/background ratio in the individual sequences and in the generated pattern. The heuristic nature of the method and distinctive properties of the pattern bring up the question of efficiency and sensitivity in the procedure. This paper presents a method of verification for this multiple sequence alignment algorithm. RESULTS: To verify the applicability of the multiple alignment approach, we constructed a set of sequences carrying the hidden pattern. The pattern was presented by weak ('signal') oscillations of occurrences of AA and TT dinucleotides along otherwise random sequences. Only a few dinucleotides of any given 145 base long sequence would correspond to the signal, appearing in about the same phase within the simulated periodic pattern. The novelty of our simulation approach is that we simulated a database as a whole, as opposed to simulating each sequence separately. The correlation between the hidden pattern and a sequence from the database is negligible on average, but our statistical multicycle alignment procedure produced the pattern with attributes very close to the simulated ones. The accuracy of the procedure was tested and calibrated. The presence in a typical sequence of as little as three dinucleotides corresponding to the signal is sufficient to generate (detect) the pattern hidden in a collection of 204 sequences.

Algorithms

Interfering contexts of regulatory sequence elements.

MOTIVATION: Although one would normally expect a given regulatory element to perform best when it fully matches its consensus sequence, this is generally far from being the case. Usually, almost none of the actual sites fits the consensus exactly, and some of those that do fit do not perform well. The main reason for that is the very nature of the sequences and the messages (codes) they contain. Normally, any given stretch of the sequence with one or another regulatory site not only carries this regulatory message, but several more messages of various types as well. These messages overlap with the regulatory element in such a way that the letter (base) which actually appears in any given sequence position simultaneously belongs to one or more additional codes. Apart from numerous individual codes (sequence patterns) specific for a given species or gene, there are many different general (universal) sequence codes all interacting with one another. These are the classical triplet code, DNA shape code, chromatin code, gene splicing code, modulation code and many more, including those that have not yet been discovered. Examples of overlapping of different codes and their interaction are discussed, as well as the role of degeneracy of the codes and the sequence complexity as a function of code density.

Base Sequence

Sequence sizes of eukaryotic enzymes.

We have shown in earlier studies that an appreciable fraction of proteins display sequence size periodicity with periods of approximately 123 aa and approximately 152 aa for eukaryotes and prokaryotes, respectively. For any firm conclusions to be made, the issue of possible bias due to an overabundance of some protein families should be addressed in more than one way. Here we present the size distributions for various sequence ensembles of eukaryotic enzymes that differ by level of data bank cleaning. The sequences were purged by applying several successive thresholds of relatedness irrespective of the sequence lengths. The previously observed preference to typical sizes is confirmed. Possible reasons for the observed excess of the typical size sequences are discussed.

Amino Acid Sequence

Periodic recurrence of methionines: fossil of gene fusion?

As we have recently shown, approximately 20% of proteins are made of uniform size units of approximately 123 aa for eukaryotes and approximately 152 aa for prokaryotes. Such regularity may reflect certain past events in protein evolution by fusion (molecular recombination) of a spectrum of standard-size protein-coding DNA segments--the early genes. Consequently, methionines, as start residues, would mark those locations in proteins that correspond to the DNA recombination sites--the borders between the fused genes. This positional preference of the methionines may still survive as a fossil of the early protein sequence organization. In this study we address the question how methionines are distributed in modern protein sequences. This analysis of eukaryotic sequences shows that methionine residues do preferentially appear at the positions corresponding to the multiples of the unit size, as predicted.

Biological Evolution

Segmented structure of protein sequences and early evolution of genome by combinatorial fusion of DNA elements.

A theory of an early stage of genome evolution by combinatorial fusion of circular DNA units is suggested, based on protein sequence "fossil" evidence. The evidence includes preference of protein sequence lengths for certain sizes--multiples of 123 aa for eukaryotes and multiples of 152 aa for prokaryotes. At the DNA level these sizes correspond to 350-450 base pairs--the known optimal range for DNA ring closure. The methionine residues repeatedly appear along the sequences with the same period of about 120 aa (in eukaryotes), presumably marking the sites of insertion of the early genes--rings of protein-coding DNA. No torsional constraint in this DNA results in very sharp estimate of the helical periodicity of the early DNA, indistinguishable from the experimental mean value for extant DNA. According to the combinatorial fusion theory, based on the above evidence, in the pregenomic, prerecombinational stage the genes and the noncoding sequences existed in form of autonomously replicating DNA rings of close to standard size, randomly segregating between dividing cells, like modern plasmids do. In the recombinational early genomic stage the rings started to fuse, forming larger DNA molecules consisting of several unit genes connected in various combinations and forming long protein-coding sequences (combinatorial fusion). This process, which involved, perhaps, noncoding sequences as well, eventually resulted in the formation of large genomes. The dispersed circular DNA--or, rather, evolutionarily advanced derivatives thereof--may still exist in the form of various mobile DNA elements.

Biological Evolution

Underlying order in protein sequence organization.

The idea of a possible standard modular structure of proteins has been known since 1929 when it was introduced by Svedberg. It still remains an idea with no quantitative confirmation of universality of such hypothetical organization. From a large collection of nonredundant protein sequences representing > 100 eukaryotic and prokaryotic species, we have obtained the protein sequence length distributions. Mere inspection of these distributions, as well as spectral analysis, shows that 15-30% of proteins, depending on species and sequence types, indeed appear to be made of sequence units with characteristic lengths of approximately 125 aa for eukaryotes and approximately 150 aa for prokaryotes. This underlying order in protein sequence organization is shown to be universal--that is, the weak regularity observed is not caused by a particular dominant species or protein group. Possible mechanisms are discussed that may be responsible for the observed regularity, including a hypothesis about the recombinational nature of such protein sequence organization.

Amino Acid Sequence

On the recombinational origin of protein-sequence-subunit structure.

Since 1929 the concept that proteins are built from subunits of certain standard size (Svedberg 1929) has been revisited several times, each time with a new demonstration that, indeed, there are certain preferred protein sizes. According to recent estimates the overrepresented sizes are close to multiples of 125 amino acid (aa) residues for eukaryotes and 150 residues for prokaryotes. To explain these preferences, a hypothesis is suggested, and quantitatively developed, on the recombinational nature of this regularity. The protein-coding sequences are assumed to evolve at some early stage via recombinational events--insertions of DNA circles of a certain optimal size. The contour lengths of the protein-coding DNA circles had to be simultaneously divisible by three and, to minimize torsional constraint, by the DNA helical repeat. With these two conditions satisfied, the calculated contour lengths of the DNA circles, 250-500 base pairs (bp), turn out to correspond well to known optimal DNA circularization sizes and to the predicted range of the protein sequence subunit sizes: 80-170 aa residues, which covers experimentally observed values. The subunit size is found to be strongly influenced by the helical repeat of DNA. The sizes 125 and 150 aa are derived when the corresponding helical repeats of DNA are set within fractions of promilles from the 10.54 bp/turn value. This fits to the experimentally estimated mean for natural mixed DNA sequences, 10.53-10.57 bp/turn.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence

CURVATURE: software for the analysis of curved DNA.

Software is presented to plot the sequence-dependent spatial trajectory of the DNA double helix and/or distribution of curvature along the DNA molecule. The nearest-neighbor wedge model is implemented to calculate overall DNA path using local helix parameters: helix twist angle, wedge (deflection) angle and direction (of deflection) angle. The procedures described proved to be very convenient as tools for investigation of a relationship between overall DNA curvature and its gel electrophoretic mobility. All parameters of the model had been estimated from experimental data. Using these wedge parameters the program takes, as input, any DNA sequence and calculates the likely degree of curvature at each point along the molecule. This information is displayed both graphically and in the form of simplified representations of curved double helices. The Software, CURVATURE, can thus be used to investigate possible roles of curvature in modulation of gene expression and for location of curved portions of DNA, which may play an important role in sequence-specific protein--DNA interactions.

Algorithms

Imported sequences in the mitochondrial yeast genome identified by nucleotide linguistics.

In addition to universally appearing mitochondrial (mt) genes, origins of replication and transcription start regions typical of all mt genome variants of the yeast Saccharomyces cerevisiae, the mt genomes of some of the strains contain variable sequences. These sequences are apparently largely dispensable. They are mainly composed of group-I and -II introns and intergenic open reading frames (ORFs). Many of the introns contain ORFs, some of which were shown by genetic and biochemical means to be involved in splicing and transposition of the mt introns. Some of the optional sequences are hypothesized to be mobile genetic elements. Nucleotide (nt) sequences of the mt genome of S. cerevisiae were examined by analyzing occurrences of oligodeoxyribonucleotide (oligo) 'words'. This linguistic technique had been found to be sensitive to both function and origin of the sequence [Pietrokovski et al., J. Biomol. Struct. Dyn. 7 (1990) 1251-1268]. A clear difference is found between the oligo vocabularies of the optional and basic yeast mt sequences. The difference is mainly located in protein coding segments of the optional sequences which contain conserved amino acid motifs, characteristic of intronic and intergenic ORFs. The use of nt linguistics to detect the sequence dissimilarity and its causes in yeast mitochondria provides fast and straightforward results, identifying the intronic and intergenic ORFs as DNA sequences of foreign, non-mt origin.

Amino Acid Sequence