Search PubMed⌕ Search

Biomedical subjects

Edward N Trifonov

Publications and source records attributed to Edward N Trifonov.

18 recordsLinked to original sources

Conserved sequences of prokaryotic proteomes and their compositional age.

A full repertoire of octapeptides which are present in at least 30 bacterial proteomes of total 131 currently available is computationally derived and filtered. An original search technique is used that, in terms of computational time and memory, is similar to the Suffix tree method. The presence of a given sequence in a large number of proteomes qualifies it as a conserved sequence. The larger the number of proteomes where it is found, the higher is the conservation. The concept of compositional age of the amino acid sequences ("compositional clock") is introduced for the first time. The compositional age is calculated on the basis of the consensus temporal order of appearance of amino acids in early evolution. The correlation between the compositional age and the sequence conservation is established.

Amino Acid Sequence↗

Gene splice sites correlate with nucleosome positions.

Gene sequences in the vicinity of splice sites are found to possess dinucleotide periodicities, especially RR and YY, with the period close to the pitch of nucleosome DNA. This confirms previously reported findings about preferential positioning of splice junctions within the nucleosomes. The RR and YY dinucleotides oscillate counter-phase, i.e., their respective preferred positions are shifted about half-period from one another, as it was observed earlier for AA and TT dinucleotides. Species specificity of nucleosome positioning DNA pattern is indicated by the predominant use of the periodical GG(CC) dinucleotides in human and mouse genes, as opposed to predominant AA(TT) dinucleotides in Arabidopsis and C. elegans.

Alternative Splicing↗

Sequence structure of van der Waals locks in proteins.

Recent works has suggested that proteins in early evolution have gone through a stage of closed loop elements with a typical contour size of 25-35 residues. These closed loops are still the elementary protein units to these days, and can be used to spell out protein sequence/structure relationship through a relatively small number of protein prototypes. In this study we aimed to identify the sequences that are used to lock the loop ends to one another, and to show how an extensive dictionary of such locking pairs can be created using positional correlation data from a large proteome database, and structural data from PDB databases. Such a dictionary can be used in reconstructing the evolutionary pathway the modern proteins have gone through, and in identifying closed loop elements in modern proteins with yet unknown 3D structure.

Amino Acid Motifs↗

Closed loops of TIM barrel protein fold.

The closed loops within the proteins of the TIM-barrel fold family are analyzed and compared sequence- and structure-wise. The size distribution of the closed loops of the TIM-barrels confirms universal preference to the standard size of 25-30 residues. 3D structural RMSD comparisons of the closed loops and presentation of their sequences in binary form suggest that the TIM-barrel proteins are built from descendants of several types of basic closed loop prototypes. Comparison of these prototypes points to a likely common ancestor--the alpha helix containing closed loops of 28 amino acids. The presumed ancestor is characterized by specific binary consensus sequence.

Amino Acid Sequence↗

Yeast nucleosome DNA pattern: deconvolution from genome sequences of S. cerevisiae.

Positional correlation analysis for the complete genome of Saccharomyces cerevisiae is performed with the aim to reveal possible chromatin-related sequence features. A strong periodicity with the period 10.4 bases is detected in the distance histograms for the dinucleotides AA and TT, with the characteristic decay distance of approximately 50 base pairs. The oscillations are observed as well in the distributions of other dinucleotides. However, the respective amplitudes are small, consistent with secondary effects, due to dominant periodicity of AA and TT. The observations are in accord with earlier data on the chromatin sequence periodicities and nucleosome DNA sequence patterns. The autocorrelations of AA and TT dinucleotides in yeast include also a counter-phase component. A tentative DNA sequence pattern for the yeast nucleosomes is suggested and verified by comparison of its autocorrelation plots with the respective natural autocorrelations. The nucleosome mapping guided by the pattern is in accord with experimental data on the linker length distribution in yeast.

Algorithms↗

Sequence periodicity of Escherichia coli is concentrated in intergenic regions.

BACKGROUND: Sequence periodicity with a period close to the DNA helical repeat is a very basic genomic property. This genomic feature was demonstrated for many prokaryotic genomes. The Escherichia coli sequences display the period close to 11 base pairs. RESULTS: Here we demonstrate that practically only ApA/TpT dinucleotides contribute to overall dinucleotide periodicity in Escherichia coli. The noncoding sequences reveal this periodicity much more prominently compared to protein-coding sequences. The sequence periodicity of ApC/GpT, ApT and GpC dinucleotides along the Escherichia coli K-12 is found to be located as well mainly within the intergenic regions. CONCLUSIONS: The observed concentration of the dinucleotide sequence periodicity in the intergenic regions of E. coli suggests that the periodicity is a typical property of prokaryotic intergenic regions. We suppose that this preferential distribution of dinucleotide periodicity serves many biological functions; first of all, the regulation of transcription.

Base Composition↗

Dinucleosome DNA of human K562 cells: experimental and computational characterizations.

Dinucleosome formation is the first step in the organization of the higher order chromatin structure. With the ultimate aim of elucidating the dinucleosome structure, we constructed a library of human dinucleosome DNA. The library consists of PCR-amplifiable DNA fragments obtained by treatment of nuclei of erythroid K562 cells with micrococcal nuclease followed by extraction of DNA and adaptor ligation to the blunt-ended DNA fragments. The library was then cloned using a plasmid vector and the sequences of the clones were determined. The dominating clones containing the Alu elements were removed. A total of 1002 clones, which comprised a dinucleosome database, contained 84 and 918 clones from the clones before and after removing Alu elements, respectively. Approximately 70% of the clones were between 300 and 400 bp in size and they were distributed to various locations of all chromosomes except the Y chromosome. The clones containing A(2)N(8)A(2)N(8)A(2) or T(2)N(8)T(2)N(8)T(2) sequences were classified into three types, Type I (N shape), Type II (V shape) and Type III (M shape) according to DNA curvature plots. The locations of experimentally determined curved DNA segments matched well with the calculated ones though the clones of Types I and III showed additional curved DNA segments as revealed by the curvature plots. The distributions of complementary dinucleotides in the nucleosome DNA, at the ends of the dinucleosome DNA clones, allowed us to predict the positions of the nucleosome dyad axis, and estimate the size of the nucleosome core DNA, 125nt. The distributions of AA and TT dinucleotides, as well as other RR and YY dinucleotides, showed a periodicity with an average period of 10.4 bases, close to the values observed before. Mapping of nucleosome positions in the dinucleosome database based on the observed periodicity revealed that the nucleosomes were separated by a linker of 7.5+ approximately 10 x n nt. This indicates that the nucleosome-nucleosome orientations are, typically, halfway between parallel and antiparallel. Also an important finding is that the distributions of AA/TT and other RR/YY dinucleotides, apparently, reflect both DNA curvature and DNA bendability, cooperatively contributing to the nucleosome formation.

Base Sequence↗

Hidden messages in the nef gene of human immunodeficiency virus type 1 suggest a novel RNA secondary structure.

The coexistence of multiple codes in the genome of human immunodeficiency virus type 1 (HIV-1) was analyzed. We explored factors constraining the variability of the virus genome primarily in relation to conserved RNA secondary structures overlapping coding sequences, and used a simple combination of algorithms for RNA secondary structure prediction based on the nearest-neighbor thermodynamic rules and a statistical approach. In our previous study, we applied this combination to a non- redundant data set of env nucleotide sequences, confirmed the conservative secondary structure of the rev-responsive element (RRE) and found a new RNA structure in the first conserved (C1) region of the env gene. In this study, we analyzed the variability of putative RNA secondary structures inside the nef gene of HIV-1 by applying these algorithms to a non-redundant data set of 104 nef sequences retrieved from the Los Alamos HIV database, and predicted the existence of a novel functional RNA secondary structure in the beta3/beta4 regions of nef. The predicted RNA fold in the beta3/beta4 region of nef appears in two forms with different loop sizes. The loop of the first fold consists of seven nucleotides (positions 494-500), with consensus UCAAGCU appearing in 79% of sequences. The other has a five-base loop (positions 495-499) with consensus CAAGC. The difference in size between these two loops may reflect the difference between respective counterparts in the hairpin recognition. This may also have an adaptive biological significance.

Algorithms↗

Evolutionary aspects of protein structure and folding.

The traditional reconstruction of molecular events of the past based on sequence conservation becomes very vague beyond one to two billion years ago. There are certain molecular features, however, such as polymer flexibility and loop closure, that are conserved merely because of their physical nature. This allows one to penetrate the earliest stages of protein evolution.

Amino Acid Sequence↗

Protein sequences yield a proteomic code.

Analysis of crystallized protein structures suggests that globular proteins are organized as consecutively connected units of 25-35 residues. These units are closed loops, that is returns of the polypeptide chain trajectory to a close contact with itself. This universal feature of apparently polymer-statistical nature is a basis for a principally novel view on the globular proteins as loop fold structures. The same unit size has been detected in protein sequences translated from complete prokaryotic genomes by positional autocorrelation analysis, which strongly indicates the evolutionary connection of the units. The units are further characterized by prototype sequences matching to their numerous derivatives in the translated genomes. The matches to five strongest prokaryotic prototypes and three prototypes of C. elegans are identified in the sequences of crystallized proteins, and their structures analyzed. Corresponding segments of the polypeptide chains in majority of cases form closed loops, though evolutionary fate of every prototype element is shown to be rather diverse. Then loop ends can be separated by a sequence-wise distant segments and stabilized by the spatial interactions in the context of the overall globular structure. The units belong to a presumably limited spectrum of the sequence prototypes, full repertoire of which would constitute a proteomic code.

Amino Acid Motifs↗

Spelling protein structure.

Recent sequence analysis of complete prokaryotic proteomes suggests that in early evolutionary stages proteins were rather small, of the size 25-35 amino acids. Corroborating evidence comes from protein crystal data, which indicate this size for closed loops--universal structural units of globular proteins. In the latest development we were able to derive and structurally characterize several sequence/structure prototypes apparently representing early protein units. Structurally the prototypes appear as closed loops stabilized by end-to-end van der Waals interactions. While nearly standard in size the loops are highly diverse in terms of their secondary structure. A presentation of the protein as an assembly of descendants of the prototypes, the first of its kind, is described in detail here. The sequence and structure of the ATP-binding subunit of histidine permease of S. typhimurium is shown to contain several modified copies of different prototype elements, closed loops, and, thus, can be spelled as: x-PI-x-PIV-PVI-PII-PVII-x, where PI-PVII are the prototype elements. This study sets up the basic principles for the sequence/structure prototype spelling of globular proteins.

ATP-Binding Cassette Transporters↗

RNA secondary structure and squence conservation in C1 region of human immunodeficiency virus type 1 env gene.

We have analyzed amino acid, nucleotide sequence, and RNA secondary structure variability in the env gene of human immunodeficiency virus type (HIV-1). In applying algorithms for computing optimal RNA-folding patterns to a nonredundant data set of 178 env nucleotide sequences, we found a conserved RNA stem-loop structure in the first conserved (C1) region of the env gene. This detailed examination also revealed the known secondary structure conservation of the Rev-responsive element (RRE). This finding is also supported by a higher third position conservation of the translatable reading frame along these subregions. The typical folding of the C1 region consists of two isolated stem-loop structures. These highly conserved structures are likely to have a biological function. This assumption is supported by the conservation of the third position along the coding region of these structures. The third position retains a conservation level above what would be statistically expected.

Algorithms↗

What positions nucleosomes?--A model.

Here we propose a new determinant for localization of nucleosomes along genomic DNA, in addition to sequence-dependent features. The new specific class of chromatin scaling signals involves curved DNA. According to the observed positional distribution of DNA curvature, the new synchronizing signal occurs once per four nucleosomes on average. This new factor in nucleosome positioning should substantially influence the efficiency of biological reactions through regulatory factors microscopically and the entire chromatin structure through the 30 nm fiber structure macroscopically. Allocation of the new type of signals is found to be fixed evolutionarily although they could be shifted in accordance with the hierarchy of functional genomic structures.

Chromatin↗

Loop fold structure of proteins: resolution of Levinthas paradox.

According to Levinthal a protein chain of ordinary size would require enormous time to sort its conformational states before the final fold is reached. Experimentally observed time of folding suggests an estimate of the chain length for which the time would be sufficient. This estimate by order of magnitude fits to experimentally observed universal closed loop elements of globular proteins - 25-30 residues.

Models, Chemical↗

Back to units of protein folding.

In response to the criticism by A. Finkelstein (J Biomol Struct Dyn 20, 311-314, 2002) of our Communication (J Biomol Struct Dyn 20, 5-6, 2002) several issues are dealt with. Importance of the notion of elementary folding unit, its size and structure, and the necessity of further characterization of the units for the elucidation of the protein folding in vivo are discussed. The criticism (J Biomol Struct Dyn 20, 311-314, 2002) on the hierarchical protein folding is also briefly addressed.

Kinetics↗

Distribution of rare triplets along mRNA and their relation to protein folding.

It is believed that pausing during mRNA translation plays some role in ensuring proper folding of newly synthesized sections of a protein chain. Such pausing occurs when rare triplets are encountered in the mRNA, as it takes additional time for the corresponding rare species of tRNA to be delivered. To determine whether pause sites are non-randomly distributed along prokaryotic mRNA (cDNA), we have located clusters of rare triplets in cDNA sequences from 21 different bacteria. From the individual profiles of local codon frequencies calculated with various windows, the positions of the clusters of the rarest codons were taken for generation of the combined histograms of positional preferences of the pause sites. The histograms show that in the prokaryotic sequences, the pause sites are located preferentially at the start positions and at about 155 triplets from the starts. To verify the generality of these observations, the data are grouped in six independent sets about 500 sequences each, all revealing the same features. A less prominent maximum is also seen at the triplet position 75. Judging by the amplitude of the peak at 155 triplets, an optimal cluster size is estimated to equal 18 triplets. The distance 155 closely corresponds to the sizes of typical protein folds and to earlier estimated prokaryotic protein sequence segments. This supports the suggestion of a role for translation pausing in the cotranslational folding of protein domains. The profiles of rare codons in mRNA can serve in the detection or prediction of boundaries between protein domains.

Amino Acid Sequence↗

Spectral analysis of distributions: finding periodic components in eukaryotic enzyme length data.

We introduce the spectral analysis of distributions (SAD), a method for detecting and evaluating possible periodicity in experimental data distributions (histograms) of arbitrary shape. SAD determines whether a given empirical distribution contains a periodic component. We also propose a system of probabilistic mixture distributions to model a histogram consisting of a smooth background together with peaks at periodic intervals, with each peak corresponding to a fixed number of subunits added together. This mixture distribution model allows us to estimate the parameters of the data and to test the statistical significance of the estimated peaks. The analysis is applied to the length distribution of eukaryotic enzymes.

Algorithms↗

Closed loops: persistence of the protein chain returns.

It has recently been discovered that globular proteins are universally built from standard loop-n-lock units of about 30 amino acid residues. The hypothesis has been put forward on the loop stage in the protein evolution when the units were autonomous. Later they joined together making longer chains. One would expect that the early individual loop-n-lock elements might still be detected in modern protein sequences as remnants of the hypothetical 30-residue sequence prototypes. Among several strong sequence motifs, extracted from protein sequences of 23 complete bacterial proteomes, one 32-residue prototype was studied here in detail. Numerous sequence segments related to the prototype are identified in the crystal structures of proteins of a PDB_SELECT database. Analysis of the respective chain trajectories for the cases with different degrees of sequence conservation confirms that the majority of the segments correspond to the closed loops. In the evolutionary diversification of the prototypes the secondary structure yields first, while the sequence is still moderately conserved. The last feature to go is the chain return property. Apparently, the opening of the loops would severely destabilize the protein fold, which explains their conservation.

Amino Acid Motifs↗